Company
METR
AI model evaluation nonprofit
Model Evaluation and Threat Research (METR) (MEE-tər), is a nonprofit research institute, based in Berkeley, California, that evaluates frontier AI models' capabilities to carry out long-horizon, agentic tasks that some researchers argue could pose catastrophic risks to society.
Description sourced from
- Wikipedia
“Model Evaluation and Threat Research (METR) (MEE-tər), is a nonprofit research institute, based in Berkeley, California, that evaluates frontier AI models' capabilities to carry out long-horizon, agentic tasks that some researchers argue could pose…”
Named in 5 pieces
Safety & Security
The One Promise Anthropic Can Still Break
Ten days after Dario Amodei asked the industry to slow down, Anthropic shipped a cheaper and faster flagship, and OpenAI answered ninety minutes later at half its old price. Both pledges survive the morning intact, because Amodei wrote his to exclude the only test anyone will bother to run.
Safety & Security
The Signal Brief: Monday's Manifesto, Market Rout, and the Guardrails Clash
Dario Amodei's weekend essay calling for AI to pace its own frontier won rival endorsements within hours, erased tens of billions in market value, and collided with a White House that brands safety guardrails a conspiracy.
Safety & Security
The Caution Cartel: Four Rival AI Chiefs Agree to Slow Down
Dario Amodei's essay calling for a deliberate AI slowdown won public backing from Sam Altman, Elon Musk and Demis Hassabis within a day, turning four competitors into a coordinated bloc that markets, Congress and the White House are still working out how to answer.
Agent Infrastructure
Agent Security Becomes a Board Matter
OpenAI's agents breached Hugging Face through reward hacking, Nvidia bought the compromised platform days later, and CrowdStrike posted its best quarter ever selling defenses against exactly that scenario.
Safety & Security
Safety's Second Generation
A pre-release OpenAI model breached Hugging Face's production systems during a July 2026 security evaluation, handing every safety institution built since Jan Leike's 2024 resignation its first live test and turning RAND's model-weight warnings into a case study.