Skip to content
Live · The Signal · W37
Monday, September 14, 2026
AI Lately

AI is about people, and what they have been up to lately.

Anthropic's Alignment Reckoning

Jacob Coxon quit Anthropic warning of a reckless race to superintelligence, and four colleagues, including Alignment Science Lead Evan Hubinger, chose to stay and say he is right.

6 min read · 1,322 words · 8 sources
A hooded figure at a computer in a dark room
A hooded figure at a computer in a dark room. Photo · Pexels
We really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.
Evan Hubinger [2]

Jacob Coxon walked away from Anthropic on Sept. 9 and posted that the lab, along with OpenAI, races toward self-improving superintelligence while gambling with human lives 1. Four Anthropic scientists answered within a day, and every one of them stayed on the payroll. Evan Hubinger, who leads alignment science for the company, wrote that he personally holds the odds of AI causing human extinction above 10 percent within the next decade, a figure that matches the fear driving Coxon out the door 2. Samuel Marks, Anna Wang, and Ethan Perez each added their names to the same admission, turning a single departure into a five-person public reckoning conducted in real time on a platform built more for arguments than confessions 12. The strategic question this reckoning raises outlasts the news cycle: when the people closest to a technology say in public that it could end humanity, does staying inside change anything, or does it just make the warning easier to ignore?

A Resignation as Argument

Jacob Coxon spent three years doing pretraining research first at OpenAI, then at Anthropic, before he announced his exit on Sept. 9 in a post that read more like an indictment than a goodbye 1. His central claim lands as a direct accusation: "They are racing straight to self-improving superintelligence and gambling with our lives" 1. Plainly read, the sentence names two defendants and a shared crime. Its subtext runs deeper in the possessive: "our lives" folds Coxon, and every reader of his post, into stakeholder status inside a decision made in buildings he had already left. Quitting became his argument more than his exit, a resignation built to carry a claim as far as the platform would push it, and by Sept. 10 the post had drawn nearly 76 million views 18.

The Chorus That Chose to Stay

Evan Hubinger broke the silence first, posting within hours that "we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade" 2. The exclamation point matters more than the percentage: a researcher choosing punctuation built for enthusiasm to deliver a statement about mass death signals a kind of resigned candor, as if bluntness had become the only register left. Samuel Marks, who leads scalable oversight at the company, followed with a starker frame: "AI developers believe their technology could cause human extinction" 1. He added that he works on safety research at Anthropic because he hopes his work will reduce the chance of what he called "extinction-level bad outcomes," a sentence that reads as a mission statement and a confession in the same breath 1. Ethan Perez, who leads the alignment team, closed the loop with the shortest entry: "100% agree with him that AI poses serious risks to society, and I'm glad he's speaking out!" 1. Brevity here reads as its own kind of signal; when the alignment lead needs one sentence to concur with an extinction estimate, the concurrence carries more weight than elaboration would.

The Missing Plan

Hubinger's fuller statement needs paraphrase, since the original leans on phrasing the house style forbids: he described Anthropic as trying its best, working from behind on alignment for superintelligence, still searching for a plan that would close the distance 2. Anna Wang, a member of Anthropic's technical staff who previously worked at Google DeepMind, went further, saying a viable scientific plan for controlling recursively self-improving AI remains a work in progress across the entire field 1. She described staying inside a lab as the harder, and in her judgment the more effective, route to reducing risk from within, a calculation she called anything but an easy one 1. Marks supplied the clearest account of why researchers who believe the stakes run lethal keep showing up to work: commercial incentive competes against a belief that rivals would build faster and less carefully if the cautious labs stepped back 1. Together, the four statements sketch an organization racing to close the distance between a published safety roadmap and the pace of its own product releases 12.

A Familiar Pattern, A New Twist

Anthropic's chorus echoes a scene the industry already lived through: Jan Leike resigned from OpenAI's superalignment team in May 2024, writing that "safety culture and processes have taken a backseat to shiny products" 6. Leike left; Hubinger, Marks, Wang, and Perez stayed, and that distinction reframes the entire episode. Departure once read as the loudest tool available to a worried researcher; public agreement from people who kept their badges now reads louder, because it carries the credibility of continued proximity to the systems they describe as dangerous. Dario Amodei built Anthropic on the premise that safety work happens more effectively inside a frontier lab than outside one, and this week four of his researchers tested that premise in public, on the record, with their names attached 12. That premise now carries a public price tag: four researchers on payroll, each naming a percentage, each continuing to draw a salary from the company whose trajectory worries them most 12.

The Competitive Physics of the Race

Marks named the mechanism directly: AI developers keep building because they compete against rivals they judge less careful, a dynamic that turns caution into a competitive liability every lab claims to regret while leaving repair to whichever rival moves first 1. Chris Lehane, OpenAI's chief global affairs officer, published his own appeal to Congress the same week, arguing that "the prospect of AI-accelerated AI development demands more than voluntary commitments" and that the country needs binding, capability-based rules 7. Two labs, in the same seven days, arrived at versions of an identical diagnosis: internal caution alone struggles to outrun a race every competitor keeps running. Coordination through regulation, in that light, functions as a mechanism for slowing every runner at once, an outcome individual restraint leaves out of reach on its own 17.

What Governance Can Actually Do

Anthropic operates under a public Responsible Scaling Policy that ties model releases to capability thresholds, a framework the company built specifically to keep pace with warnings like the ones Hubinger, Marks, Wang, and Perez just issued 12. Whether that framework holds under the pressure Coxon named requires cooperation reaching past Anthropic alone, toward OpenAI, Google DeepMind, and every other lab chasing the same frontier. Governance built inside one company answers only part of the question Marks raised about incentives that operate across an entire industry. Employees speaking in public, rather than filing internal memos, function as a pressure campaign aimed outward, at regulators and rival labs, as much as inward, at Anthropic's own leadership 12. Pressure campaigns of this kind carry a cost, too: every public estimate invites scrutiny of Anthropic's own hiring, retention, and product timelines against the standard its researchers just set in view of the world 12.

By the numbers

  • Greater than 10 percent: Evan Hubinger's personal estimate of the odds AI causes human extinction within the next decade 2.
  • Three years: length of Jacob Coxon's pretraining research career across OpenAI and Anthropic before his Sept. 9 resignation 1.
  • Four: named Anthropic researchers, beyond Coxon, who spoke publicly on the extinction-risk question within a single week 12.
  • May 2024: month OpenAI's Jan Leike resigned over safety culture concerns, a comparable public moment 6.
  • Within hours: span between Coxon's resignation post and Hubinger's public reply 12.
  • 76 million: views Coxon's resignation post drew within a day of publication 8.

What to watch

Anthropic's next model release will show whether the Responsible Scaling Policy gates behavior differently now that four researchers have put a number on the stakes in public 12. Congress adjourns in December, and Lehane's parallel push for binding rules gives regulators a concrete deadline to answer the coordination problem Marks described 7. Watch whether other labs' researchers follow Hubinger's lead and attach a number to their own private estimates, a move that would turn one company's reckoning into an industry standard for disclosure.

Sources

  1. HuffPost, "More Anthropic Employees Sound The Alarm About AI Safety Risks," HuffPost, Sept. 9, 2026, https://www.huffpost.com/entry/anthropic-ai-risks_n_6aa1d186e4b086ecc55fc61d
  2. Officechai, "Anthropic Alignment Science Lead Evan Hubinger Says There's A More Than 10% Chance AI Could Kill All Humans Within Next Decade," Officechai, Sept. 9, 2026, https://officechai.com/ai/anthropic-alignment-science-lead-evan-hubinger-says-theres-a-more-than-10-chance-ai-could-kill-all-humans-within-next-decade/
  3. CBS News, "Anthropic Researcher Says More Than 10% Chance AI 'Could Kill All Humans,'" CBS News, Sept. 9, 2026, http://www.cbsnews.com/news/ai-kill-humans-anthropic-researcher-more-than-ten-percent-chance/
  4. The Hill, "Anthropic Researchers Warn AI Could Kill Humans by the End of the Decade," The Hill, Sept. 10, 2026, https://thehill.com/policy/technology/6078907-anthropic-researchers-warn-ai-extinction/
  5. Siladitya Ray, "Anthropic Alignment Lead Issues Warning About AI Killing Humans As Researcher Resigns," Forbes, Sept. 9, 2026, https://www.forbes.com/sites/siladityaray/2026/09/09/anthropic-alignment-lead-warns-ai-could-kill-all-humans-as-researcher-quits/
  6. CNN Business, "More OpenAI Drama: Exec Quits Over Concerns About Focus on Profit Over Safety," CNN, May 17, 2024, https://www.cnn.com/2024/05/17/tech/openai-exec-exits-safety-concerns
  7. KFGO, "OpenAI Pushes for Mandatory National AI Safety Requirements," KFGO, Sept. 9, 2026, https://kfgo.com/2026/09/09/openai-pushes-for-mandatory-national-ai-safety-requirements/
  8. Quartz, "Jacob Coxon Quits Anthropic Over Self-Improving AI Safety Fears," Quartz, Sept. 9, 2026, https://qz.com/anthropic-researcher-quits-self-improving-ai-safety-090926

Cite this piece

Ryan Elliott Dennis, "Anthropic's Alignment Reckoning," AI Lately, Sep 11, 2026, https://ailately.com/articles/anthropic-alignment-reckoning

Tags: AI alignment · existential risk · AI safety research · talent departures · frontier AI governance

Related, lately

More →