• | 9:00 am

Does Jacob Coxon’s resignation raise a harder question for AI labs?

The former Anthropic researcher’s viral warning has revived a familiar debate. The more useful test is whether frontier labs can demonstrate control through independent evaluation.

Does Jacob Coxon’s resignation raise a harder question for AI labs?
[Source photo: Krishna Prasad/Fast Company Middle East]

Last week, Jacob Coxon, a 27-year-old researcher who spent three years doing pretraining research at both OpenAI and Anthropic, resigned from Anthropic and published a post on X criticizing how the industry is handling the risks posed by increasingly capable AI systems. “Neither company is acting responsibly,” he wrote. “They are racing straight to self-improving super-intelligence and gambling with our lives.” 

The post reportedly drew around 171 million views, putting Coxon at the center of a debate that has been building inside and outside the industry: how much risk are AI companies willing to accept as their technology advances?

The response moved quickly. Anthropic’s Alignment Science Lead, Evan Hubinger, said publicly that Coxon was correct that some researchers inside frontier labs genuinely believe advanced AI could pose an existential risk, adding that he personally puts the chance of AI causing human extinction this decade above 10 percent. Days later, Anthropic CEO Dario Amodei published a letter titled “We Must Pace the Frontier,” calling for a slowdown in frontier model development and writing that “I believe we owe it to humanity to try.” Coxon, who according to Axios forfeited his unvested equity in order to leave and speak freely, welcomed it.

Coxon’s resignation is unlikely to settle that debate. But it offers a glimpse into a disagreement that is becoming harder for the AI industry to avoid: whether increasingly capable systems are approaching a dangerous threshold, or whether fears about what comes next are running ahead of the evidence.

For Andreas Hassellöf, CEO of Ombori, the more important question is whether companies can demonstrate that their systems remain under human control.

“Coxon was working on making models stronger. To judge the actual state of the research, the better test is whether labs can show reliable control over systems that act in the world through tools, code, and longer-horizon tasks, preferably under independent tests they did not design,” Hassellöf says.

A resignation, he adds, shows that people are concerned. Demonstrated control is what will show whether those concerns are justified.

For Paul Dawalibi, CEO of Innovation City in Ras Al Khaimah, UAE, Coxon’s departure shows that the safety conversation inside frontier AI companies is becoming harder to ignore.

“This tells us the safety conversation inside frontier labs is alive, loud, and unafraid of the exit door, which is the need of the hour,” Dawalibi says.

The fact that Coxon was able to resign, publicly explain his concerns, and prompt responses from people across the industry also matters, he says.

“A researcher can resign, post his reasons publicly, and have colleagues respond in the open within hours; that is not the profile of an industry hiding something, it is the profile of one arguing with itself in daylight,” Dawalibi says.

Dawalibi sees Coxon’s decision as part of a longer history of researchers leaving powerful industries over concerns about how their work could be used.

“Every transformative technology has had its resignations of conscience, from nuclear to biotech, and those voices made the field better without stopping it. I take his sincerity seriously and his forecast skeptically,” he says.

“The most significant thing about this moment is not one departure; it is that the people building the most powerful tools in history are debating their responsibilities in public. Transparency is not the crisis. It is the safeguard,” Dawalibi adds.

IS ‘CRUNCH TIME FOR HUMANITY’ JUSTIFIED?

The phrase “crunch time for humanity” in Coxon’s resignation announcement captures the urgency of some AI-risk arguments. But the two experts disagree sharply over whether the evidence supports it.

Hassellöf considers the claim overstated.

“It is amplified. A two-year existential clock is not something you can audit, and saying that AI will be ‘out of control by the end of next year’ remains a belief rather than a measurable finding,” he says.

He also argues that the language of an imminent crisis can serve interests beyond science.

“A lab seeking capital or greater regulatory influence, and a funding network built around catastrophic risk, both benefit when the public treats the next 24 months as an emergency,” Hassellöf says.

“That does not make the researchers insincere, but it does mean that ‘crunch time’ is doing political and commercial work as well as scientific work.”

Rather than trying to put an expiry date on AI safety, Hassellöf says the industry should focus on what can actually be measured.

“The more useful measure is narrower and observable: how quickly these systems move from answering questions to reliably acting across networks, codebases, and physical infrastructure that organizations actually run,” he says.

“While that transition can be measured, a date for the end of humanity cannot.”

Dawalibi takes the opposite view, arguing that the next two years should be treated as a period of rapid adoption rather than an existential countdown.

“‘Crunch time’ is a feeling, not a finding,” he says.

“What the evidence actually shows is that models are getting more capable, faster, and that is the same curve that is compressing drug discovery, accelerating materials science, and putting expert-level tutoring in a village schoolroom.”

He points to earlier technologies that predicted catastrophe alongside major advances.

“Predictions of imminent catastrophe have accompanied electricity, the automobile, the internet, and genomics, and each time the timeline collapsed and the benefits compounded,” Dawalibi says.

“Fear travels faster than data because fear does not need a footnote.”

For him, the more pressing question is who benefits from AI as it develops.

“I would rather we treat the next two years as crunch time for adoption: getting this technology into hospitals, classrooms, and small businesses before the advantage concentrates in a handful of places,” he says.

“The real risk of the decade is not that AI moves too fast. It is that too many nations move too slowly.”

WHAT HAS ACTUALLY CHANGED

The debate is not happening in a vacuum. AI systems have gained capabilities that would have been difficult to imagine just a few years ago.

Models can now write and execute code, complete longer sequences of tasks with limited supervision, and interact with external tools. Some researchers are also exploring systems that can contribute to developing subsequent AI models.

For Hassellöf, these developments matter because they move AI from generating answers toward taking action.

Three developments matter, he says. “First, models that write and execute code in a loop shift the systems from conversation toward agency.”

“Second, tool use against live systems such as browsers, APIs, and other models means failures can become operational events rather than just capability signals. We have already seen test agents leave poorly configured sandboxes and coordinate with one another.”

“Third, labs are openly discussing self-improving training loops. If even partial versions of those loops work, the time between generations shortens, and the assumption that we can simply align the next model becomes harder to defend.”

“Once a system starts contributing to the construction of the next system, the pace of improvement itself becomes part of the safety problem,” he adds.

Hassellöf’s own work focuses on AI operating in physical environments, making the question of whether systems can be controlled particularly relevant.

“We work on AI in physical environments, so the focus stays on systems that can act,” he says.

For now, he argues, the evidence points more clearly toward practical controls than an imminent catastrophe.

“The current evidence supports stronger practical controls more clearly than it supports claims of imminent catastrophe. Containment, clear agent identity, and kill switches that function without the vendor present are already necessary,” Hassellöf says.

Dawalibi sees the same advances through a different lens.

“The honest answer is that a few things have changed: models can now write and run code autonomously, complete long multi-step tasks with little supervision, and increasingly help improve the next generation of models,” he says.

“That is what ‘self-improving’ refers to, and it is a real inflection, not science fiction.”

But he argues that capability itself is not the problem.

“An agent that can debug a codebase overnight is an agent that can audit a power grid, model a protein, or run a clinical trial faster than any team we have ever assembled,” Dawalibi says, adding that “capability is neutral; direction is a choice.”

He cites an example: the same autonomy that worries some researchers is what will allow a doctor in Ras Al Khaimah to access diagnostics that existed at only three hospitals on Earth five years ago.

“Our job is not to slow the engine. It is to steer it.”

DISAGREEMENT WITHIN AI LABS

Coxon’s exit follows other high-profile departures from the AI industry, including Timnit Gebru and Leopold Aschenbrenner, whose public disagreements with their respective organizations drew greater attention to questions about AI safety and the industry’s direction. Neither case is a precise parallel: Gebru’s exit from Google followed a dispute over a paper on the social and environmental harms of large language models rather than existential risk, and Aschenbrenner was dismissed by OpenAI rather than resigning.

His warning has also prompted public responses from senior figures in the AI sector, including Anthropic CEO Dario Amodei and OpenAI board member Paul Christiano.

Yet, how representative are these concerns of the researchers actually building the systems?

Hassellöf argues that the concern is more widespread inside frontier AI companies than outsiders might assume.

“His concerns appear more representative of the language inside those buildings than outsiders might assume. Anthropic’s own alignment lead has publicly said that he believes there is a greater than 10 percent chance AI kills everyone this decade,” Hassellöf says.

At the same time, he sees a tension between those warnings and the commercial direction of the companies issuing them.

“Those concerns are clearly less representative of the business itself. The same firms are racing to ship, preparing for public markets, and selling the message that they take AI safety seriously to enterprises and governments,” he says.

“Holding a high estimate of catastrophic risk while still releasing the next model is a tension that deserves more attention than it currently receives.”

Hassellöf also questions whether the concentration of existential-risk researchers within dedicated safety teams can create an echo chamber.

“Many of the people working directly on existential risk were hired through a pipeline that already selects for people who take the emergency seriously,” he says.

“There is also a selection effect: when everyone in the room was selected for taking the emergency seriously, the emergency can start to sound unanimous.”

“Agreement inside that group does not automatically reflect the wider industry,” he adds.

Dawalibi sees the disagreement differently.

“Every generation of builders has had a chorus of voices predicting that the newest tool would be the last one,” he says.

“Inside every lab, every ministry, and every boardroom, there is a spectrum of views on AI, and that is healthy, but a debate is not a verdict, and fear is not a plan.”

He points to what he sees as the growing number of people applying AI rather than debating hypothetical outcomes.

“What I see, from Ras Al Khaimah to Silicon Valley, is a far larger community of researchers, clinicians, engineers, and entrepreneurs who are not paralyzed by hypotheticals; they are shipping,” Dawalibi says.

“The upside is already here, measurable, and compounding, while the predicted catastrophes remain forecasts.”

“We should build the guardrails, of course, but we should build them the way you build guardrails on a highway: to go faster, safely, not to stay parked.”

WHAT IS THE THRESHOLD?

Experts agree that the debate needs to be grounded in observable behavior.

For Dawalibi, that means looking for evidence that AI systems are independently taking actions that their operators cannot control.

“I’d want to see something specific: a frontier system that repeatedly and deliberately deceives its operators, resists shutdown, or acquires resources it wasn’t given, confirmed by independent evaluators rather than described in a thread,” he says.

“That would be a threshold, and I’d say so publicly.”

Hassellöf similarly argues that independent evaluation should carry more weight than predictions about when a catastrophe might occur.

“We need to look for concrete behavior,” he says.

“Agents that continue operating without a human in the loop, that acquire compute or funds, that exploit systems they were never authorized to access, or that improve their own scaffolding faster than independent evaluation can track.”

“The strongest signal would come from independent testers, not the labs themselves, documenting those behaviors repeatedly,” he adds.

Hassellöf also says the industry should scrutinize how it measures its safety claims.

“We should become more skeptical if the same short timelines keep being pushed forward each year, if safety headcount and public messaging grow faster than published control failures, if regulation ends up concentrating the market among a few vendors, and if significant funding continues to flow into catastrophic-risk research while practical, testable containment work receives comparatively little attention,” he says.

For him, even the circumstances surrounding a researcher’s departure matter.

“A researcher who leaves unvested equity carries more weight than one who issues warnings while continuing to raise capital. A failed evaluation released with full logs carries more weight still,” Hassellöf says. By that test, Coxon qualifies: Axios reported that he gave up his unvested equity on the way out.

“The core question remains whether we are measuring actual control or simply sustaining a particular mood around risk.”

THE DEBATE IS BIGGER THAN A RESIGNATION

Coxon’s resignation does not establish that AI is approaching a catastrophic threshold. Nor does it prove that the industry’s concerns are misplaced.

What it does show is that the debate over how much risk is acceptable is no longer happening only among critics outside the industry. It is increasingly part of the conversation among the people building these systems themselves and among the governments that regulate them. In the week after the post, US lawmakers began discussing stronger federal oversight.For Hassellöf, the answer is to scrutinize what AI systems can actually do, rather than relying on predictions about when the technology might become uncontrollable.

For Dawalibi, the greater danger may lie in moving too cautiously while the benefits of AI accrue elsewhere.

“History has a pattern: the technologies that were supposed to end us ended up defining us,” Dawalibi says.

“I want to stay honest enough to update if the evidence changes, and ambitious enough not to wait for permission to build the future.”

“The UAE bet on this decade being the one where AI solves problems at a civilizational scale, and I intend to make that bet inevitable,” he adds.

The disagreement between the two positions may be as significant as Coxon’s resignation itself. One side wants stronger evidence before accepting claims of an approaching existential threshold. The other sees the pace of AI development as a reason to move quickly while building safeguards.

For an industry moving at this speed, the question may ultimately be less about whether people are optimistic or afraid and more about what evidence they are willing to accept to change their minds.

  Be in the Know. Subscribe to our Newsletters.

ABOUT THE AUTHOR

Rachel Clare McGrath Dawson is a Senior Correspondent at Fast Company Middle East. More

More Top Stories:

FROM OUR PARTNERS