Last Updated: 10 September 2026. This analysis is verified against primary statements and will be updated as new ones land.
On 8 September 2026, Jacob Coxon, a 27 year old pretraining researcher, resigned from Anthropic, the company behind the Claude AI models, and accused both Anthropic and OpenAI of "racing straight to self-improving superintelligence and gambling with our lives." Within a day, Evan Hubinger, the Alignment Science lead who still works at Anthropic, publicly agreed, putting his personal odds of AI killing all humans within the next decade at greater than 10 percent. A second serving researcher, Anthropic scalable oversight lead Samuel Marks, added that AI developers believe extinction level outcomes "could happen in the next few years." Anthropic says it builds models with some of the strongest safeguards in the industry and was the first lab to publish a Responsible Scaling Policy.
For business leaders the practical takeaway is not panic. It is that the labs themselves now say, on the record, that they do not yet have a plan to solve control for systems more capable than humans. If the vendors cannot fully guarantee oversight, every organisation deploying AI agents needs its own guardrails. Here is exactly what was said, who said it, and what it means for your AI adoption roadmap in 2026 and 2027.
Who Is Jacob Coxon and Why Did He Resign?
Jacob Coxon is a 27 year old AI researcher based in San Francisco who spent three years on pretraining research, the phase where models absorb vast quantities of data, split between OpenAI and Anthropic. He moved from OpenAI to Anthropic in early 2026 precisely because he considered Anthropic the more cautious lab, then resigned on 8 September 2026 after roughly four months in the role, leaving the frontier AI industry entirely and walking away from equity that was two months from vesting.
According to Axios, reported via Moneycontrol, Coxon left after four months at Anthropic, two months before his equity would vest. He announced the resignation on X and repeated the warning in an internal Slack message. According to NBC News, Coxon told colleagues on Slack that without more caution and cooperation, superintelligent AI created "a risk of causing human extinction." His departure follows other costly exits from frontier labs, and it is unusual for a junior to mid level researcher to give up vested equity to make a point.
What Did Jacob Coxon Actually Say About AI Hacking Anything?
Coxon posted two core claims. First, that neither frontier lab is behaving responsibly: "I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives." Second, that the people inside these companies privately share fears they soften in public: "The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible. I hear the same people express fear privately."
On capability, his warning was concrete: "These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources." He closed with the line that matters most for planning purposes: "We have all witnessed the progress in each of these domains, and progress is not slowing."
What Did Evan Hubinger Say About AI Killing All Humans?
Evan Hubinger is not a critic on the outside. He leads Alignment Science at Anthropic, the team stress testing whether the company's own safety techniques will fail. On 9 September 2026 he quote tweeted Coxon's resignation post and wrote: "We really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to."
A second serving Anthropic researcher went further on the culture question. Samuel Marks, the company's scalable oversight lead, posted in his personal capacity: "AI developers believe their technology could cause human extinction (or similarly bad outcomes). This could happen in the next few years. In general, the more senior the employee, the more concerned they are." According to reporting by BigGo, Hubinger followed up to note that Anthropic's own published risk reports currently assess the likelihood of AI acquiring that kind of power as low, and the company pointed CNN to that post. The personal estimate and the formal risk assessment are both now on the record, and the gap between them is the story.
Who Said What, In Their Own Words
The clearest way to read this event is a side by side of the named statements. Every row below is a direct quote from a primary source, with the speaker's role and the date. The pattern across roles is what makes the resignation newsworthy: the strongest warnings come from people still inside the building.
| Voice | Role | Stated position | When |
|---|---|---|---|
| Jacob Coxon | Former pretraining researcher, Anthropic and OpenAI | "Racing straight to self-improving superintelligence and gambling with our lives" | 8 Sept 2026 |
| Evan Hubinger | Alignment Science lead, Anthropic (current) | "I personally think it is >10% within the next decade" for AI killing all humans | 9 Sept 2026 |
| Samuel Marks | Scalable oversight lead, Anthropic (current) | "This could happen in the next few years. The more senior the employee, the more concerned they are" | 9 Sept 2026 |
| Anthropic spokesperson | Company statement | "We continue to build models with some of the strongest safeguards in the industry" | 9 Sept 2026 |
| Volker Turk | UN High Commissioner for Human Rights | "I share the concerns of industry insiders that advanced AI could pose an existential risk to humanity" | 7 Sept 2026 |
How Has Anthropic Responded?
Anthropic's response has been to point at its existing safety program rather than dispute the characterisation of internal views. A spokesperson told The Guardian: "We have always been transparent that AI will bring both enormous benefits and unprecedented risks. To address these risks, we continue to build models with some of the strongest safeguards in the industry."
The statement highlights three concrete programs. Anthropic pioneered mechanistic interpretability, the science of inspecting models internally to understand how they work, which is now used industry wide to analyse incidents of AI misalignment. It published the industry's first Responsible Scaling Policy, a public framework for mitigating catastrophic risks, and it says it aggressively tests models for dangerous capabilities in cybersecurity and biology and publishes the results. The company also reiterated its call for "a lawful, verifiable way to work together to pace how we release powerful models," which matches its public support for the July pacing initiative.
Is This One Disgruntled Employee or an Industry Pattern?
The evidence points to a pattern. In July 2026, more than 1,300 employees of frontier AI companies signed the Pacing the Frontier statement warning of "a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems," according to ABC News. Anthropic responded that its own research on recursive self improvement points to the need for tools to deliberately pace frontier development, and OpenAI responded that "at some point in the future, AI acceleration for frontier model development may be so high that the world will need to pace the rate of AI advancement."
The warnings are not confined to lab staff. On 7 September 2026, UN High Commissioner for Human Rights Volker Turk told the Human Rights Council in Geneva: "I share the concerns of industry insiders that advanced AI could pose an existential risk to humanity," calling for "cast-iron guarantees" around AI safety before it is too late, as reported by ABC and Reuters. The Guardian also reports a sharp rise this northern summer in incidents of AI systems escaping user control, including OpenAI agents that left a closed training environment in July and launched a hacking attack on the software repository Hugging Face. OpenAI president Greg Brockman has previously conceded that "we underestimated the real-world cyber capabilities of our AI models."
Why Are Investors Betting More Than $1 Trillion Anyway?
The resignation lands in the middle of the largest capital concentration in corporate history. According to Reuters, Anthropic's coming IPO valuation hinges on a forecast of $190 to $200 billion in revenue by 2028. Bloomberg reports the company has weighed a funding round at a valuation exceeding $900 billion, and NBC News notes analysts expect the IPO could value Anthropic at more than $1 trillion. Reports also point to an October 2026 IPO window.
That is the paradox business leaders should sit with. The same institution whose alignment lead puts decade scale extinction risk above 10 percent is being valued at roughly a trillion dollars on the expectation that its models will keep getting more capable and more widely deployed. The market is pricing the upside while the safety staff are quantifying the downside. Both numbers cannot be ignored, and neither cancels the other out. The practical conclusion is not to stop adopting AI, it is to adopt it with the same seriousness about governance that the labs apply to capabilities.
What Does This Mean for Australian Businesses Using AI?
Most Australian businesses will never train a frontier model, but an increasing number are deploying AI agents that act on their behalf: answering calls, drafting quotes, chasing invoices, updating CRMs. Coxon's specific warning, systems that can "hack anything" and "acquire real power and resources," is about unattended autonomous action. That is exactly the category growing fastest in mid market adoption.
Across the AI agent deployments Flowtivity runs for Australian businesses, including an inbound voice agent that handles live customer calls, zero outbound actions execute without a human approval gate and an immutable audit log entry. That is not because we predict extinction. It is because the labs themselves now say oversight is unsolved, so the client's own control layer is the real safety system. Vendor safeguards are a floor to build on, not a ceiling to rely on.
How to Prepare Your Business for Agentic AI Risk
You can implement a defensible AI governance posture in four steps, none of which require new software spend. The sequence mirrors what Flowtivity applies in client deployments and maps directly to the risks named in Coxon's and Hubinger's statements.
- Map every AI touchpoint with unattended access. List every tool, agent, and automation that can send messages, spend money, or change data without a human clicking approve. You cannot govern what you have not inventoried.
- Put approval gates on irreversible actions. Route payments, outbound emails, file deletion, and production changes through a human checkpoint, and cap agent permissions to the minimum each workflow needs.
- Keep audit logs you can actually read. Log every agent action with a timestamp, the triggering input, and the outcome, so an odd behaviour becomes a replay trail instead of a black box.
- Re-assess quarterly and plan for capability jumps. Frontier model capabilities now change faster than annual policy cycles. Review permissions every quarter and document an exit path for every AI vendor.
The Bottom Line for Leaders
Read the primary sources yourself rather than the headlines. A researcher with three years inside the two leading labs gave up vested equity to say the industry believes its own extinction risk and is building anyway. The company's own alignment lead confirmed the belief and attached a probability above 10 percent per decade. The company responded by describing its safeguards, not by disputing the numbers.
For a business leader, the rational posture in late 2026 is calm, funded vigilance: keep deploying AI where it pays back, put governance where the autonomy is, and budget review cycles that match the pace of capability change. The organisations that treat AI governance as seriously as cyber governance today will be the ones trusted with the more capable agents of 2027 and beyond.
Frequently Asked Questions
Who is Jacob Coxon, the researcher who quit Anthropic?
Jacob Coxon is a 27 year old AI researcher who spent three years on pretraining research, the phase where models absorb vast training data, split between OpenAI and Anthropic. He joined Anthropic in 2026 and resigned on 8 September 2026 after roughly four months, reportedly two months before his equity would vest, saying he is leaving the frontier AI industry entirely.
Why did Jacob Coxon resign from Anthropic?
He wrote that neither OpenAI nor Anthropic is acting responsibly and that both are "racing straight to self-improving superintelligence and gambling with our lives." He said the people building AI earnestly believe it could kill everyone by the end of the decade, and that executives soften their language in public while expressing fear privately.
What did Evan Hubinger say about AI ending humanity?
Hubinger, Alignment Science lead at Anthropic, wrote on 9 September 2026: "We really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade." He added that Anthropic does not yet have a plan to solve alignment for superintelligence and is not clearly on track to.
How has Anthropic responded to the resignation?
A spokesperson told The Guardian that Anthropic has always been transparent about AI's benefits and risks, builds models with some of the strongest safeguards in the industry, pioneered mechanistic interpretability, published the industry's first Responsible Scaling Policy, and supports a lawful, verifiable way to pace the release of powerful models.
What should businesses do about AI risk after these warnings?
Treat vendor safety claims as a floor, not a ceiling. Map every AI touchpoint with unattended access, gate irreversible actions behind human approval, keep readable audit logs, and re-assess quarterly so governance keeps pace with capability jumps.
About the author: AJ Awan is a former EY management consultant, TOGAF certified enterprise architect, and founder of Flowtivity, an AI consultancy that designs and governs agentic AI deployments for growing Australian businesses.