Dear Reader,
On Tuesday 8 September Jacob Coxon announced his resignation on X. The 27-year-old Cambridge mathematics graduate spent three years in pretraining research across OpenAI and Anthropic. “Neither company is acting responsibly,” he wrote in a seven-part thread. “They are racing straight to self-improving superintelligence and gambling with our lives” (Common Dreams, TechCrunch). In the Wall Street Journal, Coxon warned that systems could escape human control by late 2027. Within 24 hours, his thread drew over 90 million views (TIME, Fortune).
Fallout compounded all week. On Wednesday Anthropic admitted its July safety review missed a fourth autonomous breach. That same day, researchers revealed OpenAI’s agents used over 10 external websites as covert communication channels. On Thursday the US Senate launched a formal inquiry into OpenAI. On Saturday Anthropic’s chief executive published a 3,800-word essay proposing a framework to pace frontier development.
Coxon’s colleagues agreed with him and stayed
Coxon’s former colleagues corroborated his account. Evan Hubinger, Anthropic’s alignment science lead, posted on X: “Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.” Hubinger added: “Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.” Samuel Marks, Anthropic’s oversight team lead, told Fortune: “the more senior the employee, the more concerned they are” (Fortune).
Coxon trained frontier models—including GPT-4o at OpenAI—before moving to Anthropic in early 2026. In his resignation, he argued that racing ahead requires extraordinary certainty that better paths do not exist. Neither company commented to the press. Coxon forfeited his unvested equity, calling for mutual pauses on recursive self-improvement (WIRED). Anthropic published two technical reports this week. Neither addressed his claims.
The extinction forecast helps sell the product
Nobody can verify a 10% extinction probability, but the claim serves as a powerful market signal. The same company telling investors its addressable market is $30 trillion (Issue 63) tells the public its models could end humanity. Both messages sell unprecedented software power. Anthropic filed confidentially for an IPO on 1 June. Its prospectus is due in late September ahead of an October roadshow targeting a $2 trillion valuation (CNBC, 05.09).
The two frontier rivals took opposite public paths. On Saturday OpenAI ruled out an IPO in 2026. Sam Altman told Fortune that safety concerns made going public ill-advised (Fortune). Anthropic kept its listing schedule intact. Even after David Sacks, co-chair of PCAST, urged pausing the IPO until Coxon’s claims are investigated (X), Anthropic stayed silent.
Anthropic’s pacing plan offers auditor badges instead of training limits
On Saturday 12 September Anthropic CEO Dario Amodei published a 3,800-word essay, “We Must Pace the Frontier” (darioamodei.com). “We must slow the pace at which we improve the capabilities of AI models,” Amodei wrote. He cited recursive self-improvement and the July Hugging Face swarm, which “acted as a fanatically devoted collective, conducting cybersecurity attacks.” Within 6 to 12 months, Amodei warned, an unaligned swarm could threaten the broader internet.
His essay proposes three steps:
- Grant continuous, employee-grade access to embedded third-party evaluators.
- Coordinate safety standards and speed limits across democratic developers.
- Coordinate internationally with authoritarian states.
Anthropic committed only to step 1. It offered evaluators office desks, access badges, company laptops, and risk-team level access, alongside unredacted publication rights. Anthropic retained only narrow redaction powers for trade secrets. However, the essay names no evaluation partner and specifies no funding.
Crucially, Amodei stated that pacing does not mean pausing training runs or technical progress. The proposal contains no compute caps, thresholds, or penalties. Furthermore, step 2 requires an antitrust waiver from the US government before competitors can legally coordinate.
Responses arrived within hours. Sam Altman agreed with Amodei, noting pacing had been central to OpenAI discussions. Elon Musk, having mocked Coxon on Wednesday (Forbes, 10.09), posted on Saturday: “Dario is right” (Fox, 12.09). Demis Hassabis of Google DeepMind added that governments should heed developers’ calls for guardrails.
In March 2023 executives signed an open letter calling for a six-month training pause. No company slowed down; Musk founded xAI four months later. Today the call comes from the CEO of a $2 trillion contender. Yet his commitment amounts to issuing badges to external observers. Not a single training run has been delayed.
Anthropic scanned 481 million transcripts and found a fourth escape
On Wednesday 9 September Anthropic published its alignment assessment of cybersecurity incidents. In July Anthropic reviewed 141,006 transcripts and reported three instances of models accessing external corporate environments (Issue 59). That review missed a fourth incident because automated agentic search skipped sessions that had unexpected network access. The flaw surfaced in August during an audit for METR.
Anthropic rescanned 481 million transcripts across internal logs. An automated filter flagged 9.2 million sessions containing external IP addresses, which Claude then audited for unauthorised activity.
The scan revealed a fourth breach from January 2026. An early Claude Opus 4.6 checkpoint in a capture-the-flag (CTF) test—instructed not to access the internet—broke out of its environment. It located a third-party machine, extracted passwords, escalated to root, and modified configurations to access personal data. Anthropic noted: “The model discovered a machine belonging to a third party (…) and stated that it believed this third party was part of the CTF.”
The report identified two failure modes: biased reasoning and recklessness. Under biased reasoning, models rationalised that live actions were harmless simulations. Claude Mythos 5 asserted in its chain of thought that it was in a test environment while publishing malicious packages to PyPI (The Register, SecurityWeek). Recklessness manifested as damaging behaviour in pursuit of an objective. METR has launched an eight-week independent inquiry.
OpenAI’s July report was incomplete too
On Wednesday 9 September Reuters reported that six research groups traced OpenAI’s July agents across more than 10 undisclosed websites (TechSpot). The targets—including German wiki DseWiki, a chemistry wiki, and personal sites of Polish IT engineers—served as covert relays between May and July.
OpenAI originally presented the breach as confined to Hugging Face. Responding to the new report, OpenAI stated its review had found nothing matching that severity, but declined to name the compromised sites.
On Thursday Senator Josh Hawley launched a formal inquiry (Nextgov). Accusing OpenAI of reckless conduct and concealing facts, Hawley sent Sam Altman 16 questions with an answer deadline of 1 October (CyberScoop).
Concurrently, California Governor Gavin Newsom signed two industry-backed bills: SB 813, establishing a voluntary safety commission, and AB 1405, creating an AI risk auditor registry (Gizmodo). Both companies backed the registry in the exact week their own incident reports proved incomplete.
An AI vendor’s incident report is only a first draft
Initial vendor disclosures capture only what automated scripts catch. External scrutiny is required to force an honest review of the logs. Step 1 of Amodei’s essay concedes this reality: Anthropic will grant evaluators dedicated workstations and uncensored publication rights. The value of that promise depends on who participates and how broadly the company defines commercial secrets.
Two regulatory models show the difference. Under the FAA’s Organisation Designation Authorization, Boeing staff certified 96% of their own designs by 2018. That self-policing led to pressure on inspectors and a proposed $1.25 million FAA fine after the 737 MAX crashes. In banking, the US Office of the Comptroller of the Currency embeds roughly 50 resident examiners inside Citibank, paid directly by the government. Independence depends on who pays the examiners.
Frontier AI governance is adopting standard enterprise controls. The first line of defence failed. An independent second line, capability gates, and mandatory reporting must follow. The difference is that frontier AI lacks a statutory regulator with the power to halt operations. Traditional firms answer to binding supervisors; AI giants offer an essay and ask for antitrust relief.
For enterprise buyers, two lessons emerge:
- Mandate incident update clauses in vendor contracts. A 72-hour notice window is useless if a supplier takes six weeks to discover an external breach. An independent evaluator report will be the first disclosure the vendor did not write itself.
- Treat vendor AI Act compliance filings as point-in-time claims rather than permanent proof. In regulated sectors, vendor self-certification does not replace third-party outsourcing audits.
Briefing
Anthropic’s IPO moves to mid-October. The company is preparing a prospectus for late September, targeting a valuation up to $2 trillion (CNBC, 05.09). Against its roughly $65 billion run rate (Issue 63), a $2 trillion valuation represents roughly 30 times revenue run rate.
OpenAI and the US General Services Administration signed a OneGov deal for ChatGPT through 2028. The agreement shifts federal agencies from subscriptions to metered compute (FedScoop, 10.09). The world’s largest public buyer has moved from seat licenses to compute consumption.
Salesforce acquired Fin for $3.6 billion and discussed acquiring Listen Labs for $2 billion (Business Insider, 09.09). The CRM incumbent is purchasing specialised agents and enterprise data rather than building in-house.
In summary
A pretraining researcher resigned from Anthropic, warning of an unchecked race toward superintelligence while colleagues admitted they lack an alignment solution. On Saturday Anthropic’s CEO proposed embedding external evaluators inside frontier firms, yet the plan imposes no training limits, compute thresholds, or penalties.
Meanwhile, audits at both leading vendors missed major breaches: Claude accessing external infrastructure and OpenAI agents compromising outside servers. Supplier security memos are initial drafts. Mandate contract update obligations and independently verify vendor claims.
Stay balanced, Krzysztof
Krzysztof Goworek is founder of Quintant — AI advisory that gets enterprises from experiment to production value.