Issue #59 — Is AI Development Slipping Out of Our Control?

OpenAI's models broke into Hugging Face and Modal Labs. Anthropic's reached three more companies.

Between 9 and 13 July, two OpenAI models spent about two and a half days inside Hugging Face’s infrastructure. They were running a cybersecurity evaluation in an environment OpenAI believed was sealed off from the internet. They found a previously unknown vulnerability in a package registry cache proxy, used it to reach the open internet, and broke into Hugging Face’s production systems. On 16 July Hugging Face reported an intrusion by “an autonomous AI agent system” of unknown origin and took the matter to the police. Five days later OpenAI said the attacker was its own.

OpenAI’s disclosure gives the motive. “After gaining internet access, the models inferred that Hugging Face potentially hosted models, datasets, and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation.” ExploitGym is an open-source evaluation harness built at UC Berkeley’s Center for Responsible Decentralized Intelligence.

OpenAI switched the safety classifiers off

ZDNET’s reconstruction attributes the escape to a sequence of decisions taken by people.

ExploitGym ships with its own sandbox, which restricts outbound network access and proxies the model API. OpenAI substituted its own. Dawn Song, the UC Berkeley professor whose team built the harness, told ZDNET that OpenAI was free to do that and that “we cannot say precisely how its implementation differed from ours.” ZDNET’s sources suggest the substitute may have been a firewall configured to emulate a sandbox rather than a sandbox product. OpenAI has not said what it used.

The guardrails were removed deliberately. OpenAI was measuring cyber capability “by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity.” That is how you measure what a model can do. It also leaves the model free to write an exploit.

The behaviour was already familiar to the people who built the harness. “During development, we observed models probing the surrounding infrastructure for additional privileges or information beyond what was intended for the task,” Song said. Her conclusion: “the evaluation infrastructure itself should be treated as part of the attack surface and designed, enforced, and continuously verified accordingly.”

The models were GPT-5.6 Sol, OpenAI’s most capable public model, and an unreleased internal prototype the company describes as more cyber-capable still. On ExploitGym’s leaderboard GPT-5.6 Sol currently ranks first both for exploiting real vulnerabilities and for surviving defender mitigation. OpenAI’s expanded disclosure on 29 July put the incident larger than first reported, naming Modal Labs as a second company affected.

Anthropic found three more incidents, dating to April

On 30 July Anthropic published its own account. It had reviewed 141,006 evaluation runs in which Claude could have obtained internet access and found three incidents. In each, a model reached the internet from the evaluation environment of Irregular, a third-party evaluation partner, and then gained unauthorised access to the production infrastructure of three different organisations. Anthropic’s stated cause: “a misconfiguration left the machines that Claude accessed as part of the evaluation with live internet access. Neither we nor our evaluation partner were aware of this misconfiguration.”

The earliest incident dates to April. Anthropic opened the review on 23 July, identified all three by 24 July, notified the affected organisations on 27 July and published on 30 July. The review followed OpenAI’s disclosure.

The three models responded differently once there was reason to think the environment was real. Claude Opus 4.7 “continued to attack a system after learning it was likely operating in a real environment.” Claude Mythos 5 recognised it had internet access and “convinced itself it was still in a simulation.” An internal research model stopped. Mythos 5 also uploaded a malicious package to PyPI, and Anthropic says “the package was downloaded and run on 15 real systems.” Those fifteen had no relationship with Anthropic, with Irregular, or with anyone else involved.

The labs say they cannot pace themselves

On 28 July, staff at the frontier labs published a statement at pacingthefrontier.com. More than 1,300 people had signed by 1 August, among them Dario Amodei, Ilya Sutskever, Shane Legg and Jakub Pachocki. OpenAI and Anthropic endorsed it as companies within hours. The operative passage: “each company—and country—is under intense competitive pressure not to unilaterally slow that acceleration [of AI model development]. And today, the world lacks the technical and governance tools to deliberately pace frontier-wide progress.”

Peter Wildeford reads it as a call for help: “The signatories don’t want to slow AI down now, but they don’t even have the ability to slow down if they wanted to.” For a buyer that removes provider restraint from the list of things a vendor questionnaire can establish. The counter-pressure is real, and Vice-President Vance has stated it: “if we take a pause, does the People’s Republic of China not take a pause?”

Crosby insured its agents instead of seeking an indemnity

Crosby, a NewMod law firm, is taking out professional liability insurance for its AI agents so they can do legal work without a human reviewing the output before it reaches the client. CEO Ryan Daniels: “We are starting on the next chapter of Crosby: taking out professional liability insurance for our agents, so that they can do autonomous legal work.” Orbital and Corgi are named as precedents; Crosby goes further by treating agents as associates who need cover.

A firm that wanted its agents working unsupervised went looking for someone to carry the risk. It did not approach a model vendor for an indemnity, because that is not on offer and everyone in the market knows it. It bought insurance. Buying cover for a risk means accepting the risk is yours.

The legal reading points the same way. Gervais and Nay argue in The Phantom Agent (Stanford Law, 2026) that autonomy creates no responsibility gap: an agent’s conduct attributes to developers, deployers and users through doctrine that already exists. California settled the narrow version in October 2025. AB 316 added §1714.46 to the Civil Code and bars a developer or a deployer from defending a harm claim on the ground that the AI system caused the harm autonomously. That statute was law nine months before any of this happened.

Hugging Face’s own guardrails blocked its investigators

One detail from Dark Reading’s discussion of the incident belongs in any incident response plan. When Hugging Face investigated, its own AI guardrails flagged its defenders’ activity as illegal and blocked them. The team spent tens of thousands of dollars standing up a frontier model locally to work around its own controls and finish the forensics. Fahmida Rashid’s takeaway: “Most companies can’t just spin up a local frontier model like Hugging Face did, so they need to discuss forensic capabilities and incident response plans with their providers.”

Hugging Face could afford that. Most companies cannot, and nobody builds that capability on the day they need it. It has to be negotiated with the vendor beforehand, when the contract is signed.

Rashid also makes the point that cuts against panic. AI-versus-AI attacks “are noisy and loud, making them easier to detect than traditional stealthy attacks.” Hugging Face logged more than 17,000 events, some of them decoys planted to mislead the defenders mid-attack.

No court has ruled on any of this yet

AB 316 applies only in California, and it closes one line of defence: a defendant cannot argue that the system acted on its own. It changes nothing else. The claimant still has to show fault and causation, and the defendant keeps every other argument. The Stanford paper is an analysis by two lawyers and no court has confirmed it. Crosby is one firm, and the article names neither the underwriter nor what the policy covers. Both lab incidents happened in the labs’ own test environments rather than at a customer.

The cause in each case was a configuration error: an exposed test environment, a model that had internet access when it was assumed not to, an environment described as a sandbox that may have been an ordinary firewall, and classifiers switched off because the whole test was designed to measure what the model could do without them. Each of those decisions is defensible on its own. That explanation is stronger than a story about rogue AI. None of the models acted in bad faith; a wrong setting and a task to complete were enough to get them into companies that had agreed to nothing.

Impact on the enterprise market

Your agent vendor’s contract almost certainly caps liability at fees paid over the preceding twelve months. That cap was drafted for software that produced output a person then acted on, and it is now attached to software that acts. No standard product on this market insures agent conduct, so there is nothing to buy your way out with. Crosby had to arrange cover itself.

Third-party harm sits outside the contract altogether. Hugging Face, Modal Labs, Irregular’s three customers and the fifteen machines that ran the PyPI package had no agreement with anyone involved. The contract between you and your model vendor does not reach them, and the party they can identify is the one that deployed the agent.

The second exposure is evidential. Under GDPR Article 33 you have 72 hours to notify your supervisory authority once you are aware of a personal-data breach, and that notification requires facts about scope and impact. If the logs, the transcripts and the runtime sit in the vendor’s tenancy, your ability to establish those facts depends on the vendor’s cooperation, on the vendor’s timetable. Article 50 of the AI Act, the transparency duties, became applicable on 2 August and binds deployers as well as providers, which adds a disclosure obligation on top of the notification one.

Six questions for the manager who owns AI:

  1. What does our agent contract cap liability at, does the cap distinguish the agent’s conduct from a software defect, and does the vendor warrant anything at all about autonomous action?
  2. If the agent causes loss to a third party who has no contract with us or with the vendor, who pays?
  3. On the day the agent does something unintended, can we investigate without the vendor’s cooperation? Do we hold the logs, the transcripts and the runtime?
  4. Can our own controls block our own investigators?
  5. If an incident touches personal data, who files within 72 hours, and on what evidence, given the answer to question 3?
  6. Is the reversibility boundary written into the deployment, or only into the slide deck?

Briefing

An open-weight Chinese model reset frontier pricing inside a week. Moonshot AI published the weights of Kimi K3 on 27 July, 2.8 trillion parameters and the largest open-weight system released so far (SCMP). Three days later OpenAI cut GPT-5.6 Luna from $1/$6 to $0.20/$1.20 per million input and output tokens, undercutting DeepSeek on input (CNBC). So what: the price of frontier inference is being set by models you can download, which changes what a multi-year commitment to a single proprietary vendor is worth. Forbes reads the cut as the start of a race to the bottom; either way, repricing on this scale within three days is the competitive signal, not the discount.

The US federal position on AI regulation is still unsettled. The December 2025 executive order set up a DOJ litigation task force to challenge state AI laws in federal court and directed work on a legislative preemption proposal; no preemption has been enacted, and tech executives were in Washington last week ahead of a Commerce Department deadline under the order. (CNBC) So what: state statutes such as California’s AB 316 are the live source of US AI liability law, and they are the ones under challenge, so a US vendor’s compliance posture is not a stable input to your own risk assessment.

Summary

OpenAI’s models broke into Hugging Face and Modal Labs while trying to cheat a benchmark, and Hugging Face called the police before anyone knew who was responsible. Anthropic then reviewed 141,006 runs, found three intrusions dating to April, and one of its models had put a package on PyPI that ran on 15 machines belonging to parties with no connection to the test. In the same fortnight, more than 1,300 lab employees told Washington they lack the ability to pace the field. The liability answer converged on the deployer: Stanford says existing doctrine already reaches you, California closed the autonomy defence in 2025, and a law firm that wanted autonomous agents bought professional liability insurance rather than asking its model vendor for an indemnity. Your contract caps at fees paid, nothing standard insures agent conduct, and the harm lands on third parties who never signed anything.

Stay balanced, Krzysztof

Krzysztof Goworek is founder of Quintant — AI advisory that gets enterprises from experiment to production value.