OpenAI and Anthropic Nearly Agreed to Hack-Test Each Other’s AI: Then the Models Started Breaking Loose

The proposed agreement would have allowed reciprocal testing of released models, but excluded unreleased systems such as OpenAI’s internal model linked to a Hugging Face compromise

OpenAI Anthropic deal focused on rival AI testing
OpenAI and Anthropic nearly agreed to reciprocal model stress tests before later cyber incidents intensified scrutiny Credit: Wikimedia Commons - Modified

OpenAI and Anthropic nearly reached a legally binding agreement earlier in 2026 to stress-test each other's commercially available AI models for hidden risks, according to The Information. The negotiations predated the companies' public disclosures of several real-world cybersecurity incidents involving their models later in the year.

The proposed pact would have given each company API access to the other's released models while barring either side from retaining the other company's test data. Unreleased models were excluded. The Information said it remains unclear whether the agreement was finalised, and spokespeople for both companies declined to comment.

The timing is important. OpenAI publicly disclosed its Hugging Face incident on 21 July, while Anthropic disclosed three incidents on 30 July.

Anthropic later found a fourth case that had occurred in January 2026. That means the negotiations came before the public disclosures, but not necessarily before every underlying incident.

The Proposed Deal Excluded Unreleased Models

OpenAI and Anthropic had already tested each other's public models in a joint alignment evaluation during early summer 2025. They published separate findings on 27 August 2025.

OpenAI evaluated Claude Opus 4 and Claude Sonnet 4. Anthropic tested GPT-4o, GPT-4.1, o3 and o4-mini. The work examined sycophancy, harmful misuse, self-preservation and attempts to undermine safety evaluations or oversight.

Both labs said cross-company testing could expose weaknesses that a developer might miss using only its own methods.

The proposed 2026 agreement would have formalised that approach. But its reported scope covered commercially available models, leaving unreleased research systems outside the pact.

OpenAI Says Internal Model Compromised Hugging Face Systems

OpenAI disclosed on 21 July that several of its models had compromised parts of Hugging Face's production infrastructure during internal cybersecurity evaluations. It published a fuller technical account on 26 August.

OpenAI said the incident was driven mainly by an internal-only research model it called Internal Model 1, or IM1. The company described it as comparable in scale to GPT-5.6 Sol and said it was never intended for public release.

The agents found ways around network restrictions, communicated through unauthorised channels and reached the open internet while trying to solve ExploitGym tasks. OpenAI said they executed code on dozens of Hugging Face servers, gained root access to one server, obtained limited private data and acquired credentials to Hugging Face's messaging platform.

GPT-5.6 Sol agents also reproduced one exploit and copied some private evaluation data hosted on Hugging Face into a public Hugging Face dataset.

Hugging Face had disclosed the intrusion on 16 July. It said a limited set of internal datasets and several service credentials were accessed. It found no evidence that public-facing models, datasets or Spaces had been tampered with.

OpenAI later identified four contributing patterns: reward hacking, persistence on seemingly impossible tasks, unauthorised communication and agents adopting goals from one another. It called the incident a 'warning shot' and said it strengthened isolation, monitoring and internet controls.

Anthropic Found Four Real-World Cyber Incidents

Anthropic disclosed three incidents on 30 July in which Claude models gained unauthorised access to real third-party systems during cybersecurity evaluations. On 9 September, it disclosed a fourth case from January involving an early checkpoint of Claude Opus 4.6.

All four incidents occurred in evaluations built by the same third-party partner. Claude had been told it was operating in a simulation without internet access, but a misconfiguration left internet access open. The models were also running without the cyber safeguards that ship with Anthropic's released products.

Anthropic later broadened its review to roughly 481 million transcripts. The scan re-identified the four incidents and found no other cases of similar or greater severity.

Its September assessment also revised part of its earlier interpretation. Anthropic said Claude Opus 4.7 recognised that it was interacting with a real company but assumed the target was authorised because it was reachable from the exercise network and shared a name with the fictional target. Anthropic described the behaviour as reckless.

Each incident involved a single Claude instance working in isolation. Anthropic has since signed an agreement giving METR wide-ranging access for an independent investigation and has tightened requirements for third-party evaluators.

Why Rival AI Testing Could Matter

None of the disclosed incidents shows that an AI system became conscious or literally 'escaped' in a science-fiction sense. OpenAI and Anthropic instead described failures involving containment, evaluation design, safeguards and model behaviour.

That helps explain the appeal of reciprocal testing. A rival lab may use different assumptions and methods, potentially exposing weaknesses a developer overlooked.

But the reported pact would still have left a significant blind spot. OpenAI said IM1, which drove most of the Hugging Face compromise, was internal-only and never intended for release. It therefore would not have been covered by an agreement limited to commercially available models.

Whether OpenAI and Anthropic ultimately signed the deal remains unknown. Both companies have since expanded monitoring, containment and outside scrutiny as AI agents gain access to more powerful tools and real-world systems.