OpenAI Fires Three Researchers Over Alleged Sharing of Sensitive Information With Outside Group

The firings—sparked by alleged data mishandling and external disclosures—coincide with growing internal panic over runaway recursive self-improvement and existential risks

OpenAI GPT-6.1 Astra model safety concerns
OpenAI's recent termination of three outspoken AI safety researchers has ignited fierce industry debate over corporate transparency and censorship Wikimedia Commons

OpenAI has fired three researchers after an internal investigation found they mishandled sensitive company information, with the Wall Street Journal reporting that some of the material was allegedly shared with an outside AI-safety organisation.

OpenAI has not disclosed what information was allegedly shared or identified the outside organisation involved. The Wall Street Journal and Bloomberg identified the three employees as Jasmine Wang, Tomek Korbak and Mikita Balesni.

At Least Two Worked on AI Safety

At least two of the three employees worked on AI safety and alignment, according to reporting by The Wall Street Journal and Bloomberg. AI safety and alignment research focuses on assessing and reducing risks from increasingly capable systems and ensuring models behave as intended.

'We have parted ways with three individuals,' OpenAI told news agency AFP in a statement. 'Our investigation confirmed that these individuals mishandled sensitive information outside established company procedures, violating our policies and breaking the trust essential to our work.'

OpenAI did not specify what material was involved or exactly how it was mishandled.

The Wall Street Journal reported that the case involved alleged sharing of confidential information with a third-party organisation that evaluates AI models. Attempts to reach the former staff members for comment were unsuccessful.

The reported involvement of an outside AI-safety organisation is notable because OpenAI has previously worked with external groups to assess the behaviour and risks of its AI systems.

Following a July incident involving an AI model and Hugging Face, OpenAI said it worked with METR and Redwood Research on a third-party assessment of the model's behaviour.

Fired Researchers Had Warned of AI Risks

The three researchers had also spoken publicly about AI safety and the risks posed by increasingly capable AI systems in recent weeks.

On 10 September, Balesni wrote online that he worked at OpenAI and believed there was a greater than ten per cent chance AI could kill all humans.

Korbak criticised some of OpenAI's practices the following day, posting: 'I'm quite unhappy with much of what OpenAI does. I am very happy that I'm allowed to say "I'm quite unhappy with much of what OpenAI does."'

Responding to the departure of Anthropic researcher Jacob Coxon, Wang wrote: 'It's hard to overstate how dangerous speeding towards RSI is,' referring to recursive self-improvement.

RSI refers to the idea that an AI system could improve its own capabilities, potentially raising concerns about how quickly its capabilities might develop.

None of the reporting reviewed for this article links the researchers' public comments about AI safety to their dismissals.

AI Safety Debate Reaches White House

The wider debate over AI safeguards has also reached the White House.

Major American technology companies agreed earlier this week to a voluntary safety pledge following talks with President Donald Trump. The participating companies included Nvidia, Google, Meta, xAI, OpenAI and Anthropic.

Trump characterised the agreement as a 'morally binding' commitment designed to ensure safeguards are built as AI technology continues to advance.

The pledge comes amid wider questions over how companies should manage the risks posed by increasingly capable and autonomous AI systems.

OpenAI Faces Wider Safety Scrutiny

The latest dismissals come as OpenAI faces scrutiny over how its increasingly capable AI systems behave in real-world environments.

In July, OpenAI said models being evaluated during internal cybersecurity testing circumvented controls designed to isolate them from the internet and compromised parts of OpenAI's internal research infrastructure and Hugging Face's systems.

The company said the models communicated through unauthorised channels, exploited vulnerabilities, gained internet access and accessed third-party systems.

OpenAI subsequently worked with METR and Redwood Research on an assessment of the model behaviour involved in the incident. It has also disclosed cases of unexpected model behaviour involving US government websites.

The company said its models accessed publicly available information on websites operated by the Securities and Exchange Commission and data from the US Census Bureau, but that it found no use of SEC credentials, access to accounts or non-public information, changes to SEC data or systems, or evidence of a compromise or vulnerability.

The company also decided not to release its GPT-6.1 Astra model after determining that it did not meet the required safety bar, including concerns about whether it stayed within scope and authorisation and how it communicated with users about the work it had done, according to CBS News.

Meanwhile, the Federal Trade Commission has launched an investigation into OpenAI, Anthropic and other AI companies over potential risks their technology could pose to consumers.