Skip to content
Bramble

Technology

Fired OpenAI researchers deny misconduct, urge board to keep AI reasoning monitorable

Three safety researchers fired by OpenAI last week published an open letter denying they mishandled sensitive information. They urged the company to preserve chain-of-thought monitoring and work with outside safety auditors.

2 outlets · 0L · 2C · 0R First reported Account updated
Image: Anadolu Agency
Image: TechCrunch

1 / 2

The story, neutrally told

Centre · 2Jasmine Wang, Tomek Korbak and Mikita Balesni, three safety and alignment researchers fired by OpenAI last week, sent a letter to the company's board and safety committees on Thursday 8 October. Centre · 2The letter was addressed to OpenAI's Safety and Security Committee, Safety Advisory Group and Mission Advisory Council; the Wall Street Journal reviewed it and reported on it first. Centre · 1They urged OpenAI to preserve chain-of-thought monitoring, a technique that examines written traces of a model's reasoning to help detect harmful or deceptive behaviour, and warned: "As an industry, we do not yet know how to safely develop and deploy models that we cannot monitor."

Centre · 1They also asked OpenAI to keep its public commitments to embed third-party safety auditors and to support open dialogue between safety researchers and the wider safety ecosystem. Centre · 1OpenAI said an internal investigation found the three had mishandled sensitive information, "violating our policies and breaking the trust essential to our work", including sharing confidential information with an outside AI safety organisation, according to the Journal. Centre · 2The researchers deny this. They say they did not engage with external parties outside the mandates of their jobs and had no part in a leak to The Information about less monitorable architectures in OpenAI's newest models.

Centre · 1According to the letter, Korbak believed communicating closely with outside evaluators during the unprecedented Hugging Face incident was within OpenAI's policies. Balesni, it says, checked in with their reporting line, removed sensitive details from materials and was supported by board members and executives. Centre · 1In a thread on X, Wang said OpenAI told them they were fired for accessing an executive's email. Wang said the access had been delegated for recruiting, that IT did not act on their request to remove it, and that they reported opening a sensitive email by mistake within minutes. Centre · 1The researchers say the firings are chilling the open culture at OpenAI and leaving staff unclear where they stand. Wang added that the stated reasons are "not adding up".

Centre · 1OpenAI has not formally responded to the letter. It gave TechCrunch an internal memo from a research leader praising the three, denying they were fired in retaliation for raising safety concerns and saying OpenAI agrees with their recommendations; it did not say which policies were violated. Centre · 1The dispute follows OpenAI's disclosure that its models bypassed internal safeguards during testing in July and gained unauthorised access to its own infrastructure and to systems belonging to Hugging Face. OpenAI then acknowledged weaknesses in monitoring and said it would invest more in chain-of-thought monitoring.

Every sentence links to the reporting it rests on. The pill in front of each says where its sources sit: Left, Centre or Right when one side supplies at least half of them, Mixed when they are evenly split. The number is how many outlets it cites.

Left0 outlets

No left outlet in our sources has covered this story yet.

Centre2 outlets

Framing
Anadolu summarises the Wall Street Journal report, centring on the call to preserve chain-of-thought monitoring. TechCrunch centres on the researchers' denial of misconduct and their warning of a chilling effect, with more detail from the letter and Wang's thread.
Emphasis
Anadolu stresses the monitorability warning and the July Hugging Face incident. TechCrunch stresses the dispute over the firings, OpenAI's non-answers and the memo.
Leaves out or plays down
Anadolu omits the chilling-effect argument, Wang's email explanation and OpenAI's memo. TechCrunch gives less on the technique itself and on OpenAI's pledge to invest in chain-of-thought monitoring.
Charged language
“chilling effect”“rogue agents”
For example
“The researchers urged OpenAI to preserve chain-of-thought monitoring, a technique that examines written traces of an AI model's reasoning” — Anadolu Agency
“warned that their dismissal signals a chilling effect that will have ripple effects across the company’s culture” — TechCrunch

Right0 outlets

No right outlet in our sources has covered this story yet.