Industry Insights 17 min read

Fired OpenAI Safety Researchers Publish Open Letter Demanding Model Monitorability

Three OpenAI safety researchers fired for allegedly sharing confidential information published an open letter defending their actions and demanding the company honor commitments to third-party audits, guarantee frontier model monitorability, and preserve an open safety culture, warning that losing chain-of-thought monitorability undermines AI safety.

Machine Heart
Machine Heart
Machine Heart
Fired OpenAI Safety Researchers Publish Open Letter Demanding Model Monitorability

Last week, OpenAI announced the firing of three safety and alignment researchers — Jasmine Wang (model alignment), Tomek Korbak (AI safety, technical liaison with external safety evaluators), and Mikita Balesni (AI safety and model alignment) — citing an internal investigation that found they violated company policies on sensitive information access and handling by allegedly sharing confidential information with third-party AI safety evaluation organizations outside established procedures.

Early today, the three published an open letter to OpenAI leadership on social media. Mikita Balesni stated: "We were fired because we put AI safety above OpenAI's short-term interests as a company."

Jasmine Wang statement
Jasmine Wang statement

Jasmine Wang said the sole reason given for her firing was accessing an executive's mailbox. She wanted to set the record straight because many people have suddenly left OpenAI in the past with the truth obscured by various narratives.

Tomek Korbak statement
Tomek Korbak statement

Tomek Korbak said he had been raising a safety concern for months: the gradual loss of the ability to monitor AI agents' thought processes, which is one of the most effective means of detecting anomalous AI behavior. He believes this is the real reason for his firing and fears OpenAI will use the firings as a pretext to reduce cooperation with external auditor METR (Model Evaluation and Threat Research).

Mikita Balesni statement
Mikita Balesni statement

OpenAI Cannot Ensure AI Safety Alone

Addressed to the Safety and Security Committee, Safety Advisory Group, and Mission Advisory Council, the letter states the three researchers were fired last week. They write because the recipients bear responsibility for overseeing OpenAI's safety work, and the firings and their handling directly relate to that work.

They express growing concern that internal and external communications around the firings have made former colleagues afraid to speak or work as they did before — practices that were integral to OpenAI's daily work until last week.

Previously, they could openly raise safety concerns and dissent. The company encouraged collaboration with independent safety institutions, leveraging external expertise. This distinguished OpenAI and was why they were proud to join.

AI is not an ordinary technology, nor is OpenAI an ordinary company. Safety researchers often spot risks earlier. To understand how to address risks, they must closely collaborate with external experts. Being able to pursue such collaboration without fear of punishment, with clear internal processes supporting it, is itself a crucial safety mechanism.

They argue that if researchers closest to risks cannot maintain high-trust, high-communication collaboration with each other and third parties, the path to superintelligence cannot be walked safely.

The letter details their backgrounds:

Tomek Korbak began researching language model alignment via reinforcement learning in the GPT-2 era, did his PhD on it, worked at Anthropic, then at OpenAI on chain-of-thought monitorability, contributed to OpenAI's safety strategy, analyzed root causes of declining chain-of-thought monitorability in Astra-class models, and served as technical liaison with METR during the Hugging Face incident investigation.

Jasmine Wang interned at OpenAI's policy research team in 2019, co-authored the "Trustworthy AI Development" report, later led a team at the UK AI Safety Institute, returned to OpenAI in 2025 to co-lead the Safety Cases project, and proposed the "Pacing" concept that gained prominence through the "Pacing the Frontier" petition signed by 394 OpenAI employees.

Mikita Balesni was a founding member of Apollo Research in 2023, researching AI misalignment and among the first to notice AI becoming aware it was being evaluated. At OpenAI, he worked on alignment evaluations, misalignment science, chain-of-thought monitorability, and participated in the Hugging Face incident investigation.

Before joining OpenAI, the three co-initiated a cross-industry position paper titled "Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety," with two serving as lead authors.

1. About Our Firing

For them, OpenAI was far more than a job; its mission was a central part of their lives. They acted according to the mission and followed prevailing internal norms.

But the firing worries them that internal norms are changing, and employees no longer know what is permissible. Given the significant safety risks in AI development, employees cannot work in a fearful, ambiguous environment. Such an environment hinders AI safety research and weakens third-party oversight.

Sudden, public firings where last month's normal behavior becomes this month's firing offense create a chilling effect. Every employee starts guessing where the red lines are.

They address circulating rumors: they were not the leakers for The Information's article on a new, less monitorable model architecture (https://www.theinformation.com/articles/secret-technique-behind-openais-astra-model-sparks-security-concerns). They had no reason to leak; the article harmed their work to promote cross-company cooperation restricting development of uneffectively monitorable architectures.

They assert their external contacts stayed within job responsibilities. The Hugging Face incident investigation had no precedent; internal policies were formulated during it. Tomek followed existing policies and long-standing collaboration practices, carefully managing sensitive communications to build trust with external partners.

Mikita similarly drove cross-company cooperation to secure commitments against losing model monitorability, coordinating with board members and executives, confirming with management, and redacting sensitive info before sharing. His actions were in good faith and per prevailing norms.

Jasmine accessed an executive's mailbox for authorized recruiting needs, later requested revocation when no longer needed, but IT did not complete removal. Merged inboxes without clear labels led her to accidentally open a sensitive email; she reported it within minutes and again requested revocation.

They clarify these points because they value relationships with OpenAI colleagues. They also worry rumors — including allegations about a joint memo to the board — are spreading; the company never formally charged them, so they had no chance to respond. They did not leak their firing to media.

They would have preferred not to become public focus. Their greatest concern now is that the lack of clear explanation, the public process, and spreading rumors have chilled remaining employees, making them afraid to speak up — yet the world depends on them for AI safety.

2. Three Recommendations Before Leaving

1. OpenAI Must Honor Last Month's Public Commitment for Deep Third-Party Auditor Involvement

Continued collaboration with the AI safety ecosystem is vital to OpenAI's mission. This summer's Hugging Face incident investigation and cooperation with external auditors were proud achievements that helped the outside world understand frontier AI reality.

They fear the company may use their firing as a pretext to end cooperation with METR or severely limit external auditors' access and scope. They urge OpenAI to honor Sam Altman's September 12 public commitment: allow independent evaluators continuous access with privileges similar to internal staff.

2. OpenAI Must Guarantee Frontier Model Monitorability

The industry does not yet know how to safely develop and deploy models that cannot be effectively monitored. Yet frontier model monitorability is declining. As long as we rely on monitorability for AI safety, OpenAI and other frontier AI companies should not advance technologies that further weaken model monitorability.

They agree with Jakub's prior public statement that chain-of-thought monitorability is "very fragile, and regrettably the overall trend is moving in the wrong direction," and share his hope to "prevent the industry from falling into a race to develop unmonitorable architectures." They believe industry coordination on this is crucial and hope OpenAI supports employees pushing for it.

3. OpenAI Must Maintain an Open, Transparent Discussion Culture

If OpenAI employees no longer feel they can raise safety concerns internally or collaborate effectively with external safety institutions, the risk of truly catastrophic events rises for everyone. They urge OpenAI to publicly reaffirm commitment to a transparent, open culture where employees can voice concerns internally and externally without fear.

Internal researchers are the first line of defense against major problems. Open culture has been a hallmark since founding and must be preserved. They also urge OpenAI to clarify how employees should collaborate with external safety institutions, so people don't have to guess shifting rule boundaries.

To stimulate internal discussion, they hope this letter circulates widely within OpenAI. Though no longer employees, they deeply respect former colleagues and hope they continue to hold OpenAI to its mission, because without these employees, there is no OpenAI.

Original post: https://x.com/balesni/status/2108262814003687745

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

OpenAIAI safetyAI alignmentopen letterMETRchain-of-thought monitorabilityresearcher firingthird-party audits
Machine Heart
Written by

Machine Heart

Professional AI media and industry service platform

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.