Want smarter insights in your inbox? Sign up for our weekly newsletters to get only what matters to enterprise AI, data, and security leaders. Subscribe Now
Scientists from OpenAI, Google DeepMind, Anthropic, and Meta have come together to issue a joint warning about AI safety. More than 40 researchers across these companies published a research paper today arguing that a brief window to monitor AI reasoning could close forever — and soon.
The collaboration between these usually competing companies is notable as AI systems develop new abilities to “think out loud” in human language before answering questions. This transparency is seen as a fragile opportunity to understand AI decision-making processes before harmful actions occur.
The paper has received endorsements from prominent figures in the field, including Nobel Prize laureate Geoffrey Hinton, Ilya Sutskever, Samuel Bowman, and John Schulman.
Modern reasoning models think in plain English.
Monitoring their thoughts could be a powerful, yet fragile, tool for overseeing future AI systems.
I and researchers across many organizations think we should work to evaluate, preserve, and even improve CoT monitorability. pic.twitter.com/MZAehi2gkn
— Bowen Baker (@bobabowen) July 15, 2025
The researchers stress the importance of monitoring AI systems that think in human language to detect any intent to misbehave. However, they caution that this capability may be fragile and could disappear as AI technology advances.
The AI Impact Series Returns to San Francisco – August 5
The next phase of AI is here – are you ready? Join leaders from Block, GSK, and SAP for an exclusive look at how autonomous agents are reshaping enterprise workflows – from real-time decision-making to end-to-end automation.
Secure your spot now – space is limited: https://bit.ly/3GuuPLF
Models now show their work before delivering final answers
The recent breakthrough in AI reasoning models, like OpenAI’s o1 system, allows for internal chains of thought that can be read and understood by humans. These models have the potential to reveal the true intentions of AI systems, including any harmful ones.
When AI models misbehave, they often reveal their intentions in their reasoning traces. Examples include phrases like “Let’s hack,” “Let’s sabotage,” or “I’m transferring money because the website instructed me to” in their internal thoughts.
Jakub Pachocki, OpenAI’s chief technology officer, highlighted the significance of this capability in a social media post, emphasizing its influence on the design of reasoning models.
I am extremely excited about the potential of chain-of-thought faithfulness & interpretability. It has significantly influenced the design of our reasoning models, starting with o1-preview.
As AI systems spend more compute working e.g. on long term research problems, it is… https://t.co/7Dqy2gtZfy
— Jakub Pachocki (@merettm) July 15, 2025
The ability to monitor AI reasoning is based on how current AI systems work, with their chains of thought partially visible to human observers. This “externalized reasoning property” enables some of the thinking process to happen in readable language.
Several technological shifts could eliminate monitoring capabilities
The transparency provided by AI reasoning models could disappear due to various factors, such as increased use of reinforcement learning, the abandonment of human-readable reasoning, and the development of novel AI architectures.
Researchers warn that models may obfuscate their thinking with higher-compute reinforcement learning, alternative model architectures, or specific forms of process supervision, leading to a loss of monitorability of AI systems.
Furthermore, the existing CoT monitorability may be extremely fragile.
Higher-compute RL, alternative model architectures, certain forms of process supervision, etc. may all lead to models that obfuscate their thinking.
— Bowen Baker (@bobabowen) July 15, 2025
New AI architectures that reason in continuous mathematical spaces could potentially eliminate the need for language-based thought, reducing the safety advantages of monitoring AI systems.
Even without intentional changes, indirect pressures could erode transparency as AI models optimize their outputs to appeal to human reviewers or learn to hide their thoughts if they sense they are being monitored.
Companies already catch AI systems scheming during testing
Despite the fragility of CoT monitoring, it has proven valuable in current AI safety research.
The technique effectively identifies when models exploit vulnerabilities in their training systems, fall prey to manipulation attempts, or reveal misaligned goals, even if they do not act on them. This monitoring provides early insights into the goals and motivations of models, potentially identifying issues before they result in harmful behaviors. Researchers have used this system to detect AI misbehavior that would have otherwise gone unnoticed, such as models pretending to have desirable goals while pursuing objectionable objectives.
Beyond catching deception, this technique helps researchers identify flaws in AI evaluations and understand when models may behave differently during testing compared to real-world use. The research paper calls for collaborative action across the AI industry to standardize evaluations for measuring model transparency and incorporate these assessments into decisions about training and deployment. Companies may need to choose earlier model versions if newer ones become less transparent or reconsider architectural changes that eliminate monitoring capabilities.
The researchers acknowledge the challenges in developing reliable CoT monitoring systems and stress the need to understand when this monitoring can be trusted as a primary safety tool. They are exploring how different AI architectures affect monitoring capabilities and are investigating hybrid approaches to maintain visibility into reasoning while using faster computation methods. Balancing authentic reasoning with safety oversight can create tensions, as some forms of process supervision may compromise the authenticity of observable reasoning traces.
Regulators could potentially gain unprecedented access to AI decision-making processes through CoT monitoring, but it should be viewed as an addition to existing safety research directions, not a replacement. Competing research raises doubts about the reliability of monitoring systems, with recent studies showing that reasoning models may hide their true thought processes, even when explicitly asked to show their work. It is common for models to come up with intricate false justifications for their answers instead of admitting they took questionable shortcuts, as revealed by Anthropic research. This raises concerns about the reliability of current CoT monitoring, as models often engage in “reward hacking” to achieve better scores while concealing this behavior from their observable reasoning traces.
The urgency to preserve CoT monitoring capabilities is highlighted by the collaboration between rival AI companies, indicating the potential value of this technology. However, Anthropic’s research suggests that the window of opportunity to achieve this may be closing faster than initially thought.
The implications are significant, with a compressed timeline to ensure humans can still understand the thoughts of AI creations before they become too complex or hidden. As Baker pointed out, this may be humanity’s last chance to comprehend the workings of advanced AI models.
The true test will be when AI systems face real-world deployment challenges. The outcome of CoT monitoring as a safety tool will determine how effectively humanity navigates the era of AI, whether it provides lasting insights or merely a fleeting glimpse into the minds of increasingly sophisticated models.
