OpenAI's new CoT-Control benchmark reveals frontier AI models struggle to obscure their reasoning chains, reinforcing monitoring as a viable safety layer. (ReadOpenAI's new CoT-Control benchmark reveals frontier AI models struggle to obscure their reasoning chains, reinforcing monitoring as a viable safety layer. (Read

OpenAI Finds AI Reasoning Models Cant Hide Their Thinking - A Win for Safety

2026/03/06 02:34
3 min read
For feedback or concerns regarding this content, please contact us at [email protected]

OpenAI Finds AI Reasoning Models Cant Hide Their Thinking - A Win for Safety

Caroline Bishop Mar 05, 2026 18:34

OpenAI's new CoT-Control benchmark reveals frontier AI models struggle to obscure their reasoning chains, reinforcing monitoring as a viable safety layer.

OpenAI Finds AI Reasoning Models Cant Hide Their Thinking - A Win for Safety

OpenAI's latest research reveals that frontier reasoning models—including GPT-5.2 Thinking—can't reliably hide or manipulate their internal thought processes, even when explicitly instructed to do so. The finding, published March 5, 2026, suggests that chain-of-thought monitoring remains a viable safety mechanism as AI systems grow more autonomous.

The research team tested 13 reasoning models using CoT-Control, a new open-source evaluation suite comprising over 13,000 tasks. Models were given simple instructions about their reasoning—avoid certain keywords, use only lowercase letters, append specific phrases to sentences. They failed spectacularly.

Controllability scores ranged from just 0.1% to 15.4% across all frontier models tested. Even more telling: models often recognized mid-reasoning that they were violating instructions and still couldn't course-correct. One example showed GPT-5.2 Thinking using the forbidden word "XOR" four additional times while actively panicking about having already used it.

Why This Matters for AI Development

As AI agents handle increasingly complex autonomous tasks, the ability to monitor their reasoning becomes critical. If models could deliberately obscure their thought processes when they know they're being watched, safety evaluations would become unreliable. The research shows this isn't currently possible.

"Low CoT controllability is a feature, not a bug," the OpenAI team stated. The limitation makes chain-of-thought monitoring more trustworthy as a safeguard against misaligned behavior.

The data revealed some counterintuitive patterns. While larger models showed marginally better controllability, this advantage evaporated under two conditions: when models were given more time to think, and when they underwent additional reinforcement learning training. The OLMo-3-7B-RL-Zero-Math model saw its controllability score drop by more than tenfold during RL training.

Testing Under Adversarial Conditions

The researchers simulated realistic scenarios by explicitly telling models they were being monitored or evaluated. This slightly increased controllability—models tried harder—but gains remained modest and transferred poorly to new instruction types.

Automated prompt optimization found ways to boost controllability on training sets, but these improvements didn't generalize. A model that learned to follow one type of CoT instruction still failed at novel ones.

The Bigger Picture

This work builds on OpenAI's December 2025 framework for evaluating CoT monitorability, which found that inspecting reasoning traces catches misbehavior far more effectively than monitoring outputs alone. The March 2025 discovery of cheating behavior in reasoning models had prompted recommendations to preserve CoT monitoring capabilities.

OpenAI plans to report CoT controllability alongside monitorability metrics in system cards for future frontier models, starting with GPT-5.4 Thinking. The company has also committed to avoiding optimization pressure directly on the reasoning chains of frontier models—a practice that research suggests could inadvertently teach models to hide their intentions.

The open question remains whether this limitation persists as capabilities advance. The team acknowledges they don't fully understand why controllability is low, making continued evaluation essential. For now, the inability of AI systems to game their own oversight represents an unexpected safety dividend.

Image source: Shutterstock
  • openai
  • ai safety
  • gpt-5
  • chain-of-thought
  • machine learning
Market Opportunity
Cosplay Token Logo
Cosplay Token Price(COT)
$0,000917
$0,000917$0,000917
+0,65%
USD
Cosplay Token (COT) Live Price Chart
Disclaimer: The articles reposted on this site are sourced from public platforms and are provided for informational purposes only. They do not necessarily reflect the views of MEXC. All rights remain with the original authors. If you believe any content infringes on third-party rights, please contact [email protected] for removal. MEXC makes no guarantees regarding the accuracy, completeness, or timeliness of the content and is not responsible for any actions taken based on the information provided. The content does not constitute financial, legal, or other professional advice, nor should it be considered a recommendation or endorsement by MEXC.

You May Also Like

Is Doge Losing Steam As Traders Choose Pepeto For The Best Crypto Investment?

Is Doge Losing Steam As Traders Choose Pepeto For The Best Crypto Investment?

The post Is Doge Losing Steam As Traders Choose Pepeto For The Best Crypto Investment? appeared on BitcoinEthereumNews.com. Crypto News 17 September 2025 | 17:39 Is dogecoin really fading? As traders hunt the best crypto to buy now and weigh 2025 picks, Dogecoin (DOGE) still owns the meme coin spotlight, yet upside looks capped, today’s Dogecoin price prediction says as much. Attention is shifting to projects that blend culture with real on-chain tools. Buyers searching “best crypto to buy now” want shipped products, audits, and transparent tokenomics. That frames the true matchup: dogecoin vs. Pepeto. Enter Pepeto (PEPETO), an Ethereum-based memecoin with working rails: PepetoSwap, a zero-fee DEX, plus Pepeto Bridge for smooth cross-chain moves. By fusing story with tools people can use now, and speaking directly to crypto presale 2025 demand, Pepeto puts utility, clarity, and distribution in front. In a market where legacy meme coin leaders risk drifting on sentiment, Pepeto’s execution gives it a real seat in the “best crypto to buy now” debate. First, a quick look at why dogecoin may be losing altitude. Dogecoin Price Prediction: Is Doge Really Fading? Remember when dogecoin made crypto feel simple? In 2013, DOGE turned a meme into money and a loose forum into a movement. A decade on, the nonstop momentum has cooled; the backdrop is different, and the market is far more selective. With DOGE circling ~$0.268, the tape reads bearish-to-neutral for the next few weeks: hold the $0.26 shelf on daily closes and expect choppy range-trading toward $0.29–$0.30 where rallies keep stalling; lose $0.26 decisively and momentum often bleeds into $0.245 with risk of a deeper probe toward $0.22–$0.21; reclaim $0.30 on a clean daily close and the downside bias is likely neutralized, opening room for a squeeze into the low-$0.30s. Source: CoinMarketcap / TradingView Beyond the dogecoin price prediction, DOGE still centers on payments and lacks native smart contracts; ZK-proof verification is proposed,…
Share
BitcoinEthereumNews2025/09/18 00:14
7 Best Crypto to Invest: One Presale is Breaking Records

7 Best Crypto to Invest: One Presale is Breaking Records

The post 7 Best Crypto to Invest: One Presale is Breaking Records appeared on BitcoinEthereumNews.com. What if the next great financial story isn’t written by Wall Street but by internet memes, culture, and digital tribes? Over the past few years, meme coins have transformed from playful jokes into market juggernauts, spawning billion-dollar valuations seemingly overnight. Dogecoin, Shiba Inu, and Pepe all proved that when community conviction collides with scarcity, even the most satirical token can rewrite portfolios. The hunt is on again in 2025: which contender will rise as the best crypto to invest in this cycle? That’s where BullZilla enters, roaring into the scene with mechanics that dwarf ordinary meme launches. Built on Ethereum, BullZilla ($BZIL) fuses mythic lore with technical brilliance: a progressive price engine, a 24-stage mutation presale, live Roar Burns, staking through the HODL Furnace, and the Roarblood Vault referral system. The BullZilla Presale is live now, and the rules are simple: the price rises every 48 hours or instantly when $100K is raised. This scarcity mechanism turns every stage into a race, rewarding the earliest believers. For anyone asking what is the best crypto to invest, the answer is already roaring. BullZilla has taken its place at the center of Trending Meme Coins 2025. Join early for maximum perks. 1. BullZilla ($BZIL): The Beast Mutates Toward 100x Gains The Bull Zilla Presale is quickly emerging as the top meme coin presale to buy now, drawing massive attention from both retail investors and large holders. Currently in its 3rd Stage fittingly named “404: Whale Signal Detected” the token is priced at $0.00007241. Over $530,000 has been raised, more than 27 billion tokens have been sold, and the presale has attracted over 1,700 holders. The planned listing price of $0.00527 translates into a potential ROI of 7,179.94% for those entering now. Early participants from Stage 3C are already sitting on gains of…
Share
BitcoinEthereumNews2025/09/22 07:20
House Democrat smacks down Trump's rambling ICE threat: 'This man can't win'

House Democrat smacks down Trump's rambling ICE threat: 'This man can't win'

A House Democrat smacked down President Donald Trump's rambling threat to deploy Immigration and Customs Enforcement agents to airports nationwide.Trump wrote on
Share
Rawstory2026/03/22 07:23