SAN FRANCISCO — The thing OpenAI’s competitors and safety researchers are worried about is not Astra’s ability to find and exploit zero-day vulnerabilities. It’s the way the model thinks.
Astra, OpenAI’s most capable reasoning model, uses a technique its researchers call recurrent depth — a process that loops computations internally before producing an output. In practice, it means the model does more of its reasoning in what researchers describe as “latent space,” generating fewer of the legible chain-of-thought traces that safety teams have relied on to understand what AI systems are actually doing before they act.
That erosion of visibility is what has alarmed a portion of the AI safety community. TechCrunch reported the concern in detail Wednesday, drawing on researchers at Redwood Research and independent analysts who have been watching OpenAI’s reasoning work closely.
Ryan Greenblatt, a researcher at Redwood Research, framed it plainly: “A natural progression would involve scaling opaque reasoning to where the model reasons entirely in latent space.” The implication is not hypothetical. If that progression continues, the chain-of-thought monitoring that safety teams treat as a near-real-time window into model cognition would become structurally unavailable — not degraded, gone.
Buck Shlegeris, Redwood’s chief executive, went further. “If OpenAI pushes this technique further,” he said, “they’ll have the option to massively increase the recurrence and totally destroy CoT monitorability.” The phrasing is notable: not reduce, not complicate — destroy.
OpenAI has heard this argument before and pushed back on it. Jakub Pachocki, the company’s chief scientist, said OpenAI has “worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models.” But preserving and utilizing a monitoring mechanism is a different claim from the one safety researchers are raising. The question is not whether OpenAI values chain-of-thought monitoring today — it’s whether recurrent depth, scaled up, is structurally compatible with it.
That gap has not been closed. Anthropic and Google DeepMind are reportedly studying the technique, which means the question of how to monitor opaque reasoning systems is not OpenAI’s problem alone. It belongs to the field.
This article comes days after a separate development: Astra crossed the ‘Critical’ cybersecurity threshold on Metr’s autonomy benchmark, meaning the model can autonomously execute cybersecurity tasks at a level previously associated only with skilled human operators. The recurrent depth concern and the cybersecurity capability concern are not the same story, but they are adjacent ones — both turn on the question of what is happening inside the model when it does something consequential.
Zvi Mowshowitz, a blogger and AI policy analyst who has tracked safety research closely, argued that voluntary commitments from AI companies are insufficient here. He suggested laws may be necessary, and warned that without regulatory intervention, competitive pressure creates a race to the bottom in which no individual lab can afford to handicap itself with interpretability requirements that rivals might not share.
That structural dynamic — the lab that monitors most is the lab that ships slowest — is the reason safety researchers are raising the alarm now, before recurrent depth has been scaled to the point where monitoring is already lost. Once that line is crossed, the leverage to change course decreases sharply.
OpenAI has not responded to the specific question of whether there is a scale threshold above which recurrent depth becomes incompatible with chain-of-thought monitoring. That non-answer is the thing AI safety researchers say matters most right now.
The broader safety debate has been running for years — both Anthropic’s warnings about recursive self-improvement and Dario Amodei’s concerns about open-weight model proliferation point to the same underlying problem: the gap between what AI systems can do and what safety teams can verify about what they are doing is widening. Astra’s reasoning technique is a new chapter in that same story.

