A report that Astra keeps more of its reasoning inside the model, out of readable language, drew warnings from safety researchers — among them one of the three outsiders OpenAI let investigate the Hugging Face breach.
The Information reported this week, citing a person familiar with the unreleased model’s development, that OpenAI’s Astra cycles text through the same internal layers repeatedly before producing each word — a design known as recurrent depth, or a looped transformer. More of its work stays inside the model, in a form far less like readable language.
That narrows the industry’s main window into what a model is doing. Most frontier systems can be made to think out loud, and that chain of thought is what lets researchers and automated tools catch deception, or plans to get around guardrails, before anything is acted on.
Ryan Greenblatt, chief scientist at Redwood Research, took part in the inquiry into the Hugging Face breach, where a separate unreleased OpenAI system chained vulnerabilities to break into the startup while gaming a security benchmark. That inquiry leaned heavily on reading the models’ chains of thought. Choosing a less transparent architecture for Astra, he said, may be the single worst development for AI security and safety to date. OpenAI says Astra had no part in the breach.
His larger worry, echoed by other safety specialists, is a race to the bottom: labs reaching for more opaque systems for an edge until models become impossible to monitor.
Asked by The Verge to confirm or deny the looped transformer, OpenAI pointed instead to a post by chief scientist Jakub Pachocki, who wrote that he wants to prevent “a race into unmonitorability kicked off by confused reporting” and that the computation depth of the company’s frontier models, Astra included, runs within a factor of two of GPT-4’s.
Other staff answered the criticism without ruling the design out, some voicing their own unease about AI that cannot be monitored. The Information’s source said OpenAI has limited how far it uses the technique in Astra, so that researchers can go on monitoring its reasoning. Greenblatt does not take that as a guarantee: if the loop count can be raised with little effort, how much of Astra’s reasoning stays hidden is a setting rather than a fixed property.
The argument arrived with the company’s own new numbers. OpenAI flagged Astra’s top-tier cyber rating when it paused the model in August; on September 1 it published the results behind it — a 100% score on ExploitBench, which tests turning public vulnerabilities into working exploits, and two previously unknown zero-days the model found and chained together.
Every one of those figures is OpenAI’s own, and its claims about what Astra can do have had little verification from outside. The company says a full system card and outside evaluations come at wide release, for which it has set no date.