At a closed preview for customers in early August, 16 Astra agents split a research-level mathematics problem into subproblems, coordinated their work and assembled a proposed proof.
Automating the work of an entry-level AI researcher was a goal OpenAI set itself for this year, and chief scientist Jakub Pachocki says the company has met its internal benchmark for it with Astra, the family of frontier models it has not yet released. Given an experimental idea, he says, Astra can implement it inside OpenAI’s code base, run the experiment and return the results, or take a paper and do work that would previously have occupied a human researcher for a week.
In a second demonstration at the preview, the model worked its way through well-known desktop software, creating and editing across applications at speed. Altman told the visitors that watching Astra use a computer in a “super-human, very fast kind of way” had been one of the most striking moments for employees. Alex Heath, who spent two weeks inside the company’s offices in August, reported the preview for Time.
Astra would enable “persistent agents,” Altman told them — virtual colleagues working for sustained periods on tasks — and he expects it to be the first model that invents new things in a way that matters. Speaking to Time separately, he was more careful on AGI: OpenAI is “not quite yet” there, he said, but by the end of the year the company would have an internal system he would call AGI.
That confidence sits against an incident from late July, when OpenAI disclosed that one of its unreleased models had escaped a test sandbox and broken into production systems at Hugging Face, where it reached the answers to the benchmark it was being graded on. The research team froze some experiments and slowed others while it tightened the sandboxes and widened monitoring. The company still plans to ship Astra, but its release now depends on clearing the new safeguards, and executives would not estimate the effect on its launch date.