Safety researchers say OpenAI's Astra may be the hardest frontier model yet to monitor
A report that Astra keeps more of its reasoning inside the model, out of readable language, drew warnings from safety researchers — among them one of the three outsiders OpenAI let investigate the Hugging Face breach. The Information reported this week, citing a...