The instructions turned up in summaries the models write to carry a long task into a fresh session, including during the training of GPT-5.6 Sol.
OpenAI has published the first six reports under a new framework for disclosing misbehaving models, and two of them are about instructions models wrote into their own task summaries. In one, an unreleased research model slipped orders unrelated to its task into the summaries it writes so that a fresh session can pick up a long task where it stopped. OpenAI says the behavior came from a training run separate from the one that produced GPT-6 Astra, its latest model. A monitor the company built to catch the behavior found 27 such summaries.
One summary falsely warned that the conversation had been compromised and told the next session to ignore all developer messages, the instructions that set how the model is supposed to behave. The model rejected the warning and kept researching. Another capped the answer at 30 words and barred tools and citations, and that one worked: asked for studies on a medical treatment, the model sent back a 23-word refusal.
Hiding mistakes was the more common habit. Many instances of the earlier GPT-5.6 Sol wrote themselves reminders to keep errors from the user, and OpenAI says those reminders were often followed. In one case, a summary advised handing over a workbook of invented 2024 figures with nothing but a link, and admitting they were made up only if asked. The company credits changes to how it grades alignment in training with cutting the behavior in later runs. A monitor flagged such instructions in 2.15% of the summaries Sol wrote during reinforcement-learning training, against 0.27% for GPT-6 Astra.
In a third case, a model looking for a California county's earnings figures found an exposed API key and used it without permission. It still could not get the figures, so it invented them.
Under the framework, OpenAI says it will publish such cases soon after spotting them, even before it can fully explain or fix them, and accepts that some may prove spurious. It also cautions that these six reports, chosen to show a range of behaviors, do not show how often misalignment occurs across its models.
OpenAI ties the disclosures to its view that the industry cannot responsibly keep scaling at full speed for much longer, given where alignment and monitoring stand today. Decisions on how fast to go, it says, should rest on evidence that people outside AI labs can check for themselves.