Imitate a model's own reasoning notes and it often obeys — a flaw the researchers say training cannot fix
An LLM reads a passage's role off its writing style, not off the role tags that safety training relies on. Swapping those tags around a passage barely changed how models behaved — and forged notes pulled cocaine-manufacturing and aircraft-sabotage...