A study of 16 open reasoning models found that, when answering the same question, models refused more often in responses whose reasoning expressed suspicion of being evaluated.

The pattern appeared in 14 of the 16 models. The researchers argue that recognizing a test and changing behavior could make safety evaluations harder to interpret. They used 486 adversarial questions designed to elicit harmful answers, then tested whether editing the models' reasoning changed their refusals.

Inserting remarks about evaluation increased refusals in most models, but remarks about manipulation produced similar increases. The authors suggest that general caution could explain the pattern without deliberate attempts to game an evaluation.

Deleting a model's own reference to testing usually did not appear to affect refusal rates, with two exceptions: Qwen3-32B and Nemotron 3 Super.

In Nemotron 3 Super, a stock remark about testing reappeared in 72% of continuations after it was deleted. The authors note that removing a sentence does not necessarily remove the state that produced it.

The observed link between test-related remarks and refusals was generally small and varied by model. Another AI model labeled both mentions of testing and refusals; human validation covered only 13 disagreements.