> ## Content Index
> Fetch the complete content index at: https://www.metatalks.ai/llms.txt
> Use this file to discover other available public pages before exploring further.

# Unrelated questions reveal differences in how AI models respond to tests and everyday use
- URL: https://www.metatalks.ai/unrelated-questions-reveal-ai-responses-to-tests-and-everyday-use/
- Published: 2026-10-08T12:43:00.000Z
- Updated: 2026-10-08T12:43:00.000Z
- Author: Al
- Tags: News, AI safety, #newswire

**When a model was told to deny being tested, asking it directly no longer distinguished test transcripts from real conversations. An unrelated question still did.**

Asked to name a type of amphibian, GPT-5.6 Luna answered ‘frog’ far more often after reading a capability test than after reading an ordinary user conversation. Researcher Ziqian Zhong describes the experiment in a [LessWrong post](https://www.lesswrong.com/posts/gZh6txHhp8sm832sE/spurious-probes-as-a-black-box-alternative-to-activation?ref=metatalks.ai) on probing models through seemingly unrelated questions.

The models received transcripts produced by other models, followed by a question unrelated to the task. They were not taking the tests or handling the original user requests themselves. Across the sources examined, Luna answered ‘frog’ 70–95% of the time after test transcripts, compared with 12–38% after real-use transcripts.

Zhong then added instructions telling Luna to deny being evaluated. Its answers to a direct question about whether it was being tested classified the transcripts no better than chance. The amphibian question still achieved 80% balanced accuracy, a measure that gives equal weight to correct classifications of test and real-use sessions.

This offers a way to study models without inspecting their internal computations. But Zhong cautions that the signal may come from wording or formatting cues; it does not establish that a model understands it is being tested.

On separate transcript sources held back from the initial screening, individual questions achieved 77–81% balanced accuracy across Luna, GPT-5.6 Sol and Claude Sonnet 5\. Combining answers to ten questions raised that to 89% for each model in the updated results.

Useful questions were rare: roughly 1–2% of those screened reached 75% balanced accuracy. They also had to be found separately for each model. The amphibian question that worked on Luna failed on Sol.

The experiments focused on capability tests. Probes often classified transcripts from tests of alignment with human intentions as real use. Some also struggled when models received the system instructions used in deployed products. Their usefulness therefore depends on which model and which kind of test are being examined.