A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal
The authors introduce the Probe of Internal Recognition (PIR), a forensic-inspired method that reads a model's internal states to identify which candidate answer it recognizes as correct. Tested across eight models from five families (Gemma, Qwen, Llama, Mistral, and Phi), PIR recovers recognized answers at 0.70 to 0.87 balanced accuracy, outperforming the 0.28 to 0.40 unknown-item baseline. The method remains robust across various forms of concealment including prompted deception and trained sandbagging, thus supporting model audits and unlearning verification.
PIR recovers the recognized answer at 0.70 to 0.87 balanced accuracy, well above the 0.28 to 0.40 unknown-item baseline and the 0.25 chance rate.
It stays readable across every form of concealment we test, from prompted deception and trained sandbagging to external password-locked and circuit-broken checkpoints, with recognition between 0.85 and 0.93.