CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — evaluation 9 upvotes

A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal

QUESTION — How can we determine whether a language model genuinely lacks knowledge or is simply concealing it?

The authors introduce the Probe of Internal Recognition (PIR), a forensic-inspired method that reads a model's internal states to identify which candidate answer it recognizes as correct. Tested across eight models from five families (Gemma, Qwen, Llama, Mistral, and Phi), PIR recovers recognized answers at 0.70 to 0.87 balanced accuracy, outperforming the 0.28 to 0.40 unknown-item baseline. The method remains robust across various forms of concealment including prompted deception and trained sandbagging, thus supporting model audits and unlearning verification.

PIR recovers the recognized answer at 0.70 to 0.87 balanced accuracy, well above the 0.28 to 0.40 unknown-item baseline and the 0.25 chance rate.

It stays readable across every form of concealment we test, from prompted deception and trained sandbagging to external password-locked and circuit-broken checkpoints, with recognition between 0.85 and 0.93.

hisku · 18 Sept 2026 read the original ↗
↑