Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models
QUESTION — How can hidden chain-of-thought traces be extracted and characterized from closed-source frontier language models via standard APIs?
This research leverages a simple custom tool registered through a standard API feature to induce frontier models to externalize intermediate reasoning that is otherwise hidden in closed systems. By comparing against native CoT on open models and extending to systems like GPT-6 Astra, the authors demonstrate that the extracted reasoning matches native reasoning performance and substantially outperforms no-reasoning baselines across competition mathematics, science, and code generation. Structural analysis reveals that models like Astra exhibit token-efficient directed reasoning, resolving elementary steps internally while externalizing only crucial reasoning.
Xiaoyuluoit97 · 22 Sept 2026
read the original ↗