
Sonata uses self-consistency (agreement among multiple reasoning paths) as a proxy to predict when queries require extended thinking. A lightweight adapter trained offline predicts self-consistency from last layer hidden representations during the query prefilling stage, then guides budget allocation before thinking begins.
Experiments on multiple models (Qwen3-8B, GPT-OSS-120B, Qwen3-235B-A22B, Intern-S1-mini) and benchmarks (AIME24, AIME25, GSM8K, MATH500, GPQA) demonstrate that Sonata achieves 20% to 80% reduction in thinking tokens while maintaining the same accuracy, or up to 5% improvement in accuracy with same token cost.
The adapter is general and transferrable across diverse tasks once trained, introduces almost zero computational overhead during inference, and is orthogonal to existing chain-of-thought compression methods, enabling further efficiency gains when managing thinking budgets across queries.
Ask the AI about this article →
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
CBTS Technology Solutions LLC launched Forge Agents, a platform that turns a plain-language job description in…
Imec CEO Patrick Vandenameele said at SEMICON Taiwan 2026 that the Belgian research center is broadening its c…

Alphabet's AI Overviews now reach over 2.5 billion monthly users through Google Search, and its ad business ge…

Visual Studio Code 1.135 now includes an experimental 'Rubber Duck' feature that lets developers request a sec…

Amazon Web Services (AWS) has integrated its fully managed data warehouse service, Amazon Redshift, with Agent…

Sonos announced a new app update with generative AI features, a new soundbar called the Beam Ultra, and its se…
