AIToday
Large Language Modelsr/MachineLearningPublished: Aug 22, 2026, 13:00 JST2 min read

Telling LLMs to be concise cuts costs 1.5× without losing accuracy

Telling LLMs to be concise cuts costs 1.5× without losing accuracy

Key takeaway

  • Researchers tested nine LLMs and found that instructing models to write concisely cuts costs by roughly 1.5× while keeping accuracy steady.

  • Shortening input prompts, by contrast, did not reduce costs effectively.

  • The finding is backed by testing across five datasets, eleven languages, and longer-form summarization.

3 Key Points

  1. What happened

    Researchers tested nine LLMs—GPT-4o, GPT-4, Claude Haiku, Claude Sonnet, Qwen2.5-VL-7B, Qwen3.5-9B, DeepSeek-R1-Distill, Gemma-4-E4B, and Kimi-K2.6—across five short-answer datasets, eleven languages, and longer-form summarization tasks. They measured the cost and accuracy impact of two approaches: shortening the input prompt versus instructing the model to produce shorter output.

  2. Why it matters

    Shortening output instructions saved approximately 1.5× in cost on average while maintaining accuracy, whereas compressing the input prompt did not yield the same benefit. This finding is relevant for organizations running large-scale LLM inference, where token consumption directly drives operational expense. Claude Code recently shipped a "concise output style" feature based on this work.

  3. What to watch

    The study evaluated performance across multiple reduction levels and confirmed that shortened text still matched the model's unconstrained response quality, suggesting the cost savings do not come at the expense of answer fidelity.

Ask the AI about this article →

Context & Analysis

The paper addresses a practical constraint in LLM operations: while users can control what they feed into a model and how they ask it to respond, the model's natural verbosity drives up inference costs through token consumption. Claude Code's recent "concise output style" feature reflects growing awareness in the industry that output length is a lever companies can pull. The researchers' systematic approach—testing across five reduction levels, multiple datasets, eleven languages, and a longer-form task—suggests they designed the study to capture real-world variation rather than a single use case. The critical finding is the asymmetry: telling a model to be concise works; telling it a concise prompt does not. This implies that the model's instruction-following capability in the output channel is more reliable than the intuition that fewer input tokens automatically reduce downstream verbosity. The accuracy preservation across all nine models tested indicates the finding is not model-specific.

FAQ

Which models were tested?
The study evaluated GPT-4o, GPT-4, Claude Haiku, Claude Sonnet, Qwen2.5-VL-7B, Qwen3.5-9B, DeepSeek-R1-Distill, Gemma-4-E4B, and Kimi-K2.6.
How much can you save by asking for concise output?
Shortening the output saved approximately 1.5× in cost on average while keeping accuracy about the same.
Does shortening the input prompt save money too?
No. The study found that compressing the input prompt does not yield the same cost-saving benefit as instructing the model to produce shorter output.
r/MachineLearningRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Visko raises $10M, launches live AI video model OrbisSiliconANGLE AI · 1h ago
  • Runway unveils Solaris, an AI that generates app interfaces in real timeTHE DECODER · 1h ago
  • Google AI Search flags Facebook users as dangerTHE DECODER · 1h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleNscale seeks $3B IPO, eyes 11 gigawatts of AI data center capacity