
What happened
Anthropic published a blog post on September 8, 2026 (US time), outlining six common prompting anti-patterns that waste Claude tokens, and tested its new /claude-api prompt-audit command, which reduced costs by an average of 14.6% and improved accuracy by an average of 5.3%.
Why it matters
The company says the cost-versus-performance tradeoff is not inevitable; targeting just three areas—advanced tool use, prompt caching, and effort adjustment—can deliver both savings and maintained or improved performance, and it has packaged this guidance into a skill available in Claude Code.
What to watch
The recommendations hinge on prompt caching working as intended, so the specific configuration choices Anthropic describes—such as keeping the tools and system prompt unchanged, avoiding mid-conversation changes to the tools or system prompt, and placing static content at the beginning of the prompt—are likely to determine whether readers see similar gains. The article also notes that prompt cache expires, so cache efficiency may require continuous management.
WHO IT HITSEngineering and product teams building applications on Claude Platform—especially those already using Claude Code and its API—can apply the anti-pattern audit and the three focus areas to reduce token costs without sacrificing accuracy. Teams that maintain prompt caches will need to watch for configuration changes that invalidate the cache, which the body says can raise costs.
Summaries like this, in your inbox every morning.
Anthropic's blog post and the accompanying commands—/claude-api prompt-audit, /claude-api hillclimb, and /claude-api cost-optimize—reflect a push to make Claude more efficient at a time when model providers face pressure on both price and quality. The company's tests show that applying its prompt-audit command during a shift from Opus 4.8 to Opus 5 delivered both lower costs and higher accuracy, and that hillclimb explored model, effort level, and prompt updates to reach 98.9% accuracy on a training set at about 2.6 cents per thousand tokens. The guidance also acknowledges that prompt caching, while valuable, depends on exact byte-level prefix matches and a limited cache lifetime. For teams building on Claude Platform, the practical question is whether Anthropic's own test results transfer to their specific applications, since the body does not claim they will. The commands and skills are already available, so the main uncertainty is whether users can match the configuration conditions Anthropic describes for cache preservation and cost reduction.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
OpenAI launched Astra for Law, wrapping GPT-6 Astra in a legal search index covering US case law, statutes, re…
Anthropic detailed three metrics — AI-led R&D, oversight of autonomous AI agents, and compute allocation — dis…
Google Labs opened its experimental AI agent CC to households of up to six people
Shiseido Japan's AI agent for ingredient discovery cut search time by 95% and increased proposed ingredient ca…

Microsoft is phasing in a consolidation of its Microsoft 365 Copilot app and experience into Microsoft Copilot…

Nvidia CEO Jensen Huang rejected the notion that the industry must choose between safety and speed, saying "yo…
