
What happened
With GPT‑6, OpenAI launched prompt caching giving higher cache hit rates by default and discounts for eligible shared prefixes reused within a 30-minute window.
Why it matters
Higher default hit rates and discounts for reused prefixes can lower response times and costs for applications that repeatedly carry forward the same instructions, tools, and context.
What to watch
The gains hinge on how well developers keep tool definitions and instructions stable, since changes can cause cache misses; new diagnostics and dashboards are meant to help them spot and fix those misses.
WHO IT HITSDevelopers and teams building AI agents or API-driven applications that make repeated calls with shared context stand to see lower latency and costs if they keep their prompts stable. Teams whose tools and instructions change frequently may see less benefit.
Summaries like this, in your inbox every morning.
GPT‑6 is designed to support persistent agents that work for hours on complex tasks, from refactoring codebases to producing well-researched documents. Those agents make a series of API requests that build on each other, often carrying forward the same instructions, tool definitions, and context. OpenAI caches that shared context to reuse computation across requests, reducing response times and giving developers discounts on cached input tokens.
The new prompt caching system builds on that engine by delivering higher hit rates by default and adding tools to help developers see and fix misses. A dashboard shows how much of an application's input is served from cache, and a diagnostics tool compares a request with a recent response to identify what changed. Developers can also choose explicit cache breakpoints, adjust reasoning effort without breaking cache, keep tool definitions stable, and prewarm the cache.
The value of these controls hinges on how well developers manage stability in their prompts and tools. Teams that keep their setups consistent may capture more of the default performance gains, while those with frequent changes may need the new diagnostics to understand where reuse is lost.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, LINE, or Slack.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
A New York Times columnist gave control of everyday tasks to Meta's AI agent and came away impressed

Palo Alto Networks announced Unit 42 Continuous Frontier AI Defense, an annual-subscription service that pairs…

ServiceNow raised its 2026 subscription revenue midpoint, citing strong demand for its workflow platform, AI-f…

Axios exclusively reported that a major cybersecurity vendor has introduced a new service aimed at fighting AI…

Palo Alto Networks unveiled an AI-powered cybersecurity service that uses Claude and GPT models

DEEPX, a South Korean AI semiconductor company, is reportedly conducting a proof of concept with a US smart gl…
