AIToday
Large Language ModelsAI Coding AssistantsITmedia AI+Published: Oct 1, 2026, 13:00 JST

AWS ties Kiro prompt tuning to LLM judge, A/B tests

AWS ties Kiro prompt tuning to LLM judge, A/B tests

3 Key Points

  1. What happened

    AWS said on August 25, 2026 that it evaluates Kiro's system prompt by scoring real internal conversations with an LLM judge, and in the first comparison cut Kiro CLI unspecified errors 5% and Kiro IDE style mismatches 54%.

  2. Why it matters

    System prompts weaken as models, tools, codebases, tasks and users change, so AWS's practices may help teams keep agents working as intended, though AWS notes benchmarking cannot cover every environment.

  3. What to watch

    AWS says prompt changes that merely fit one model version may not carry over, so it now retunes prompt and model as a pair; watch whether the same approach holds as models update.

WHO IT HITSEngineering and platform teams running internal AI agents should expect to own prompt evaluation and periodic retuning, since AWS describes prompt decay as an ongoing operational task rather than a one-time setup.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

AWS's disclosure centers on a problem it frames as inherent to agents like Kiro, which combine a model, tools, a codebase, tasks and users. Because those ingredients change, the company argues no prebuilt test suite can validate every combination. Its answer is a four-stage loop: diagnosing recurring failure patterns from internal traffic, designing changes around recall, control, tools or coding, testing them with control and cohort groups, and reviewing LLM-judge scores per cohort.

The LLM judge scores conversations on 15 criteria, including task completion, whether Kiro split off a conclusion before verification, and whether it ran code or checked compilation. AWS separates results into textual dissatisfaction and product-issue signals, and it says judgments are revised only when there is objective evidence, to limit hallucinations.

In past validation, AWS screened 27 changes at once and found that applying them together instantly degraded performance, which is why it favors hash-based splitting and A/B comparison. It also applied the mechanism to reasoning-effort settings, lowering the setting to see whether quality improvements from more effort stopped. The initial prompt-and-structure comparison produced the reported reductions in unspecified errors and style mismatches, but prompt-only gains from a model update were limited to 4%. AWS's own reading is that the new model simply made fewer mistakes, leaving less room for prompt improvement, and it later found cases where the model followed instructions too literally and produced unintended side effects. The practical implication is that prompt quality may look like a one-time setup but behaves like an operational cycle, and the test of AWS's approach is whether retuning the model and prompt together keeps improvements from backsliding.

FAQ
How does AWS evaluate Kiro's system prompt?
AWS scores internal conversation sessions with an LLM judge on 15 criteria, then assigns conversations to a control group and a cohort using the changed prompt for A/B comparison.
What results did AWS report from its initial evaluation?
The first comparison reduced Kiro CLI unspecified-error signals 5% and task-completion failures 10.6%, while Kiro IDE reduced task-completion failures 20% and style mismatches 54%.
What does AWS advise when a new model is released?
AWS says a new model's higher raw capability may leave little room for prompt improvement, so it treats the new model and prompt as a pair and retunes them together.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleBull Theory: midterm loss could pop AI bubble