
What happened
AWS said on August 25, 2026 that it evaluates Kiro's system prompt by scoring real internal conversations with an LLM judge, and in the first comparison cut Kiro CLI unspecified errors 5% and Kiro IDE style mismatches 54%.
Why it matters
System prompts weaken as models, tools, codebases, tasks and users change, so AWS's practices may help teams keep agents working as intended, though AWS notes benchmarking cannot cover every environment.
What to watch
AWS says prompt changes that merely fit one model version may not carry over, so it now retunes prompt and model as a pair; watch whether the same approach holds as models update.
WHO IT HITSEngineering and platform teams running internal AI agents should expect to own prompt evaluation and periodic retuning, since AWS describes prompt decay as an ongoing operational task rather than a one-time setup.
Summaries like this, in your inbox every morning.
AWS's disclosure centers on a problem it frames as inherent to agents like Kiro, which combine a model, tools, a codebase, tasks and users. Because those ingredients change, the company argues no prebuilt test suite can validate every combination. Its answer is a four-stage loop: diagnosing recurring failure patterns from internal traffic, designing changes around recall, control, tools or coding, testing them with control and cohort groups, and reviewing LLM-judge scores per cohort.
The LLM judge scores conversations on 15 criteria, including task completion, whether Kiro split off a conclusion before verification, and whether it ran code or checked compilation. AWS separates results into textual dissatisfaction and product-issue signals, and it says judgments are revised only when there is objective evidence, to limit hallucinations.
In past validation, AWS screened 27 changes at once and found that applying them together instantly degraded performance, which is why it favors hash-based splitting and A/B comparison. It also applied the mechanism to reasoning-effort settings, lowering the setting to see whether quality improvements from more effort stopped. The initial prompt-and-structure comparison produced the reported reductions in unspecified errors and style mismatches, but prompt-only gains from a model update were limited to 4%. AWS's own reading is that the new model simply made fewer mistakes, leaving less room for prompt improvement, and it later found cases where the model followed instructions too literally and produced unintended side effects. The practical implication is that prompt quality may look like a one-time setup but behaves like an operational cycle, and the test of AWS's approach is whether retuning the model and prompt together keeps improvements from backsliding.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
HENNGE said on October 1 it set up HENNGE AI, a subsidiary with only two directors and no other employees, whe…

Writing for Robotics and Automation News, Unbox Robotics CEO Pramod Ghadge said reverse logistics needs AI dec…

GUGA said on October 1 it will significantly revise the syllabus for its 生成AIパスポート certification, applying it…

AMD announced AMD Ross, an agentic AI assistant for embedded design and development that runs across the AMD E…

DataSnipper CEO Vidya Peters said audit faces a staffing crisis, with two to two and a half times as many jobs…

Yann LeCun, the 2018 Turing Award winner, said he has "zero concerns" about rogue AI incidents like OpenAI age…
