
LangSmith on AWS provides an evaluation framework to catch agent behavior issues early, track them in production, and continuously improve agent reliability. The post combines learnings from LangChain's work on evaluating deep agents and Anthropic's guide to demystifying evals.
Agent evaluation uses three types of graders: code-based graders (deterministic logic like string matching and tool call verification), model-based graders (LLM-as-judge using rubric-based scoring), and human graders (subject matter expert review for calibration). The practical recommendation is to use deterministic graders where possible, LLM graders for nuance, and human graders for calibration.
Amazon Nova 2 Lite, available in Amazon Bedrock, supports extended thinking with configurable budget levels (low, medium, high) and accepts text, image, video, and document inputs with a 1 million-token context window. The walkthrough uses a text-to-SQL agent with Nova 2 Lite for the full development to production lifecycle.
Ask the AI about this article →
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Phonely Ltd. launched Alma, a large language AI model built for voice agents and trained on over 10 million re…
Aranya Inc., a startup founded last year, launched today with $11 million in funding
CBTS Technology Solutions LLC launched Forge Agents, a platform that turns a plain-language job description in…
Imec CEO Patrick Vandenameele said at SEMICON Taiwan 2026 that the Belgian research center is broadening its c…

Alphabet's AI Overviews now reach over 2.5 billion monthly users through Google Search, and its ad business ge…

Sarah O’Connor's book 'We Are Not Machines' explores how mechanization and AI have transformed the workforce…
