
Apple's new pipeline, Agent Seer, builds test scenarios automatically from tool specifications. It requires no examples or live access.
Tests across seven specs showed strong quality.
Parameter schema complexity matters most.
What happened
Apple researchers introduced Agent Seer, a pipeline that automatically creates realistic test scenarios for AI agents that use external tools, starting from a single Model Context Protocol (MCP) specification. It requires no examples, no live tool access, and no domain-specific tuning.
Why it matters
Hand-built test scenarios don't scale and become outdated as APIs change, so this could make AI tool evaluation faster and more current. Tests on seven MCP specifications showed strong quality across all domains, including complete tool coverage on small and medium specs.
What to watch
Parameter schema complexity was the strongest factor in quality variation, with tool-suite size playing a smaller role. Argument value accuracy was the main failure mode, a detail not captured by coarse name-match metrics. The paper was accepted at the Fifth Workshop on Natural Language Generation, Evaluation, and Metrics at ACL 2026.
Ask the AI about this article →
The paper addresses a practical problem in evaluating AI agents that use external tools: creating test scenarios by hand requires deep expertise and doesn't scale, especially as APIs evolve. Agent Seer leverages the latent information in tool specifications, which include function names, descriptions, and parameter schemas, to synthesize realistic scenarios without manual curation. The evaluation on seven MCP specifications showed strong quality, with complete tool coverage on smaller suites, but also revealed that parameter schema complexity is the main driver of quality variation, not the number of tools. The findings suggest that improving argument value accuracy—often missed by simple name-match metrics—could be a key area for refinement. The paper's acceptance at a workshop at ACL 2026 signals growing interest in automating evaluation of tool-using AI systems, which may reduce the need for bespoke benchmarks.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Nvidia beat all expectations for revenue and its stock rose almost 9% Thursday
Anthropic held discussions to acquire AI chip startup MatX for roughly US$7 billion, according to Reuters

Cerebras Systems is expanding its data centers outside the US, with facilities in France, Finland, Norway, and…

Meta announced it is using closed-loop liquid cooling across the majority of its newest AI-optimized data cent…

Google DeepMind is launching the first double-blind evaluation of a proprietary frontier AI model

A federal judge ruled the Pentagon acted illegally by punishing AI company Anthropic for its criticism of the…
