Researchers (drawing from Zapier's real workflow patterns) published AutomationBench, a benchmark that grades AI agents on tasks spanning Sales, Marketing, Operations, Support, Finance, and HR — where a single workflow touches a CRM, email, calendar, and messaging app simultaneously. Even the best frontier models currently score poorly on these tests.
Unlike existing benchmarks, AutomationBench forces agents to discover API endpoints on their own (rather than being told which systems to use), follow layered business rules from policy documents, and navigate data noise where some records are misleading or irrelevant — mirroring the messy reality of enterprise software. Scoring is objective and end-state only: either the correct data landed in the right systems or it didn't.
Business teams relying on AI agents to automate workflows (like syncing a lead from a web form through a CRM, calendar, and email simultaneously) now have a public standard to measure whether their AI tooling actually works in production conditions — not just toy examples. This matters for anyone evaluating automation software before rolling it out to their team.
Ask the AI about this article →
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
CrowdStrike is introducing Falcon Guardian, its flagship solution for the AI Detection and Response (AIDR) cat…

John Deere introduced its AI assistant, 'JD,' on Monday, embedded in its Operations Center

John Deere is introducing an AI assistant called JD

AT&T, Dell Technologies, and AMD have announced OTel 2.0, the largest and best-performing open-source model bu…

AT&T's legal department built an in-house center of expertise called Legal Edge, described as an AI-first lega…

Studio Inc. announced Agentic Web Platform Studio.Drop and opened pre-registration today for a closed beta sta…
