
What happened
Mercor had 12 licensed CPAs work through simplified APEX Accounting Benchmark tasks. The best models now solve them almost flawlessly, versus scoring below the accountants' roughly 37 percent average eighteen months ago.
Why it matters
AI models are now faster, more accurate, and far cheaper than human accountants on structured bookkeeping tasks, per the study — a reversal from 18 months ago, when accountants beat the models.
What to watch
Mercor admits its tasks test exactly what AI does best — hunting down details and following instructions precisely — and left out client calls, colleague check-ins, and years of built-up context, so whether AI can close the books without oversight is still undecided.
WHO IT HITSStaff accountants and bookkeeping teams at firms that run structured, rules-based ledger work are the ones most directly exposed: the study's own numbers put AI ahead on speed, accuracy, and cost for those tasks, while client-facing and context-heavy parts of the job remain out of scope.
Summaries like this, in your inbox every morning.
The Mercor study is notable less for the topline that AI beats accountants than for how quickly the gap closed. Eighteen months ago, the best models still scored below the accountants' average of about 37 percent on simplified tasks drawn from the APEX Accounting Benchmark. Today, those same tasks are solved almost flawlessly.
But the full benchmark tells a different story from the simplified version. Across 160 tasks spanning 10 simulated companies, Claude Opus 5.5 leads with 61.8 percent of grading criteria met, followed by Fable 5.1 at 61.0 percent and GPT-6 Astra at 57.9 percent — and Mercor says no model fully solved almost 60 percent of the tasks. Mercor also concedes the study's tasks test exactly what AI does best: hunting down details and following instructions precisely. The parts of the job that involve talking with clients, checking in with colleagues, and drawing on years of accumulated context were left out entirely.
That combination — near-perfect performance on the narrow tasks, large gaps on the fuller set — is why Mercor says accountants cannot be replaced, even as it expects major productivity gains across the industry. The reading is that AI is becoming a strong assistant for structured bookkeeping rather than an autonomous book-closer, and the benchmark's own design suggests the boundary is as much about what the tasks leave out as about what the models can do.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
NetApp and Iterate.ai are packaging the AIPod Mini with Iterate's Generate platform and an embedded LLM, so en…
Testing Azure API Management's llm-token-limit policy at 800 tokens per hour, actual consumption hit 1,472 tok…

Qwen released Qwen3.8-Flash-Next on August 27, 2026, calling it a preview of the architecture planned for Qwen…

A Zenn article floated a hackathon where participants get the theme on the day, use no PC, internet, smartphon…

Anthropic's Message Batches API offers a 50% off rate, takes up to 10,000 requests per batch, and returns resu…

A developer testing Cloudflare's Clef and Clef-flash, released October 1 under Apache 2.0, found Jev's officia…
