AIToday
Large Language ModelsAI Coding AssistantsVercel AI BlogPublished: Aug 22, 2026, 04:01 JST3 min read

Ora benchmarks all major AI agents on Vercel, finds eve outperforms rivals

Ora benchmarks all major AI agents on Vercel, finds eve outperforms rivals

Key takeaway

  • Ora tested all major AI agent frameworks on live websites and found Vercel's eve outperformed Claude Code.

  • Eve used 7% fewer steps and achieved 2× native success on customer sites.

  • The startup estimates 99% of the web cannot yet handle agents that sign up and pay, making benchmarking critical for adoption.

3 Key Points

  1. What happened

    Ora, a platform that tests how well AI agents can navigate and transact on live websites, ran side-by-side benchmarks of every major agent framework—Claude Code, ChatGPT, Gemini, Hermes, OpenClaw, and Vercel's eve—on customer sites. Eve outperformed Claude Code with 7% fewer steps to reach goals, 2× native success (twice as many tasks finished on the customer's own site instead of falling back to web search), and 9% more valid endpoints.

  2. Why it matters

    Ora estimates 99% of the web cannot yet handle agents that sign up, integrate, and pay. By running agents on live sites and recording where and why they fail, Ora shows companies what to fix—a critical step because agents often stumble in real-world workflows. The benchmark is significant because no two agent frameworks use the same infrastructure; Ora built separate runtimes for each and traces every step, letting customers see exactly which step stalled and what the agent tried.

  3. What to watch

    Ora is now building its own product on eve instead of treating it as just another test subject. The team runs 16 engineers shipping hundreds of commits a day, with coding agents handling day-to-day infrastructure work on Vercel. Ora is splitting its platform into microservices, all on Vercel, so new services and internal eve agents will deploy with no extra configuration needed.

Ask the AI about this article →

Context & Analysis

Ora's founding is rooted in a gap Assaf Elovic discovered at his previous company, Tavily, which Nebius acquired earlier this year. Tavily built a web search engine for AI agents—solving half the problem. But an agent that finds a product still must actually use it. Elovic and co-founder Liad Yosef started Ora to measure how ready the web is for agents and to fix what isn't.

The core challenge Ora solved is infrastructure fragmentation. Every major agent framework—Claude Code, ChatGPT, Gemini, and others—expects its own environment and exposes steps differently. Running them side by side required separate runtimes for each harness, with full tracing of every step. This depth is what Ido Finder calls "one of the most valuable things ora brings to its customers": without the trace, a benchmark score is meaningless. When an agent stalls, the customer sees exactly which step failed and what the agent tried.

Ora's decision to run entirely on Vercel—front end, back end, and agent runtime sharing the same deployment path, logs, and authentication—became the foundation for both testing and internal operations. That consolidation feeds back into Ora's own engineering: 16 people shipping hundreds of commits a day, with coding agents handling day-to-day infrastructure work on eve. Finder saves a few hours a week from this setup; Elovic credits similar savings to how well coding agents build with Vercel's libraries. As Ora scales, it is splitting its platform into microservices, all on Vercel, so internal eve agents will deploy as one more service with no extra configuration.

FAQ

Which agents did Ora test, and what was the comparison?
Ora benchmarked Claude Code, ChatGPT, Gemini, Hermes, OpenClaw, and eve. The initial detailed comparison was between eve and Claude Code, running the same models (Claude Fable 5 and Haiku 4.5) across hundreds of real journeys. Eve achieved 7% fewer steps, 2× native success, and 9% more valid endpoints than Claude Code.
What is Ora's core finding about web readiness for agents?
By Ora's estimate, 99% of the web is not agent-ready. The platform shows customers where agents fail on their sites and what changes are needed to make them work.
Why did Ora choose to build on eve after testing it?
Eve follows the Next.js paradigm with little configuration needed, and features like the sandbox override let Ora swap in its instrumented environment to trace every step without building anything new. Ido Finder, Ora's AI Lead, called eve "the easiest setup I've experienced."
Vercel AI BlogRead Original Article

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articlePanasonic Avionics cuts aircraft diagnostics time from hours to minutes using AWS AI agents

The AI news that matters, in one minute each morning.

Sign up free