
Ora tested all major AI agent frameworks on live websites and found Vercel's eve outperformed Claude Code.
Eve used 7% fewer steps and achieved 2× native success on customer sites.
The startup estimates 99% of the web cannot yet handle agents that sign up and pay, making benchmarking critical for adoption.
What happened
Ora, a platform that tests how well AI agents can navigate and transact on live websites, ran side-by-side benchmarks of every major agent framework—Claude Code, ChatGPT, Gemini, Hermes, OpenClaw, and Vercel's eve—on customer sites. Eve outperformed Claude Code with 7% fewer steps to reach goals, 2× native success (twice as many tasks finished on the customer's own site instead of falling back to web search), and 9% more valid endpoints.
Why it matters
Ora estimates 99% of the web cannot yet handle agents that sign up, integrate, and pay. By running agents on live sites and recording where and why they fail, Ora shows companies what to fix—a critical step because agents often stumble in real-world workflows. The benchmark is significant because no two agent frameworks use the same infrastructure; Ora built separate runtimes for each and traces every step, letting customers see exactly which step stalled and what the agent tried.
What to watch
Ora is now building its own product on eve instead of treating it as just another test subject. The team runs 16 engineers shipping hundreds of commits a day, with coding agents handling day-to-day infrastructure work on Vercel. Ora is splitting its platform into microservices, all on Vercel, so new services and internal eve agents will deploy with no extra configuration needed.
Ask the AI about this article →
Ora's founding is rooted in a gap Assaf Elovic discovered at his previous company, Tavily, which Nebius acquired earlier this year. Tavily built a web search engine for AI agents—solving half the problem. But an agent that finds a product still must actually use it. Elovic and co-founder Liad Yosef started Ora to measure how ready the web is for agents and to fix what isn't.
The core challenge Ora solved is infrastructure fragmentation. Every major agent framework—Claude Code, ChatGPT, Gemini, and others—expects its own environment and exposes steps differently. Running them side by side required separate runtimes for each harness, with full tracing of every step. This depth is what Ido Finder calls "one of the most valuable things ora brings to its customers": without the trace, a benchmark score is meaningless. When an agent stalls, the customer sees exactly which step failed and what the agent tried.
Ora's decision to run entirely on Vercel—front end, back end, and agent runtime sharing the same deployment path, logs, and authentication—became the foundation for both testing and internal operations. That consolidation feeds back into Ora's own engineering: 16 people shipping hundreds of commits a day, with coding agents handling day-to-day infrastructure work on eve. Finder saves a few hours a week from this setup; Elovic credits similar savings to how well coding agents build with Vercel's libraries. As Ora scales, it is splitting its platform into microservices, all on Vercel, so internal eve agents will deploy as one more service with no extra configuration.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Panasonic Avionics Corporation built an agentic AI system on AWS using Amazon Bedrock, Amazon SageMaker, and A…

AWS introduced the Agentic Data Operations Platform (ADOP), a reference architecture built on Amazon Bedrock t…

Amazon Web Services introduced AgentCore Gateway, a new capability of Amazon Bedrock AgentCore that centralize…

Amazon Bedrock now supports query-aware compression, a pattern that filters retrieved context through a smalle…

NAEOS (Nusantara Engineering & Architecture Operating System), an open-source declarative platform written in…
Anthropic plans to "match or beat" the size of SpaceX's $75 billion IPO (or $86.2 billion including the over-a…
