
Artificial Analysis launched the Search Index, a benchmark that evaluates search API providers used by AI agents on three dimensions: quality, cost, and speed.
Testing seven providers with the same AI model, the benchmark found that search quality directly correlates with lower total costs—despite higher per-query fees, better search results reduce token consumption and improve agent accuracy.
Parallel, Exa, and Firecrawl led the quality rankings (75, 74, and 73 points respectively), and Parallel's advanced version achieved a 40 percent reduction in token use and lower overall task costs.
What happened
Artificial Analysis released the Search Index, a benchmark testing seven search API providers (Parallel, Exa, Firecrawl, You.com, Tavily, Keenable, and Brave) using the same AI model (GPT-5.6 Luna) in identical agent setups. The benchmark combines three equally weighted test suites: DeepSearchQA (900 research questions), BrowseComp (200 hard-to-find facts), and AA-Omniscience (600 questions across six domains).
Why it matters
Search quality directly affects both AI agent performance and cost. Without search, the model scored just 33 points; with search, scores ranged from 65 to 75. Better search quality reduces total token use—Parallel's advanced version cut token consumption by over 40 percent compared to its Basic version, lowering total task cost from $0.11 to $0.084 despite higher per-search fees. This means businesses can optimize both accuracy and spending by choosing the right search provider.
What to watch
Parallel, Firecrawl, and Parallel (turbo) achieved the best cost-performance mix, though raw speed per query does not always improve total task time—Parallel's turbo version responded fastest per query (0.51 seconds vs. 1.03 seconds for Basic) but lower quality (67 vs. 73) forced more agent passes, keeping overall time flat. The methodology is public and other providers can apply to join.
Ask the AI about this article →
Artificial Analysis built the Search Index to address a practical gap: how do you compare search APIs for AI agents when quality, cost, and speed all matter? By standardizing the model (GPT-5.6 Luna), the framework (Stirrup), and the test scenarios (900–600 questions per suite), the benchmark isolates the search provider as the only variable. This controlled approach yields a counterintuitive insight: the fastest search per query does not guarantee the fastest overall task completion. Parallel's turbo version delivers the shortest response time (0.51 seconds), but its lower quality forces the agent to run additional passes, negating the speed advantage. Meanwhile, Parallel's advanced version and Exa prioritize quality over raw speed, reducing the number of queries needed and cutting token consumption—which translates to lower total cost despite higher per-search fees. The benchmark thus reveals that for AI agents, quality is an economic lever: fewer, better-informed queries cost less overall than many cheap, low-quality ones.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
As AI technology matures, the bottleneck in the industry is moving beyond semiconductor constraints like GPUs…

OpenAI has launched an Apple Messages plug-in for ChatGPT that lets users connect their Messages inbox to the…

Amazon Bedrock now supports OpenAI GPT-5.6 models (Sol, Terra, and Luna variants) across more than 25 AWS Regi…

Slack introduced Slack Code, a new feature that lets teams collaborate with AI coding agents (Claude, Devin, G…

Cisco is transforming its digital customer experience (DCX) strategy by embedding AI throughout customer journ…

Mastercard CEO Michael Miebach introduced "Agent Pay" last April, a payment framework that allows AI agents to…
