
Artificial Analysis has released Optima, a benchmarking platform that lets users test AI models against their own data and workflows instead of relying on general-purpose benchmarks.
Early testers have used it to find models that cut costs significantly for finance and accounting tasks, and to match specific writing styles in legal work.
The platform charges based on actual token usage and evaluation costs, addressing a known limitation of public benchmarks that often fail to capture the specifics of individual use cases.
What happened
Artificial Analysis, known for independent LLM evaluations, has launched Optima, a platform that lets users build custom benchmarks to compare AI models using their own data and workflows. Users can upload datasets, describe use cases, or provide sample inputs, then test models on quality, cost per task, and time per task. The platform is available now.
Why it matters
Public benchmarks use predefined tasks that may not reflect a specific business need. Optima addresses this by letting organizations test models against their actual workflows—finance teams found models that cut costs by a factor of ten without major quality loss, and legal teams tested which model best matched their writing style. For agentic applications, cost per completed task often matters more than raw token price, since a cheaper model can end up more expensive if it fails more often.
What to watch
Optima charges only actual token costs with no markup—rubric-based evaluations cost $0.125 per criterion per model, and pairwise evaluations cost $0.375 per comparison. However, the platform's usefulness still depends on how precisely capabilities are defined and how representative test cases are; even cheap and fast outputs may need heavy rework if they add little business value.
Artificial Analysis, which operates well-known LLM evaluations including GDPval-AA and AA-Briefcase, has released a new platform called Optima designed to close a gap in AI model comparison. While public benchmarks test models on predefined tasks, they often don't reveal which model actually performs best for a particular organization's needs. Optima allows users to build their own benchmarks using their specific data and workflows, then run tests across leading current models and compare results on quality, cost per task, and time per task.
The platform offers three paths to building a custom benchmark. Users can upload existing evaluation datasets from their own files or from Hugging Face, provide AI agent traces from platforms like Arize, Braintrust, or Langfuse, or install a skill that gathers information from their coding environment and past sessions. For those without ready-made data, Optima accepts a description of the intended use case along with sample inputs and outputs, then generates suggested test inputs, evaluation criteria, and example tasks for users to review and refine. Two scoring approaches are available: rubric-based evaluation against objective criteria, or a pairwise comparison method that matches the approach Artificial Analysis uses for benchmarks like GDPval-AA and AA-Briefcase, where users evaluate a sample of response pairs and indicate which answer they prefer before Optima derives the full ranking across the test dataset.
Pricing is straightforward: Optima charges only the actual token costs of models used, with no markup. Rubric-based evaluations cost $0.125 per criterion per model, and pairwise evaluations cost $0.375 per comparison. At the start of benchmark creation, each benchmark run, and each evaluation round, the platform holds a balance based on a cost estimate, then bills based on actual usage and evaluation costs incurred.
The need for custom benchmarking reflects known shortcomings in general-purpose AI evaluations. Research by Epoch AI demonstrated that benchmark results depend heavily on implementation details rarely disclosed—different prompt wording and temperature settings caused the same model to score noticeably differently, and for agentic benchmarks like SWE-bench, swapping the agent's control software and tool environment accounted for up to 15 percentage points of difference. A broader study examining 445 benchmark papers from leading AI conferences found that nearly all had methodological weaknesses in at least one area, including unclear definitions, unrepresentative samples, and missing statistical validation. Only about 10 percent used complete real-world tasks reflecting actual application scenarios. Early testers of Optima have already found value: finance and accounting teams built benchmarks to find which model could cut costs by a factor of ten without major quality loss, and others tested which model best matched the writing style of lawyers or most accurately identified elements in a proprietary image dataset.
Optima addresses a well-documented problem in AI evaluation: public benchmarks often fail to predict which model works best for a specific business need. Research by Epoch AI showed that small implementation details—different prompt wording, temperature settings, or changes to an agent's control software—can shift model scores noticeably. A broader study of 445 benchmark papers from leading conferences found that nearly all had methodological weaknesses, with only about 10 percent using complete real-world tasks that reflected actual application scenarios.
By letting users test models on their own data and workflows, Optima sidesteps this gap. Early users have found concrete value: finance and accounting teams discovered models that cut costs by a factor of ten without major quality loss, while legal teams identified which model best matched their writing style. The platform's focus on cost per completed task and time per task is particularly relevant for agentic applications, where a cheaper model that fails more often or requires extra cleanup work can end up costing more overall than a pricier alternative.
However, Optima does not eliminate the deeper challenges of benchmarking itself. The platform's usefulness still hinges on how precisely target capabilities are defined, how representative test cases are, and how the evaluation is implemented. Even outputs that are fast and cheap may require heavy rework if they add little business value to the process they serve.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Amazon Web Services posted 37% revenue growth in the third quarter, up from roughly 20% in prior years, outpac…

Analysts have raised price targets for Amazon to the US$320 to US$365 range, citing strong AWS (Amazon Web Ser…

A survey by Epoch AI and Ipsos of 1,106 employed US adults (conducted July 10–19, 2026) found that 20 percent…

Ventrova, an AI-agent-operated business, is offering Sentinel Scan, a $249 one-time authorized security audit…

Widen is a new open-source, native PostgreSQL GUI for macOS 14+ that lets users ask questions in English and g…

Privibe is a new open-source fork of Mistral Vibe, a CLI coding agent, redesigned to run entirely on local mac…

The AI news that matters, in one minute each morning.
Sign up free