AIToday
Large Language ModelsAI Business & IndustryTHE DECODERPublished: Aug 16, 2026, 16:00 JST5 min read

Artificial Analysis launches Optima for custom AI model benchmarks

Artificial Analysis launches Optima for custom AI model benchmarks

Key takeaway

  • Artificial Analysis has released Optima, a benchmarking platform that lets users test AI models against their own data and workflows instead of relying on general-purpose benchmarks.

  • Early testers have used it to find models that cut costs significantly for finance and accounting tasks, and to match specific writing styles in legal work.

  • The platform charges based on actual token usage and evaluation costs, addressing a known limitation of public benchmarks that often fail to capture the specifics of individual use cases.

3 Key Points

  1. What happened

    Artificial Analysis, known for independent LLM evaluations, has launched Optima, a platform that lets users build custom benchmarks to compare AI models using their own data and workflows. Users can upload datasets, describe use cases, or provide sample inputs, then test models on quality, cost per task, and time per task. The platform is available now.

  2. Why it matters

    Public benchmarks use predefined tasks that may not reflect a specific business need. Optima addresses this by letting organizations test models against their actual workflows—finance teams found models that cut costs by a factor of ten without major quality loss, and legal teams tested which model best matched their writing style. For agentic applications, cost per completed task often matters more than raw token price, since a cheaper model can end up more expensive if it fails more often.

  3. What to watch

    Optima charges only actual token costs with no markup—rubric-based evaluations cost $0.125 per criterion per model, and pairwise evaluations cost $0.375 per comparison. However, the platform's usefulness still depends on how precisely capabilities are defined and how representative test cases are; even cheap and fast outputs may need heavy rework if they add little business value.

In Depth

Read the full story

Artificial Analysis, which operates well-known LLM evaluations including GDPval-AA and AA-Briefcase, has released a new platform called Optima designed to close a gap in AI model comparison. While public benchmarks test models on predefined tasks, they often don't reveal which model actually performs best for a particular organization's needs. Optima allows users to build their own benchmarks using their specific data and workflows, then run tests across leading current models and compare results on quality, cost per task, and time per task.

The platform offers three paths to building a custom benchmark. Users can upload existing evaluation datasets from their own files or from Hugging Face, provide AI agent traces from platforms like Arize, Braintrust, or Langfuse, or install a skill that gathers information from their coding environment and past sessions. For those without ready-made data, Optima accepts a description of the intended use case along with sample inputs and outputs, then generates suggested test inputs, evaluation criteria, and example tasks for users to review and refine. Two scoring approaches are available: rubric-based evaluation against objective criteria, or a pairwise comparison method that matches the approach Artificial Analysis uses for benchmarks like GDPval-AA and AA-Briefcase, where users evaluate a sample of response pairs and indicate which answer they prefer before Optima derives the full ranking across the test dataset.

Pricing is straightforward: Optima charges only the actual token costs of models used, with no markup. Rubric-based evaluations cost $0.125 per criterion per model, and pairwise evaluations cost $0.375 per comparison. At the start of benchmark creation, each benchmark run, and each evaluation round, the platform holds a balance based on a cost estimate, then bills based on actual usage and evaluation costs incurred.

The need for custom benchmarking reflects known shortcomings in general-purpose AI evaluations. Research by Epoch AI demonstrated that benchmark results depend heavily on implementation details rarely disclosed—different prompt wording and temperature settings caused the same model to score noticeably differently, and for agentic benchmarks like SWE-bench, swapping the agent's control software and tool environment accounted for up to 15 percentage points of difference. A broader study examining 445 benchmark papers from leading AI conferences found that nearly all had methodological weaknesses in at least one area, including unclear definitions, unrepresentative samples, and missing statistical validation. Only about 10 percent used complete real-world tasks reflecting actual application scenarios. Early testers of Optima have already found value: finance and accounting teams built benchmarks to find which model could cut costs by a factor of ten without major quality loss, and others tested which model best matched the writing style of lawyers or most accurately identified elements in a proprietary image dataset.

Context & Analysis

Optima addresses a well-documented problem in AI evaluation: public benchmarks often fail to predict which model works best for a specific business need. Research by Epoch AI showed that small implementation details—different prompt wording, temperature settings, or changes to an agent's control software—can shift model scores noticeably. A broader study of 445 benchmark papers from leading conferences found that nearly all had methodological weaknesses, with only about 10 percent using complete real-world tasks that reflected actual application scenarios.

By letting users test models on their own data and workflows, Optima sidesteps this gap. Early users have found concrete value: finance and accounting teams discovered models that cut costs by a factor of ten without major quality loss, while legal teams identified which model best matched their writing style. The platform's focus on cost per completed task and time per task is particularly relevant for agentic applications, where a cheaper model that fails more often or requires extra cleanup work can end up costing more overall than a pricier alternative.

However, Optima does not eliminate the deeper challenges of benchmarking itself. The platform's usefulness still hinges on how precisely target capabilities are defined, how representative test cases are, and how the evaluation is implemented. Even outputs that are fast and cheap may require heavy rework if they add little business value to the process they serve.

FAQ

What types of data can I use to build a benchmark in Optima?
Users can upload existing evaluation datasets from their own files or from Hugging Face, provide AI agent traces from platforms like Arize, Braintrust, or Langfuse, or describe their use case with sample inputs and outputs—Optima will then generate suggested test inputs and evaluation criteria for review.
How much does it cost to run a benchmark?
Optima charges only the actual token costs of models used with no markup. Rubric-based evaluations cost $0.125 per criterion per model, and pairwise evaluations cost $0.375 per comparison.
What metrics does Optima measure beyond model quality?
Optima tracks cost per task and time per task as standalone comparison dimensions, making it possible to check whether a performance gain justifies the higher cost or longer processing time of a given model.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleData Centers' Power Hunger Turns Energy Stocks Into AI Bets

The AI news that matters, in one minute each morning.

Sign up free