
What happened
Researchers shared results from an open-source benchmark measuring cost per correct answer, cost per passing agent task, and rubric-graded professional deliverables across OpenAI models on Amazon Bedrock.
Why it matters
After a July 30, 2026 price cut — luna −80 percent, terra −20 percent — luna cost $0.0021 per correct AIME answer versus mini's $0.0139, and $0.010 per passing GDPval deliverable versus mini's $0.030.
What to watch
The sample sizes are small (48–198 questions), so treat close gaps as directional; the authors say to reproduce the evaluation on your own workload before choosing a model.
WHO IT HITSThis lands on teams already running or evaluating OpenAI models on Amazon Bedrock — especially those using cost-efficient models like gpt-5.4-mini or nano for high-volume, agentic, or document-production workloads — who may need to re-check their model choice against outcome costs, not token prices.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
The starting point for this work is a common habit among organizations building generative AI applications: comparing models on dollars per million tokens because it is the number on every pricing page. The article argues that production workloads actually buy outcomes, and between the pricing page and the outcome sit multipliers the sticker price ignores — how often the model is right, how many tokens it needs to get there, and, for agentic workloads, how many turns it takes. Each extra turn re-sends the growing conversation, so a model that finishes in five turns instead of eight can save more than the reduction in turns alone suggests.
The benchmark harness, called openai-on-aws/benchmarks-openai, runs one identical code path (the OpenAI Responses API) against both Amazon Bedrock and the OpenAI API, switching the backend and model ID while holding evaluation logic constant. The article is explicit that the results reflect differences in the models, provider infrastructure, and model-specific configuration, and that the Amazon Bedrock models ran with reasoning disabled as a deliberate cost floor. Grading combines deterministic checks with an LLM judge using frozen prompts whose hashes are recorded in every result file.
The findings across single-call accuracy, agent trajectories, and professional deliverables consistently point to token and turn efficiency as decisive pricing variables that are invisible on the pricing page. The article also notes that pricing pages move — the July 2026 GPT-5.6 reductions being a case in point — and that the durable part of the work is the methodology, not any single ranking. Whether luna remains the lowest-cost option for a given team hinges on whether its observed efficiency holds on that team's own tasks and configuration, which is why the authors recommend re-running the evaluation rather than assuming the sample results transfer.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Much of the attention on AI infrastructure buildouts is now tied to sheer compute power, with dominance define…

Barron's reported September 10 that Kepler Computing emerged from stealth with a memory architecture using fer…

Dynatrace acquired Arize AI, adding AI observability, evaluation and agent monitoring to its application obser…
Reuters reported September 10 that inference-chip startup d-Matrix will use Nvidia's NVLink Fusion to connect…

A Daily Dose of Data Science test kept LoRA adapters separate from a shared 7B base model, cutting 100 fine-tu…

A report by Spencer Kitts, Thomas Larsen and Sydney Von Arx says an OpenAI agent swarm very likely ran an atta…
