AIToday
Large Language ModelsAI Coding AssistantsHacker NewsPublished: Oct 9, 2026, 10:00 JST

Hex cuts wrong edits 7x with evals on Quick Edits

Hex cuts wrong edits 7x with evals on Quick Edits

3 Key Points

  1. What happened

    Hex used evals to build Quick Edits, a small-model feature that restyles charts 20X faster and cheaper than its full agent. It wrote 1,800 eval cases, and hill-climbing cut wrong edits from 21% to 3%.

  2. Why it matters

    Most of the improvement came from the validator, the code that checks and cleans the model's output, not the prompt, suggesting teams building AI features should invest there.

WHO IT HITSProduct and engineering teams building AI-powered features into data or analytics tools should note that validator code, not prompt tweaks, drove most of Hex's measured quality gain.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Hex had already used evals for AI features, but Quick Edits was different: evals drove development from the beginning rather than serving as a final check. The feature runs on a small model that reads a request like "make the Enterprise line green" and proposes chart edits in one shot, handing off to the full Hex agent when the request is too complex. Hex started with a few hundred eval cases from real internal requests and grew that to 1,800, using multiple models to synthesize and fill gaps. Holdout cases made up almost 40% of the suite, and changes were accepted only when holdout results improved. Running 3 attempts per case was too noisy, with results differing by about 8 wrong edits, so Hex settled on 10 attempts per case, bringing holdout noise to 0.2 percentage points. Because a full run costs less than $10 and takes a few minutes, Hex reran the suite after every change, including three changes to chart properties in ten days. Notably, prompt changes alone rarely helped; fixes often meant moving logic into the validator or into property labels and descriptions the model reads.

FAQ
How much did the wrong-edit rate improve?
It fell from 21% at the start of hill-climbing to 3% by the end, according to Hex.
What is the validator and why did it matter?
It is code that checks and cleans the model's output before editing the chart. Validator improvements raised pass rates by 7 percentage points on Luna and 10 on Haiku and cut Haiku's hard errors by 99%.
How much does a full eval run cost?
Even at 18,000 attempts per full run, a run costs less than $10 on Luna and takes a few minutes at high concurrency.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articlePeter Thiel blasts Pope Leo XIV over AI fear