
What happened
Hex used evals to build Quick Edits, a small-model feature that restyles charts 20X faster and cheaper than its full agent. It wrote 1,800 eval cases, and hill-climbing cut wrong edits from 21% to 3%.
Why it matters
Most of the improvement came from the validator, the code that checks and cleans the model's output, not the prompt, suggesting teams building AI features should invest there.
WHO IT HITSProduct and engineering teams building AI-powered features into data or analytics tools should note that validator code, not prompt tweaks, drove most of Hex's measured quality gain.
Summaries like this, in your inbox every morning.
Hex had already used evals for AI features, but Quick Edits was different: evals drove development from the beginning rather than serving as a final check. The feature runs on a small model that reads a request like "make the Enterprise line green" and proposes chart edits in one shot, handing off to the full Hex agent when the request is too complex. Hex started with a few hundred eval cases from real internal requests and grew that to 1,800, using multiple models to synthesize and fill gaps. Holdout cases made up almost 40% of the suite, and changes were accepted only when holdout results improved. Running 3 attempts per case was too noisy, with results differing by about 8 wrong edits, so Hex settled on 10 attempts per case, bringing holdout noise to 0.2 percentage points. Because a full run costs less than $10 and takes a few minutes, Hex reran the suite after every change, including three changes to chart properties in ten days. Notably, prompt changes alone rarely helped; fixes often meant moving logic into the validator or into property labels and descriptions the model reads.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Stanford and Carnegie Mellon researchers tracked 1182 Character.AI users from September 2024 to August 2025, t…

A study presented at USENIX Security Symposium 2026 examined 196,682 resumes from hiring platforms and detecte…

Google Cloud announced Gemini Agent, a cloud-based multi-agent tool

Google Cloud announced Gemini agent, a universal agent that handles answers, knowledge work, media creation, a…

The guide shows code that pulls AAPL prices with yfinance, averages tweet sentiment via TextBlob, then fits a…

The open-source ax-engine for Apple Silicon Macs measured over 76 tok/s decode on one M5 Max setup, and about…
