
LLM as a Judge uses a second AI to rate outputs of another AI. dotData shared its method for making it work.
Traditional software tests are not enough for generative AI.
What happened
dotData explained its approach to using a separate LLM (an AI that understands and generates text) to rate another AI's outputs for quality, in a blog post published on July 8, 2026.
Why it matters
Traditional software tests don't work for generative AI because outputs are not deterministic, so subjective quality like "is this output appropriate" must be judged. dotData emphasized that human-defined evaluation criteria and data, not the AI itself, are the key to reliable quality management.
What to watch
dotData uses a scoring system with a hard rule (OK/NG), a soft rule (0–3 points), and a three-part setup: evaluation rubric, evaluation dataset, and evaluation AI. For example, in a demo, "instant gross profit" scored 3 points (Perfect), while "gross margin" was 2 points (Good) and "unnecessary tax" was 0 points (Failure).
Ask the AI about this article →
The article highlights a fundamental shift in software quality assurance as generative AI becomes more common in business. Unlike traditional software, where inputs yield predictable outputs, generative AI's outputs vary with probability, making simple correctness checks insufficient. To address this, dotData proposes a structured approach where humans define what constitutes good output, and an evaluation AI applies those criteria consistently.
The three essential elements—evaluation rubric, evaluation dataset, and evaluation AI—work together. The rubric clarifies what output should achieve, the dataset includes good and bad examples, and the evaluation AI scores outputs with reasons. dotData notes that while building a prototype achieving about 80 points is relatively easy, maintaining quality through model updates is a challenge.
A practical example shows how scores are assigned: for a retail data analysis, "instant total gross profit" scores 3 points (Perfect), "gross margin" scores 2 points (Good), and "unnecessary tax" scores 0 points (Failure). This demonstrates that even with clear criteria, interpretation matters, and human review of the scores and reasons remains essential.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
ITmedia's ICT Research Division tested ChatGPT, Gemini, and Claude Sonnet 5 in July 2026

Anthropic's Claude AI service is currently experiencing an outage, as indicated by the article title

Snowflake introduced dynamic model routing in Cortex AI Gateway, which automatically selects the most affordab…

Snowflake's Observe platform found that AI assistance can speed up incident investigation by 3x-10x, with an a…

Snowflake CoCo now offers per-user AI credit quotas that are generally available, and introduced three new gov…

OpenAI Group PBC CEO Sam Altman said in an interview with podcaster David Senra that he is worried AI technolo…