AIToday
Large Language ModelsAI Safety & AlignmentITmedia AI+Published: Aug 24, 2026, 16:01 JST2 min read

AI judges AI output quality with LLM-as-a-Judge

AI judges AI output quality with LLM-as-a-Judge

Key takeaway

  • LLM as a Judge uses a second AI to rate outputs of another AI. dotData shared its method for making it work.

  • Traditional software tests are not enough for generative AI.

3 Key Points

  1. What happened

    dotData explained its approach to using a separate LLM (an AI that understands and generates text) to rate another AI's outputs for quality, in a blog post published on July 8, 2026.

  2. Why it matters

    Traditional software tests don't work for generative AI because outputs are not deterministic, so subjective quality like "is this output appropriate" must be judged. dotData emphasized that human-defined evaluation criteria and data, not the AI itself, are the key to reliable quality management.

  3. What to watch

    dotData uses a scoring system with a hard rule (OK/NG), a soft rule (0–3 points), and a three-part setup: evaluation rubric, evaluation dataset, and evaluation AI. For example, in a demo, "instant gross profit" scored 3 points (Perfect), while "gross margin" was 2 points (Good) and "unnecessary tax" was 0 points (Failure).

Ask the AI about this article →

Context & Analysis

The article highlights a fundamental shift in software quality assurance as generative AI becomes more common in business. Unlike traditional software, where inputs yield predictable outputs, generative AI's outputs vary with probability, making simple correctness checks insufficient. To address this, dotData proposes a structured approach where humans define what constitutes good output, and an evaluation AI applies those criteria consistently.

The three essential elements—evaluation rubric, evaluation dataset, and evaluation AI—work together. The rubric clarifies what output should achieve, the dataset includes good and bad examples, and the evaluation AI scores outputs with reasons. dotData notes that while building a prototype achieving about 80 points is relatively easy, maintaining quality through model updates is a challenge.

A practical example shows how scores are assigned: for a retail data analysis, "instant total gross profit" scores 3 points (Perfect), "gross margin" scores 2 points (Good), and "unnecessary tax" scores 0 points (Failure). This demonstrates that even with clear criteria, interpretation matters, and human review of the scores and reasons remains essential.

FAQ

How does LLM as a Judge work?
It uses a separate evaluation AI that references a rubric and an evaluation dataset to score the output of a functional AI. The evaluation AI returns a score and a reason, which the team reviews.
What is the difference between hard rule and soft rule in evaluation?
A hard rule uses OK/NG to decide if the output meets required conditions. A soft rule rates quality on a scale like 0–3 points, such as 3 points for Perfect and 2 points for Good.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleTaiwan indicts nine over AI server exports to China

The AI news that matters, in one minute each morning.

Sign up free