AIToday
THE DECODERPublished: Jul 22, 2026, 06:00 JST

Pakistani judges clear 1,848 extra cases yearly with AI training, $38.50 return per dollar

Pakistani judges clear 1,848 extra cases yearly with AI training, $38.50 return per dollar

3 Key Points

  1. What happened

    Researchers from ETH Zurich, the New Economic School, and Imperial College London ran a randomized trial with 1,559 judges across 118 Pakistani courts—roughly half of all Pakistani trial court judges. Judges given access to JudgeGPT, an AI assistant built on OpenAI's GPT-4, plus targeted training (six 90-minute lectures over three weeks) resolved about 1,848 more cases per year per district, a 6.3 percent increase. Judges with training used the tool four times as much as those given only general technology seminars.

  2. Why it matters

    The study found appeal rates and ruling quality held steady or improved, with no increase in gender or religious bias in judicial language. The researchers estimate a return of $38.50 per dollar invested based on the cost to hire additional judges to match the same output; even conservative estimates put the return at 'at least' $10 per dollar. The results suggest AI can boost public sector productivity when paired with targeted training—but training proved essential, as AI access alone made little difference.

  3. What to watch

    Trained judges were steered toward limited support tasks like editing and summarization, where language models are more reliable, rather than broad legal questions where hallucination risk is higher. Only about a fifth of requests involved 'substantive AI delegation' where judges asked the tool to evaluate decisions on its own. The researchers note that today's reasoning models 'write better, hallucinate less, and handle complex tasks more reliably' than GPT-4, suggesting productivity gains in the study may be a floor rather than a ceiling.

Not sure about something? Ask the AI

Summaries like this, in your inbox every morning.

Context & Analysis

The trial reveals a nuanced picture of AI adoption in the public sector. Raw access to an AI tool—in this case JudgeGPT, built on GPT-4 and searching a database of 129,235 documents including 128,292 court rulings and 943 Pakistani laws—produced minimal gains. Judges without training averaged only about 20 logins and fewer than 50 prompts over 40 weeks. The decisive factor was the targeted training program, which shifted both usage volume (trained judges averaged nearly 60 logins and over 200 prompts) and usage pattern (toward reliable tasks like text editing, away from risky broad legal questions). This behavioral steering proved crucial: trained judges resolved significantly more cases—1,848 extra per year per district on average, with even bottom-quartile districts clearing about 616 more—while maintaining or improving ruling quality and avoiding bias drift.

The economic case is striking. A return of $38.50 per dollar invested (or conservatively at least $10 per dollar) reflects the cost savings of avoiding new judge hires to achieve the same output. Notably, the gains came without degrading the judicial process: appeal rates per 1,000 resolved cases fell slightly, judges worked the same hours, and work-life balance did not change. The researchers emphasize that these findings do not support replacing judges with AI, but rather demonstrate that thoughtfully deployed AI with proper training can amplify human productivity in constrained public systems. The implication—that training steers users toward appropriate, lower-hallucination tasks while preserving human decision-making—may apply beyond the judiciary to other sectors deploying language models in high-stakes settings.

FAQ
How much training did judges receive?
Judges in the trained group received six 90-minute lectures over three weeks, taught by ETH Professor Elliott Ash after court hours. The training covered which tasks suited the tool, where it fell short, and how to check its output.
Did AI use change the quality of rulings or introduce bias?
An LLM-based quality check validated by two Pakistani lawyers showed a slight improvement: rulings from trained judges were rated better in 59 percent of pairwise comparisons, up from 42 percent in the control group. The study found no evidence that AI use increased gender or religious bias in judicial language.
What tasks did judges use JudgeGPT for most?
Legal research, text editing, and text generation were the most common tasks. Around 60 percent of queries sought information about laws, procedures, or legal concepts. Trained judges used the tool more for editing and summarizing, tasks where language models are more reliable, and asked fewer broad legal questions where hallucination risk is higher.

Get AI news like this every morning

For example, today's edition would include:

  • SpaceX launches Grok 4.7 at $4.69 per task, beating rivalsSiliconANGLE AI · 40m ago
  • Amazon blocks Meta's Muse agent from its marketplaceSiliconANGLE AI · 40m ago
  • Tesla audits China suppliers for Optimus mass productionDIGITIMES Asia · 40m ago

AI-summarized, only the topics you pick — one digest a day via Email, LINE, or Slack.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Next articleOpenAI models hacked Hugging Face to cheat on security test