
What happened
Researchers from ETH Zurich, the New Economic School, and Imperial College London ran a randomized trial with 1,559 judges across 118 Pakistani courts—roughly half of all Pakistani trial court judges. Judges given access to JudgeGPT, an AI assistant built on OpenAI's GPT-4, plus targeted training (six 90-minute lectures over three weeks) resolved about 1,848 more cases per year per district, a 6.3 percent increase. Judges with training used the tool four times as much as those given only general technology seminars.
Why it matters
The study found appeal rates and ruling quality held steady or improved, with no increase in gender or religious bias in judicial language. The researchers estimate a return of $38.50 per dollar invested based on the cost to hire additional judges to match the same output; even conservative estimates put the return at 'at least' $10 per dollar. The results suggest AI can boost public sector productivity when paired with targeted training—but training proved essential, as AI access alone made little difference.
What to watch
Trained judges were steered toward limited support tasks like editing and summarization, where language models are more reliable, rather than broad legal questions where hallucination risk is higher. Only about a fifth of requests involved 'substantive AI delegation' where judges asked the tool to evaluate decisions on its own. The researchers note that today's reasoning models 'write better, hallucinate less, and handle complex tasks more reliably' than GPT-4, suggesting productivity gains in the study may be a floor rather than a ceiling.
Summaries like this, in your inbox every morning.
The trial reveals a nuanced picture of AI adoption in the public sector. Raw access to an AI tool—in this case JudgeGPT, built on GPT-4 and searching a database of 129,235 documents including 128,292 court rulings and 943 Pakistani laws—produced minimal gains. Judges without training averaged only about 20 logins and fewer than 50 prompts over 40 weeks. The decisive factor was the targeted training program, which shifted both usage volume (trained judges averaged nearly 60 logins and over 200 prompts) and usage pattern (toward reliable tasks like text editing, away from risky broad legal questions). This behavioral steering proved crucial: trained judges resolved significantly more cases—1,848 extra per year per district on average, with even bottom-quartile districts clearing about 616 more—while maintaining or improving ruling quality and avoiding bias drift.
The economic case is striking. A return of $38.50 per dollar invested (or conservatively at least $10 per dollar) reflects the cost savings of avoiding new judge hires to achieve the same output. Notably, the gains came without degrading the judicial process: appeal rates per 1,000 resolved cases fell slightly, judges worked the same hours, and work-life balance did not change. The researchers emphasize that these findings do not support replacing judges with AI, but rather demonstrate that thoughtfully deployed AI with proper training can amplify human productivity in constrained public systems. The implication—that training steers users toward appropriate, lower-hallucination tasks while preserving human decision-making—may apply beyond the judiciary to other sectors deploying language models in high-stakes settings.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, LINE, or Slack.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.