
A randomized trial with Pakistani judges showed that pairing AI access with targeted training delivered substantial gains: judges with training used JudgeGPT four times more than those without, and their districts resolved about 1,848 extra cases per year—a 6.3 percent increase. Ruling quality held steady or improved, and the return on investment was estimated at $38.50 per dollar spent. The key finding is that training mattered critically; AI access alone produced little benefit.
Summaries like this, in your inbox every morning.
Sign up free →What happened
Researchers from ETH Zurich, the New Economic School, and Imperial College London ran a randomized trial with 1,559 judges across 118 Pakistani courts—roughly half of all Pakistani trial court judges. Judges given access to JudgeGPT, an AI assistant built on OpenAI's GPT-4, plus targeted training (six 90-minute lectures over three weeks) resolved about 1,848 more cases per year per district, a 6.3 percent increase. Judges with training used the tool four times as much as those given only general technology seminars.
Why it matters
The study found appeal rates and ruling quality held steady or improved, with no increase in gender or religious bias in judicial language. The researchers estimate a return of $38.50 per dollar invested based on the cost to hire additional judges to match the same output; even conservative estimates put the return at 'at least' $10 per dollar. The results suggest AI can boost public sector productivity when paired with targeted training—but training proved essential, as AI access alone made little difference.
What to watch
Trained judges were steered toward limited support tasks like editing and summarization, where language models are more reliable, rather than broad legal questions where hallucination risk is higher. Only about a fifth of requests involved 'substantive AI delegation' where judges asked the tool to evaluate decisions on its own. The researchers note that today's reasoning models 'write better, hallucinate less, and handle complex tasks more reliably' than GPT-4, suggesting productivity gains in the study may be a floor rather than a ceiling.
A multinational research team from ETH Zurich, the New Economic School, and Imperial College London conducted a large-scale field experiment in Pakistan's judiciary, enrolling 1,559 judges across 118 courts—representing roughly half of all Pakistani trial court judges. The study tested JudgeGPT, an AI assistant built on OpenAI's GPT-4 that uses retrieval augmented generation to search a database of 129,235 documents: 128,292 court rulings and 943 Pakistani laws. When a judge enters a query, the tool retrieves the ten most relevant passages and generates a cited answer.
Researchers divided judges into three groups. The first received JudgeGPT access plus targeted training: six 90-minute lectures delivered over three weeks by ETH Professor Elliott Ash, held after court hours. The curriculum taught judges which tasks suited the tool, where it fell short, and how to verify its output. A second group received the same AI access but only a general seminar on technology and law. The control group attended the technology seminar with no JudgeGPT access. After 40 weeks, the contrast was stark: trained judges averaged nearly 60 logins and over 200 prompts, while the comparison group (with AI access but only the general seminar) averaged about 20 logins and fewer than 50 prompts. AI access alone produced minimal usage.
The results on case resolution were substantial. Districts with more trained judges resolved significantly more cases. At moderate exposure levels, trained judges' districts cleared about 1,848 extra cases per year per district, a 6.3 percent increase. Even districts in the bottom quartile—the least exposed to trained judges—still resolved about 616 more cases annually. Importantly, these gains came without compromising quality or working conditions. The appeal rate per 1,000 resolved cases fell slightly, judges worked the same hours, and work-life balance remained unchanged. An analysis of roughly 4,000 court judgments found that while AI-flagged text appeared more frequently (as expected), readability, length, and the number of legal arguments held steady. An LLM-based quality check validated by two Pakistani lawyers showed a slight improvement: rulings from trained judges were rated better in 59 percent of pairwise comparisons, compared to 42 percent in the control group. Crucially, the study found no evidence that AI use increased gender or religious bias in judicial language.
Analysis of anonymized chat logs from about 1,500 judges revealed that legal research, text editing, and text generation were the most common tasks, with around 60 percent of queries seeking information about laws, procedures, or legal concepts. Training shifted behavior significantly: trained judges used JudgeGPT more for editing and summarizing text—tasks where language models are reliable—and asked fewer broad legal questions, where hallucination risk is higher. Only about a fifth of requests involved what researchers call 'substantive AI delegation,' where judges asked the tool to evaluate decisions or draft reasoning independently. Training made judges more likely to decide cases themselves and use AI only to write up their reasoning. Economically, researchers estimate a return of $38.50 per dollar invested, based on the cost to hire additional judges to produce the same output; even conservative estimates place the return at 'at least' $10 per dollar. The researchers stress that their findings do not support replacing judges with AI, but instead demonstrate that AI can boost public sector productivity when paired with training that directs users toward appropriate tasks. They note that JudgeGPT ran on GPT-4, a pre-reasoning model that was the best available for the project at the time, and that today's reasoning models 'write better, hallucinate less, and handle complex tasks more reliably,' suggesting the productivity gains in this study may be a floor rather than a ceiling.
The trial reveals a nuanced picture of AI adoption in the public sector. Raw access to an AI tool—in this case JudgeGPT, built on GPT-4 and searching a database of 129,235 documents including 128,292 court rulings and 943 Pakistani laws—produced minimal gains. Judges without training averaged only about 20 logins and fewer than 50 prompts over 40 weeks. The decisive factor was the targeted training program, which shifted both usage volume (trained judges averaged nearly 60 logins and over 200 prompts) and usage pattern (toward reliable tasks like text editing, away from risky broad legal questions). This behavioral steering proved crucial: trained judges resolved significantly more cases—1,848 extra per year per district on average, with even bottom-quartile districts clearing about 616 more—while maintaining or improving ruling quality and avoiding bias drift.
The economic case is striking. A return of $38.50 per dollar invested (or conservatively at least $10 per dollar) reflects the cost savings of avoiding new judge hires to achieve the same output. Notably, the gains came without degrading the judicial process: appeal rates per 1,000 resolved cases fell slightly, judges worked the same hours, and work-life balance did not change. The researchers emphasize that these findings do not support replacing judges with AI, but rather demonstrate that thoughtfully deployed AI with proper training can amplify human productivity in constrained public systems. The implication—that training steers users toward appropriate, lower-hallucination tasks while preserving human decision-making—may apply beyond the judiciary to other sectors deploying language models in high-stakes settings.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion





Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime
1 minute a day. The AI essentials.
200+ sources · Email / LINE / Slack