Daily Tech Briefing
AI 科技速览
每天 5 分钟内学习 AI。获取最新的人工智能新闻,理解其重要性,并学习如何将其应用于您的工作。
The Decoder · 2026/7/21 19:12:20

An AI system helped Pakistani judges clear massive backlogs at $38.50 return per dollar invested
AI 中文解读
巴基斯坦法院用AI清理积压案件,每投入1美元竟赚回38.5美元!这场由苏黎世联邦理工、帝国理工等机构主导的大规模实验,让1559名法官用上了名为JudgeGPT的AI助手。不过,光给AI工具没用——只有接受过近三周针对性培训的法官才真正用起来,每周登录近60次、输入200多个指令。结果令人振奋:每个地区法院每年多处理1848起案件,工作效率提升6.3%,而且判决质量不降反升,上诉率还略有下降。法官的工作时长和生活平衡也没受影响。这背后是AI通过检索12.9万份法律文件,帮法官快速找到相关判例和法律条文。对普通人来说,这意味着打官司的等待时间有望大幅缩短,司法系统不再“案等人”,政府也能用更低成本提供更高效的公共服务。
An AI system helped Pakistani judges clear massive backlogs at $38.50 return per dollar invested
Matthias Bastian
View the LinkedIn Profile of Matthias Bastian
Jul 21, 2026
Nano Banana Pro prompted by THE DECODER
Can AI make government institutions more productive? A study from researchers at ETH Zurich, Imperial College London, and the New Economic School delivers the strongest experimental evidence yet.
Pakistan has fewer than two judges per 100,000 residents, according to the authors. The EU has 22. England and Wales have 30. At the end of 2024, 2.26 million cases were pending, 82 percent of them in trial courts. Judges work with bare-bones tech and no support staff. Before the experiment, only about 25 percent had ever used a large language model like ChatGPT.
Researchers from ETH Zurich, the New Economic School, and Imperial College London ran a large-scale field experiment with Pakistan's judiciary. The randomized trial covered 1,559 judges across 118 courts, roughly half of all Pakistani trial court judges.
The tool was JudgeGPT, an AI assistant built on OpenAI's GPT-4 and designed for Pakistani trial courts. It uses retrieval augmented generation to search a database of 129,235 documents, including 128,292 court rulings and 943 Pakistani laws. When a judge enters a query, JudgeGPT picks the ten most relevant passages and generates a cited answer.
Trained judges resolve 1,848 more cases per year per district
The researchers split judges into three groups. One got JudgeGPT access plus targeted training: six 90-minute lectures over three weeks, taught by ETH Professor Elliott Ash after court hours. Judges learned which tasks suited the tool, where it fell short, and how to check its output.
A second group got the same AI access but only a general seminar on technology and law. The control group attended that seminar with no JudgeGPT access.
AI access alone did little. Judges with targeted training used JudgeGPT four times as much as those in the general seminar group. After 40 weeks, trained judges averaged nearly 60 logins and over 200 prompts. The comparison group averaged about 20 logins and fewer than 50 prompts.
Districts with more trained judges resolved more cases. At moderate exposure levels, that meant about 1,848 extra cases per year per district, a 6.3 percent bump. Even districts in the bottom quartile still cleared about 616 more cases.
Judgment quality ticks up without added bias
Ruling quality held steady or improved. The appeal rate per 1,000 resolved cases fell slightly, and judges worked the same hours with no change in work-life balance. The researchers estimate savings of about $38.50 per dollar invested, based on what it would cost to hire enough extra judges to match the same output. Even conservative estimates put the return at "at least" $10 per dollar.
A review of roughly 4,000 court judgments found more AI-flagged text, as expected. But readability, length, and the number of legal arguments held steady. An LLM-based quality check, validated by two Pakistani lawyers, showed a slight improvement. Rulings from trained judges were rated better in 59 percent of pairwise comparisons, up from 42 percent in the control group. The study found no evidence that AI use increased gender or religious bias in judicial language.
Hands-on training shapes how judges use AI
The researchers reviewed anonymized chat logs from about 1,500 judges. Legal research, text editing, and text generation were the most common tasks. Around 60 percent of queries sought information about laws, procedures, or legal concepts.
Trained judges used JudgeGPT more for editing and summarizing text, tasks where language models are more reliable. They asked fewer broad legal questions, where hallucination risk is higher. The researchers say the training
分享
阅读原文 ↗