An AI system helped Pakistani judges clear massive backlogs at $38.50 return per dollar invested

Posted on

The judiciary of Pakistan has become the site of one of the world’s largest and most comprehensive field experiments regarding the integration of artificial intelligence into the public sector. In a collaborative effort between researchers from ETH Zurich, the New Economic School, and Imperial College London, a randomized controlled trial was conducted involving 1,559 judges across 118 courts. This cohort represents approximately half of all trial court judges currently serving in Pakistan. The study, titled "Courts of Tomorrow: Evidence from a Nationwide Rollout of Generative AI," provides a rigorous empirical foundation for understanding how Large Language Models (LLMs) can transform legal productivity, provided they are accompanied by specialized human training.

The core of the experiment centered on the deployment of JudgeGPT, a bespoke AI assistant built upon OpenAI’s GPT-4 architecture. Unlike standard consumer-facing AI, JudgeGPT was specifically tailored for the Pakistani legal landscape using Retrieval-Augmented Generation (RAG). This system allows the AI to query a massive, localized database consisting of 129,235 documents, which includes 128,292 historical court rulings and 943 specific Pakistani laws. By processing these documents, the tool provides judges with the ten most relevant legal passages for any given query, generating a synthesized answer complete with citations. This ensures that the AI’s output is grounded in actual Pakistani jurisprudence rather than general patterns found in global training data.

The Critical Role of Targeted Training

The experiment was structured to isolate the impact of training on AI adoption. The 1,559 judges were divided into three distinct groups to measure varying levels of intervention. The first group received full access to JudgeGPT along with a "targeted training" regimen. This curriculum, led by ETH Professor Elliott Ash, consisted of six 90-minute lectures delivered over a three-week period. These sessions were conducted after court hours and focused on practical applications: identifying tasks suitable for AI, understanding the limitations and risks of LLMs (such as hallucinations), and mastering the art of verifying AI-generated output.

The second group was granted the same access to JudgeGPT but received only a general seminar on the intersection of technology and law, lacking the specific hands-on instruction provided to the first group. Finally, a control group attended the general technology seminar but was given no access to the JudgeGPT tool.

The results highlighted a stark disparity in engagement. Judges who received the targeted training utilized the AI tool four times as often as those who were merely given access. Over a 40-week period, the trained judges averaged nearly 60 logins and over 200 prompts, whereas the group with only general information averaged roughly 20 logins and fewer than 50 prompts. This finding suggests that simply providing advanced technology to public sector employees is insufficient; specialized training is the primary driver of adoption and effective utilization.

Quantifiable Productivity Gains and Economic Impact

The implementation of JudgeGPT led to measurable improvements in the efficiency of the Pakistani court system, which has long struggled with a significant backlog of cases. In districts where judges received targeted training, the resolution rate of cases saw a marked increase. At moderate levels of exposure to the tool, districts resolved an average of 1,848 additional cases per year, representing a 6.3 percent increase in total output. Even in the bottom quartile of districts—those with the lowest levels of AI engagement—there was still a notable increase of approximately 616 resolved cases per year.

Beyond the raw number of cases resolved, the researchers performed a cost-benefit analysis to determine the economic viability of the AI rollout. They estimated a return on investment (ROI) of approximately $38.50 for every dollar spent on the system and training. This figure is based on the comparative cost of hiring the additional human judges that would be required to achieve an equivalent increase in case throughput. Even when applying the most conservative estimates and accounting for various overheads, the researchers concluded that the return remains at least $10 for every dollar invested, making a compelling case for the fiscal efficiency of AI in the public sector.

An AI system helped Pakistani judges clear massive backlogs at $38.50 return per dollar invested

Maintaining Judicial Quality and Mitigating Bias

A primary concern regarding the use of AI in the legal system is the potential for a decline in the quality of justice or the introduction of algorithmic bias. To address this, the researchers conducted a detailed review of approximately 4,000 court judgments. While they found an expected increase in AI-flagged text within the rulings of the treatment group, other indicators suggested that the quality of work remained stable or improved.

Key findings regarding quality included:

  • Appeal Rates: The rate of appeals per 1,000 resolved cases showed a slight decrease, suggesting that the AI-assisted rulings were no less legally sound than those produced through traditional methods.
  • Readability and Argumentation: The length of judgments, the complexity of legal arguments, and general readability scores held steady, indicating that judges were not using AI to "shortcut" the depth of their legal reasoning.
  • Peer Validation: An LLM-based quality assessment, which was cross-validated by two experienced Pakistani lawyers, indicated a slight uptick in the perceived quality of rulings. In pairwise comparisons, rulings from trained judges were preferred in 59 percent of cases, compared to 42 percent in the control group.
  • Bias Analysis: Perhaps most significantly, the study found no evidence that the use of JudgeGPT increased gender or religious bias in the language of judicial rulings. This is a critical finding for a diverse nation like Pakistan, where the equitable application of the law is paramount.

Behavioral Shifts: How Judges Use AI

The researchers gained unique insights into the "black box" of judicial work by analyzing anonymized chat logs from approximately 1,500 judges. The data revealed that legal research, text editing, and text generation were the primary use cases. About 60 percent of all queries were directed toward retrieving information about specific laws, procedures, or legal concepts.

However, the nature of these queries shifted based on training. Trained judges were more likely to use JudgeGPT for linguistic tasks such as editing and summarization—areas where LLMs are known to excel with high reliability. Conversely, they were less likely to ask broad, open-ended legal questions that carry a higher risk of "hallucination" (the generation of false or misleading information).

The study also tracked "substantive AI delegation," defined as instances where a judge asks the AI to evaluate a decision or draft the core reasoning of a case independently. Only about 20 percent of requests fell into this category. Interestingly, the targeted training made judges less likely to delegate the final decision to the AI. Instead, they used the tool to assist in writing up the reasoning for decisions they had already reached independently. This suggests that training fosters a "human-in-the-loop" approach, where the AI serves as a high-powered clerk rather than a replacement for the judge’s own cognitive and moral labor.

Broader Implications for Global Governance

The success of the JudgeGPT experiment in Pakistan offers a blueprint for other nations facing judicial backlogs and administrative inefficiencies. The researchers emphasized that the findings do not suggest that AI can or should replace human judges. Rather, the experiment demonstrates that AI can serve as a potent "productivity multiplier" in the public sector when the rollout is managed with a focus on human skill development.

One of the most promising aspects of the study is the suggestion that these productivity gains may represent a "floor" rather than a "ceiling." The experiment was conducted using GPT-4, which has since been surpassed by "reasoning" models like OpenAI’s o1 series. These newer models are designed with better internal logic chains, lower hallucination rates, and a superior ability to handle complex, multi-step legal reasoning. As the underlying technology improves, the potential for even greater efficiency gains in the legal sector grows.

The Pakistan field experiment proves that the bottleneck for AI adoption in government is not necessarily the technology itself, but the lack of institutional knowledge on how to use it safely and effectively. By investing in targeted training, public institutions can unlock significant value, clearing backlogs and improving the delivery of services without compromising the integrity of their core missions. For the global legal community, the results from Pakistan serve as a landmark case study in the responsible and effective integration of generative AI into the halls of justice.

Leave a Reply

Your email address will not be published. Required fields are marked *