Questions with a reference answer.
Legal-exam questions from four jurisdictions, scored for substantive consistency with the reference answer on a 0–4 scale.
- United States
- 603
- China
- 94
- United Kingdom
- 86
- Australia
- 78
AIMS WORKSHOP · COLM 2026
Measuring Exam-to-Case Transfer
in LLM Legal Reasoning
How much does success on a legal exam tell us
about reasoning through a real case?
1 UC Santa Barbara2 Yale University3 UC Berkeley
* Corresponding author
Downloadable now: 78 source records + benchmark metadata · Data availability and release notes
THE RESEARCH QUESTION
Public legal exams are widely used to evaluate legal reasoning in large language models, yet it remains unclear whether exam performance reflects the ability to reason over real-case facts. We introduce LegalScope, a dual-track benchmark that pairs 861 legal-exam questions with 276 lawyer-reviewed issue-stance prompts derived from de-identified Chinese judgments. LegalScope evaluates models under both reference-answer scoring and case-based rubric scoring, enabling direct analysis of exam-to-case transfer.
Across 28 model groups, public-exam scores correlate with real-case scores but do not fully predict them; citation relevance receives lower mean scores than argument validity under both automatic and lawyer evaluation. We further show that automated evaluation aligns strongly with human review on exam answers but less reliably on real-case legal analysis, highlighting the need for expert-grounded case evaluation.
Explore the workshop paperBENCHMARK
Two tracks, evaluated on the same 28 model groups.
Legal-exam questions from four jurisdictions, scored for substantive consistency with the reference answer on a 0–4 scale.
138 legal issues from 56 de-identified Chinese judgments. Each issue becomes a supporting and an opposing prompt, preserving a closed factual record.
Seven categories: Tort, Contract, Criminal, Intellectual Property, Administrative, Civil Procedure and Property.

Does the selected authority respond to the issue?
Does the response follow the stance, facts and task constraints?
Does the conclusion follow from the law and the supplied facts?
The real-case rubric adapts the three evaluation dimensions of CourtReasoner (Han et al., EMNLP 2025) to closed-book Chinese case analysis.
Read the input, inspect an answer, and see what the scoring protocol measures.
AUSTRALIA · VICTORIAN BAR · EXTVICBAR01A
A barrister is representing a celebrity client who issued proceedings for breach of contract as a result of sexual harassment from their employer. During the course of the retainer, the client informs the barrister that the client is considering divorcing their spouse.
(a) Could the barrister disclose the information about the potential divorce to other people? Why or why not? [1 mark]
Question 1(a), 13 October 2024 entrance exam, p. 3. © 2024 Victorian Bar Inc. Original publication · CC BY-NC-ND 4.0. Line breaks normalized; no endorsement implied.
Whether a model reaches a compatible conclusion and explains the decisive legal point: here, the duty of confidentiality and its application to information obtained during the retainer. Evaluation compares substance with the reference; it does not require identical wording.
CHINESE JUDGMENT · PAPER EXAMPLE RV038
In a medical-malpractice dispute, an appraisal had already been issued. The plaintiff challenged it and applied for re-appraisal. The court refused that request and relied on the existing appraisal materials.
The prompt does not supply the detailed appraisal procedures or the specific medical records reviewed.
Write a two-paragraph Chinese legal analysis supporting the court’s refusal to grant re-appraisal. Use only the supplied facts.
First summarize the relevant facts and dispute; then apply the legal authorities. Do not add medical records, appraisal defects, treatment details or outside case information.
The paper’s scoring packet identifies Article 40 of the Supreme People’s Court provisions on civil evidence: dissatisfaction with a conclusion alone does not establish grounds for re-appraisal. The answer must connect the applicable threshold to the facts actually supplied.
Abbreviated research example from the current manuscript’s RV038 scoring packet and case study. No original judgment or identifying case details are reproduced.
A plausible legal argument can still rely on invented factual support. Separate scores make that boundary violation visible instead of hiding it inside a single overall rating.
RESULTS
0.817
Model-level Pearson correlation for automatic scores; Spearman correlation is 0.708. Rankings and variant gains do not transfer uniformly.
66.9 / 72.6
Mean automatic scores on a 0–100 scale. Lawyers show the same direction descriptively; the pooled lawyer interval crosses zero.
0.910 / 0.312
Answer-level Pearson correlations between automatic and human scores. Case scores use the equal-weight mean of two lawyers.
Loading published results…
| Automatic evaluation | Human evaluation | ||||||
|---|---|---|---|---|---|---|---|
All scores are on a 0–100 scale. Automatic scores cover the full tracks. Human scores cover 80 exam items and 10 case prompts; case scores pool two lawyers. Best and second-best marks use the displayed one-decimal values across all 28 groups; equal values share a mark, and filtering does not change their rank. The initial order follows the paper. CSV/Parquet additionally include the two descriptive overall means. Highlights are descriptive, not significance tests. Scores do not equate difficulty across tracks.



RESOURCES
Research resources and current release information.
Method, full experimental results, rubric calibration and expert validation.
OpenReview02Documentation, original paper figures, aggregate results and workbook helpers.
GitHub0378 licensed source excerpts and selected candidate answers, model results, benchmark composition and dataset card.
Hugging FaceData availability. Hugging Face provides 78 Victorian Bar source-excerpt records and selected candidate answers under CC BY-NC-ND 4.0, alongside metadata and aggregate results. The four dataset configurations load independently; original CSV and JSONL files are also available. Source excerpts preserve attribution and license notices and are not exact replacements for the paper’s evaluation prompts. The full 861-item exam track, the 276-prompt case dataset, full model responses and lawyer review sheets are not included. Case prompts and answers are being considered for item-level source and privacy review; this is not a blanket prohibition on publishing anonymized research material. These materials cannot rerun the complete evaluation. Source-specific release details · Loading examples, fields and answer types.
CITATION
AI Measurement Science Workshop
at COLM 2026
@inproceedings{wang2026legalscope,
title = {{LegalScope}: Measuring Exam-to-Case
Transfer in {LLM} Legal Reasoning},
author = {Wang, Hongyu and Han, Rilyn R. and
Zhao, Yilun and Zhao, Xuandong and
Cohan, Arman},
booktitle = {AI Measurement Science Workshop
at COLM 2026},
year = {2026},
url = {https://openreview.net/forum?id=BNx62Wx1ej}
}