AIMS WORKSHOP · COLM 2026

LegalScope

Measuring Exam-to-Case Transfer
in LLM Legal Reasoning

How much does success on a legal exam tell us
about reasoning through a real case?

1 UC Santa Barbara2 Yale University3 UC Berkeley

* Corresponding author

Downloadable now: 78 source records + benchmark metadata · Data availability and release notes

861Legal-exam questions
276Issue–stance prompts
28Model groups
56De-identified judgments

THE RESEARCH QUESTION

Abstract

Public legal exams are widely used to evaluate legal reasoning in large language models, yet it remains unclear whether exam performance reflects the ability to reason over real-case facts. We introduce LegalScope, a dual-track benchmark that pairs 861 legal-exam questions with 276 lawyer-reviewed issue-stance prompts derived from de-identified Chinese judgments. LegalScope evaluates models under both reference-answer scoring and case-based rubric scoring, enabling direct analysis of exam-to-case transfer.

Across 28 model groups, public-exam scores correlate with real-case scores but do not fully predict them; citation relevance receives lower mean scores than argument validity under both automatic and lawyer evaluation. We further show that automated evaluation aligns strongly with human review on exam answers but less reliably on real-case legal analysis, highlighting the need for expert-grounded case evaluation.

Explore the workshop paper

BENCHMARK

Benchmark overview

Two tracks, evaluated on the same 28 model groups.

PUBLIC EXAMS861

Questions with a reference answer.

Legal-exam questions from four jurisdictions, scored for substantive consistency with the reference answer on a 0–4 scale.

United States
603
China
94
United Kingdom
86
Australia
78
REAL CASES276

Arguments on both sides of an issue.

138 legal issues from 56 de-identified Chinese judgments. Each issue becomes a supporting and an opposing prompt, preserving a closed factual record.

SupportOppose

Seven categories: Tort, Contract, Criminal, Intellectual Property, Administrative, Civil Procedure and Property.

LegalScope pipeline: source collection and de-identification; two-track evaluation across 28 models; independent human validation; and reliability audits.
Figure 1. Benchmark construction and evaluation. Historical controls and fixed-answer audits retain their own evaluation pools.
View full-resolution figure ↗
A

Citation relevance

Does the selected authority respond to the issue?

B

Constraint extraction

Does the response follow the stance, facts and task constraints?

C

Argument validity

Does the conclusion follow from the law and the supplied facts?

The real-case rubric adapts the three evaluation dimensions of CourtReasoner (Han et al., EMNLP 2025) to closed-book Chinese case analysis.

Task examples

Read the input, inspect an answer, and see what the scoring protocol measures.

AUSTRALIA · VICTORIAN BAR · EXTVICBAR01A

Barrister confidentiality

Reference-answer scoring

Facts supplied in the exam

A barrister is representing a celebrity client who issued proceedings for breach of contract as a result of sexual harassment from their employer. During the course of the retainer, the client informs the barrister that the client is considering divorcing their spouse.

The question

(a) Could the barrister disclose the information about the potential divorce to other people? Why or why not? [1 mark]

Question 1(a), 13 October 2024 entrance exam, p. 3. © 2024 Victorian Bar Inc. Original publication · CC BY-NC-ND 4.0. Line breaks normalized; no endorsement implied.

What LegalScope evaluates

Whether a model reaches a compatible conclusion and explains the decisive legal point: here, the duty of confidentiality and its application to information obtained during the retainer. Evaluation compares substance with the reference; it does not require identical wording.

0.817

Exam–case association

Model-level Pearson correlation for automatic scores; Spearman correlation is 0.708. Rankings and variant gains do not transfer uniformly.

66.9 / 72.6

Citation / argument

Mean automatic scores on a 0–100 scale. Lawyers show the same direction descriptively; the pooled lawyer interval crosses zero.

0.910 / 0.312

Exam / case judge agreement

Answer-level Pearson correlations between automatic and human scores. Case scores use the equal-weight mean of two lawyers.

Model performance

28 model groups · Reported benchmark results

Download CSV
Best in columnSecond best in columnAll 28 model groups · ties share rank

Loading published results…

LegalScope model scores. Best and second-best displayed values are marked separately within each column over the complete roster.
Automatic evaluationHuman evaluation

All scores are on a 0–100 scale. Automatic scores cover the full tracks. Human scores cover 80 exam items and 10 case prompts; case scores pool two lawyers. Best and second-best marks use the displayed one-decimal values across all 28 groups; equal values share a mark, and filtering does not change their rank. The initial order follows the paper. CSV/Parquet additionally include the two descriptive overall means. Highlights are descriptive, not significance tests. Scores do not equate difficulty across tracks.

Explore the score distributions and transfer plots

RESOURCES

Paper, code and data

Research resources and current release information.

Data availability. Hugging Face provides 78 Victorian Bar source-excerpt records and selected candidate answers under CC BY-NC-ND 4.0, alongside metadata and aggregate results. The four dataset configurations load independently; original CSV and JSONL files are also available. Source excerpts preserve attribution and license notices and are not exact replacements for the paper’s evaluation prompts. The full 861-item exam track, the 276-prompt case dataset, full model responses and lawyer review sheets are not included. Case prompts and answers are being considered for item-level source and privacy review; this is not a blanket prohibition on publishing anonymized research material. These materials cannot rerun the complete evaluation. Source-specific release details · Loading examples, fields and answer types.

CITATION

Citation

AI Measurement Science Workshop
at COLM 2026

@inproceedings{wang2026legalscope,
  title = {{LegalScope}: Measuring Exam-to-Case
           Transfer in {LLM} Legal Reasoning},
  author = {Wang, Hongyu and Han, Rilyn R. and
            Zhao, Yilun and Zhao, Xuandong and
            Cohan, Arman},
  booktitle = {AI Measurement Science Workshop
               at COLM 2026},
  year = {2026},
  url = {https://openreview.net/forum?id=BNx62Wx1ej}
}