> Source repository: harveyai/harvey-labs > Trending data: +87 stars today | Language: Python | Category: AI legal benchmark
The Blind Spot in AI Benchmarks
AI benchmarks are everywhere. MMLU measures knowledge breadth, SWE-bench measures coding ability, MATH measures mathematical reasoning. But one field has long lacked a solid benchmark: legal work.
The reason is simple. Legal work is not "answering questions" or "writing code" — it involves finding issues in piles of real documents, cross-referencing them, and producing structured deliverables according to rules. The complexity lies not in the knowledge itself, but in the process and the environment.
Harvey AI has just open-sourced the Legal Agent Benchmark (LAB), specifically designed to measure how well LLM agents perform real legal work.
What LAB Is
LAB consists of two parts:
1. A task dataset: agent instructions, relevant documents, and scoring rubrics 2. An execution framework: tools to run agents, evaluate results, and generate reports
The sample task is an M&A data room assignment, and its design is strikingly realistic:
- The agent receives a data room (a set of M&A-related documents)
- The agent must complete specific tasks per instructions (e.g., reviewing contract clauses, identifying risks, generating reports)
- The agent can use tools (retrieval, computation, file operations)
- Evaluation is scored against rubrics, not simple right/wrong answers
- Multi-model adapters: plug in different LLMs (GPT-4, Claude, Gemini, etc.)
- Tool calling: agents can retrieve documents, manipulate files, and compute
- Report generation: each run produces detailed reports, including pass/fail status for every rubric item
- Batch sweeps: run many tasks across multiple models at once and generate comparison reports
Why Legal Work Is Hard to Evaluate
Legal work differs from programming work in three fundamental ways, which is why SWE-bench-style methods don't apply:
First, inputs are documents, not code. Coding tasks come with function signatures and test cases — highly structured. Legal tasks come with dozens of pages of contracts, due diligence reports, and regulatory filings — poorly structured. The agent must understand the documents before it can even begin.
Second, outputs are not binary but a quality spectrum. Coding tasks have test suites: pass or fail. Legal outputs are reports or documents whose quality exists on a continuous spectrum. LAB quantifies this with rubrics — each rubric item is pass/fail, but overall quality is a weighted combination of multiple items.
Third, the process is not linear. Coding usually follows "read the problem → write code → run tests." Legal work may require repeatedly retrieving documents, cross-referencing, and revising earlier judgments. Agent trajectories are far more complex than in coding tasks.
LAB's Evaluation Method
LAB uses all-pass rubric scoring: a task has N rubric items, and the agent must pass all of them to count as complete. This is stricter than proportional scoring — a 90% pass rate in a legal context might mean 10% of clauses went unreviewed, which is unacceptable.
LAB also uses an LLM as judge to evaluate agent outputs. This raises a question: how reliable is the judge itself? LAB's documentation specifically discusses LLM judge behavior analysis, including consistency and bias across different rubric items.
The Execution Framework
LAB's harness supports:
Implications for Legal AI Applications
The open-sourcing of LAB sends several important signals:
First, legal AI has moved from "can it work" to "how well does it work." In 2023, people asked whether AI could do legal work at all; in 2025, the question is how good it is. LAB gives "how good" a quantifiable answer.
Second, agent capability matters more than model capability. LAB measures not a single LLM's legal knowledge, but how LLM agents perform within legal workflows. Gaps in agent architecture (tool calling, document retrieval, multi-step reasoning) may outweigh gaps in the underlying models themselves.
Third, realistic environments are key. LAB's tasks are based on real legal scenarios (an M&A data room), not artificial law exam questions. Benchmark results therefore track actual application performance more closely than academic test scores.
A Question Worth Considering
LAB's rubrics were designed by Harvey's legal experts. That means benchmark quality depends on the designers' domain expertise.
If another team wants to extend LAB — adding litigation or compliance tasks, for example — they need matching legal experts. That differs from SWE-bench, where test cases come from real GitHub PRs and any programmer can contribute.
The barrier to entry for legal benchmarks is inherently higher. Will that limit community contribution velocity? Or does the legal domain's uniqueness precisely require such a high bar to guarantee quality?
Harvey's choice: open-source the framework and task set, but retain expert review — a design that balances openness with quality.
---
*Project: github.com/harveyai/harvey-labs* *Announcement: harvey.ai/blog/introducing-harveys-legal-agent-benchmark*