Model benchmarks

Which model.
For which task.

Last updated: Jul 2026

Frontier models publish benchmark scores. Most benchmarks were not designed around professional document work.

From this analysis
Claude Opus 4.7 leads FinanceBench at 50.4%.
GPT-5 mini leads Financial QA 10K at 79.7%.
Claude models hold the top four positions on CUAD clause extraction.
Claude Fable 5 leads LegalBench contracts at 92.0%.
Domain
Finance

Performance on financial document tasks

These datasets measure frontier models on key financial workflows.

FinanceBench
Finance
Where this matters:Earnings analysis, investor research, SEC filing review

Tests AI accuracy on questions drawn from real SEC filings. Covers numerical reasoning and document comprehension across income statements, balance sheets, and footnotes.

No model exceeds 53.5%. A 6-point spread separates first from last.
Composite score · 9 models · 5 metrics
Claude Opus 4.7
50.4%
Claude Opus 4.8
50.0%
Claude Fable 5
49.0%
Groq Llama 3.3 70B
49.0%
Groq Llama 4 Scout
48.5%
GPT-5
47.8%
Claude Haiku 4.5
47.2%
Claude Sonnet 4.6
45.6%
GPT-5 mini
44.8%
Financial QA 10K
Finance
Where this matters:10-K prose review, credit analysis, research automation

Tests AI performance on question-answer pairs from 10-K annual reports. Evaluates whether models can accurately locate and interpret financial disclosures.

GPT-5 mini leads. On prose extraction, the cheaper model wins.
Composite score · 9 models · 5 metrics
GPT-5 mini
79.7%
GPT-5
78.8%
Claude Opus 4.8
75.5%
Claude Fable 5
75.1%
Groq Llama 4 Scout
74.3%
Claude Opus 4.7
73.0%
Groq Llama 3.3 70B
71.6%
Claude Haiku 4.5
66.0%
Claude Sonnet 4.6
65.7%
TAT-QA
Finance
Where this matters:Financial modelling support, table extraction, multi-step calculation

Tests AI performance on questions requiring both table and text comprehension from real financial reports. Covers multi-step numerical reasoning and structured data extraction.

Opus 4.7 leads by 0.9 points. Model choice matters more here than on prose tasks.
Composite score · 9 models · 5 metrics
Claude Opus 4.7
63.6%
Claude Opus 4.8
62.7%
Claude Fable 5
62.3%
Groq Llama 3.3 70B
60.2%
GPT-5 mini
60.1%
Groq Llama 4 Scout
57.7%
GPT-5
56.3%
Claude Sonnet 4.6
54.3%
Claude Haiku 4.5
53.4%
Legal

Performance on contract and legal document tasks

These datasets measure frontier models on contracting and legal tasks.

CUAD
Legal
Where this matters:Contract review, clause extraction, M&A due diligence

Tests AI accuracy on contract clause extraction across 150 contract samples. Covers identification and location of key legal provisions relevant to M&A and commercial review.

Claude models hold the top four positions. An 11-point spread separates first from last.
Composite score · 9 models · 3 metrics
Claude Opus 4.7
66.4%
Claude Opus 4.8
65.5%
Claude Fable 5
62.5%
Claude Haiku 4.5
61.4%
Claude Sonnet 4.6
60.3%
GPT-5
58.7%
GPT-5 mini
58.0%
Groq Llama 4 Scout
55.9%
Groq Llama 3.3 70B
55.1%
LegalBench (contracts)
Legal
Where this matters:Contract obligation review, compliance flagging, NDA analysis

Tests AI on contract obligation classification. Evaluates whether models can correctly identify and categorise contractual terms relevant to compliance and NDA review.

Leader at 92.0%. Narrow spread: mostly easy Yes/No classification.
Composite score · 9 models · 3 metrics
Claude Fable 5
92.0%
Groq Llama 3.3 70B
89.3%
Claude Opus 4.7
89.3%
Claude Opus 4.8
88.7%
Claude Haiku 4.5
85.3%
Groq Llama 4 Scout
84.7%
Claude Sonnet 4.6
83.3%
GPT-5
82.7%
GPT-5 mini
79.3%
ContractNLI
Legal
Where this matters:NDA review, contract obligation identification, compliance automation

Tests AI on natural language inference across NDA text. Evaluates whether models can determine whether a contractual hypothesis is supported, contradicted, or unaddressed by a given clause.

A 11-point spread across nine models. Model selection has limited impact at this difficulty level.
Composite score · 9 models · 3 metrics
Claude Fable 5
86.7%
Claude Opus 4.7
83.3%
GPT-5
82.7%
Claude Opus 4.8
82.7%
Claude Sonnet 4.6
82.0%
Groq Llama 4 Scout
82.0%
GPT-5 mini
80.7%
Claude Haiku 4.5
80.0%
Groq Llama 3.3 70B
76.0%
Methodology

All evaluations run independently by the Lab. 150 rows per dataset, identical inputs across all models. Composite scores use a per-dataset metric mask: only metrics appropriate to the task type are included. Finance datasets use five metrics; legal classification and CUAD extraction use three. Numbers reflect the latest evaluation run.