Last updated: Jul 2026
Frontier models publish benchmark scores. Most benchmarks were not designed around professional document work.
These datasets measure frontier models on key financial workflows.
Tests AI accuracy on questions drawn from real SEC filings. Covers numerical reasoning and document comprehension across income statements, balance sheets, and footnotes.
Tests AI performance on question-answer pairs from 10-K annual reports. Evaluates whether models can accurately locate and interpret financial disclosures.
Tests AI performance on questions requiring both table and text comprehension from real financial reports. Covers multi-step numerical reasoning and structured data extraction.
These datasets measure frontier models on contracting and legal tasks.
Tests AI accuracy on contract clause extraction across 150 contract samples. Covers identification and location of key legal provisions relevant to M&A and commercial review.
Tests AI on contract obligation classification. Evaluates whether models can correctly identify and categorise contractual terms relevant to compliance and NDA review.
Tests AI on natural language inference across NDA text. Evaluates whether models can determine whether a contractual hypothesis is supported, contradicted, or unaddressed by a given clause.
All evaluations run independently by the Lab. 150 rows per dataset, identical inputs across all models. Composite scores use a per-dataset metric mask: only metrics appropriate to the task type are included. Finance datasets use five metrics; legal classification and CUAD extraction use three. Numbers reflect the latest evaluation run.