Skip to content

Financial-statement review benchmark

Every model reads a full set of consolidated financial statements with deliberately planted inconsistencies, and has to find them. We score what each model actually catches, and penalise what it invents.

Leaderboard

No published results yet. The leaderboard appears once a run is published.

Methodology

The task. Each task pack is a real set of IFRS consolidated financial statements with a fixed number of deliberately injected inconsistencies: arithmetic that does not foot, figures that disagree between the statements and the notes, and totals that do not reconcile to their components. A model receives the full document and returns the inconsistencies it finds.

Scoring. A separate judge model matches each finding against the known list. A finding counts when it names the same inconsistency at the same location. The headline number is F1, the harmonic mean of recall (how many planted errors were found) and precision (how many of the model's findings were real). Findings that matched nothing are reported as hallucinations.

Contamination. Statements are anonymised and rescaled so a model cannot recognise a public filing or recall it from training, and runs execute without web access. The answer keys are never published.