Historical-text OCR · IMPACT-BHL ground truth

BHL OCR Leaderboard

How well OCR and vision-language models transcribe 18th–19th-century book pages — a broken-out scorecard, not a single grade. Ranked on micro-averaged character error rate, read as confidence intervals, not ranks.

The test material is a ground truth made by hand long before today’s models: 2,165 pages from six natural-history volumes, expert-transcribed for the EU’s IMPACT digitisation programme (with BHL-Europe) in English, German, French and Latin. This board reuses that older ground truth to ask a current question — which of the new models actually reads old books best — and breaks the answer out, because models fail differently: some misread letters, some invent text on near-empty pages, some cannot stop repeating themselves.

clara intra Philoſophiæ Thalamos admiſsâ luce clara intra philosophiæ thalamos admissâ luce one line of the ground truth as each lane scores it — diplomatic keeps the long ſ, reading folds it · Piscium querelae, 1708
Type
Size
Reads
click a row for per-language CER
Model Type i Params i CER · reading i CER · dip i Content CER i Sparse CER i Recall i Over-x i Loop % i Reads i

Sparse and blank pages are where models diverge most: strong content OCR can coexist with runaway hallucination on near-empty pages, and the aggregate CER shows it.

Size vs. accuracy

params (log) vs CER · lower-left is better

Reading the board. Where 95% bootstrap CIs overlap, differences are within noise — that overlap is the finding, not a defect. CIs resample volumes (six books), so they are wide by design.