How well OCR and vision-language models transcribe 18th–19th-century book pages — a broken-out scorecard, not a single grade. Ranked on micro-averaged character error rate, read as confidence intervals, not ranks.
The test material is a ground truth made by hand long before today’s models: 2,165 pages from six natural-history volumes, expert-transcribed for the EU’s IMPACT digitisation programme (with BHL-Europe) in English, German, French and Latin. This board reuses that older ground truth to ask a current question — which of the new models actually reads old books best — and breaks the answer out, because models fail differently: some misread letters, some invent text on near-empty pages, some cannot stop repeating themselves.
| Model ▾ | Type i | Params i ▾ | CER · reading i ▾ | CER · dip i ▾ | Content CER i ▾ | Sparse CER i ▾ | Recall i ▾ | Over-x i ▾ | Loop % i ▾ | Reads i |
|---|
Sparse and blank pages are where models diverge most: strong content OCR can coexist with runaway hallucination on near-empty pages, and the aggregate CER shows it.
Reading the board. Where 95% bootstrap CIs overlap, differences are within noise — that overlap is the finding, not a defect. CIs resample volumes (six books), so they are wide by design.