How well OCR and vision-language models transcribe Icelandic documents — a broken-out scorecard, not a single grade. Ranked on micro-averaged character error rate, read as confidence intervals, not ranks.
The test material is a ground truth sample drawn from Sigurdur/icelandic-ocr-benchmark, pages of Icelandic-language documents. The scoring core is ported, unchanged, from finebooks/bhl-ocr-eval — same diplomatic/reading CER lanes, same bootstrap CIs, same fail-closed eligibility rule.
| Model ▾ | Type i | Params i ▾ | CER · reading i ▾ | CER · dip i ▾ | Content CER i ▾ | Sparse CER i ▾ | Recall i ▾ | Over-x i ▾ | Loop % i ▾ |
|---|
Sparse and blank pages are where models diverge most: strong content OCR can coexist with runaway hallucination on near-empty pages, and the aggregate CER shows it.
Reading the board. Where 95% bootstrap CIs overlap, differences are within noise — that overlap is the finding, not a defect. CIs resample volumes (source documents), so with few volumes they are wide by design.
icelandic-ocr-eval,
ported from finebooks/bhl-ocr-eval.
Ground truth: Sigurdur/icelandic-ocr-benchmark.