Tuesday, 22 September 2026

Search
Latent Digest

TECHNOLOGY, TRACKED ACROSS DISCIPLINES

Research Digest

Historical OCR: VLMs Face Cost and Hallucination Hurdles

Three arXiv papers examine whether vision-language models can replace traditional OCR for historical documents, with cost and hallucination emerging as key barriers.

· 1 min read · 5 sources

Three recent arXiv papers converge on the same question: how useful are vision-language models (VLMs) for OCR, especially on historical documents? All three acknowledge that VLMs have posted impressive or strong results, but they diverge on whether that translates into practical deployment.

One paper argues that large VLMs are too computationally expensive and too dependent on large-scale pretraining for historical text recognition, and instead proposes adapting the traditional PP-OCRv6 pipeline as a "free lunch." A second paper pushes back on that pessimism, presenting HunyuanOCR-1.5, a lightweight OCR-specialized VLM designed to be faster and better while unifying document parsing, text spotting, information extraction, and related tasks.

A third paper adds a different caveat: on historical Uruguayan documents, VLMs can achieve low character error rates yet still hallucinate. That means standard accuracy metrics may not capture the reliability needed for archival digitization. Together, the papers suggest that the choice between traditional OCR and VLMs is not settled, and that cost, speed, and hallucination all need to be weighed.

Sources · 5

  1. 01From Retrieval to Recognition:How Vision--Language Models Become OCR SpecialistsarXiv
  2. 02RAVE: Re-Allocating Visual Attention in Large Multimodal ModelsarXiv
  3. 03A Free Lunch? Adapting PP-OCRv6 for Historical Text RecognitionarXiv
  4. 04HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and BetterarXiv
  5. 05When Low CER is Not Enough: An Analysis of Hallucinations in Vision-Language OCR Systems on Historical Uruguayan DocumentsarXiv

More in Research Digest