Reposted by @antoniakrm.bsky.social
I will personnaly stay far away from VLM for the time being, for anything remotely philological, following our experiments: arxiv.org/abs/2605.27750
I can do with OCR error, but plausible text is too dangerous for me.
(we got a great OCR model btw if you need it: htrmopo.inria.fr/models/10.52... )
arxiv.org
Reading or Guessing? Visual Grounding Failures of Vision-Language Models for OCR in Ancient Greek Editions
Recent work has shown that Vision-Language Models (VLMs) used for optical character recognition (OCR) can generate plausible but visually unsupported text, suggesting reliance on language priors. Comp...