A new paper on arXiv examines how well large language models handle numerical claim verification. The authors report that LLMs are widely used for this task, yet remain brittle when the numbers involved are altered. Even small changes in a value can sharply degrade accuracy, according to the abstract.

The finding holds for frontier LLMs, meaning the most advanced models available also exhibit this weakness. The paper is titled "Towards Robust Numerical Claim Verification" and is listed as arXiv:2610.00689. The abstract does not detail proposed solutions, but the stated goal is to move toward more robust verification of numerical claims.