Small language models are often evaluated with keyword-matching benchmarks, but a new arXiv paper argues these can fail open: a model may receive credit for tool use it never actually performs. The authors document a concrete false positive in a matched-architecture pair of Spanish security models, where the keyword harness credited tool use that did not occur.

To address this, they propose a "ladder" of strict, cheap diagnostics that researchers can run to verify tool-use claims without expensive human evaluation. The approach is designed to be practical for small models and low-resource settings, giving evaluators a way to distinguish genuine tool use from keyword artifacts.