A paper posted on arXiv challenges the standard way of measuring generalization in large language models. It argues that generalization is not about achieving high accuracy on a fixed test set, but about producing consistent, semantically stable outputs when the same input is expressed in different ways.

The paper notes that existing work typically evaluates generalization through accuracy on held-out data. That approach can miss cases where a model answers correctly for one phrasing but incorrectly for a near-synonymous one. The authors propose a multi-axis evaluation that treats stability across input variations as the primary signal.

This reframing has practical implications: if a model is stable across paraphrases, it is more likely to be relying on robust reasoning rather than surface patterns. The abstract, as released, does not include experimental details, so the evidence base for this proposal is not yet clear from the announcement alone.