The abstract of a new arXiv preprint points to a persistent problem in multilingual NLP: reported grammatical competence of large language models shifts depending on the measurement approach. The authors note that claims vary sharply with the evaluation paradigm, yet the interaction between evaluation paradigm, post-training, and language resource availability has not been fully explored.

The paper appears to target that gap, examining how these three factors jointly shape linguistic ability across languages. Because the abstract is truncated, the specific results and methodology are not available in the source, but the framing suggests that single-metric evaluations may be misleading.

The implication for researchers and practitioners is that cross-language comparisons should be interpreted with caution. The preprint underscores the need for evaluation designs that account for resource disparities and post-training choices.