LLM arenas rank models by pitting answers against each other and letting humans vote, but those votes may be swayed by style as much as substance. A new paper takes a stylometric approach to 137,293 decisive French-language votes from the Compar:IA arena, looking at how formatting, length, and lexical diversity affect outcomes.

The authors argue that preferences in such arenas reflect not only what a model says but how it presents the answer. Their analysis focuses on measurable textual traits, suggesting that even strong content can lose if it is poorly formatted or too terse.

Because the abstract is truncated, the full methodology and conclusions are not yet available, but the core finding is clear: presentation matters in human evaluation of LLM outputs. The study adds to a growing body of evidence that arena rankings are not purely about factual quality.