Direct-decision models are increasingly used to turn text into structured labels and scores, offering low-latency solutions for classification and automatic evaluation. But a new arXiv preprint argues that their reliability cannot be measured by accuracy alone. The authors point to a specific failure mode: models may not correctly use the ordinal nature of the scale they are asked to output.
The paper, titled "More Choices, Fewer Decisions: Ordinal-Scale Bias in JEV-like Direct-Decision Models," examines how these models handle ordinal scales—such as rating systems where order matters, not just category membership. The abstract suggests that a model must do more than predict the right label; it must also respect the ordering implied by the scale. When a model ignores that ordering, it introduces a bias that accuracy metrics fail to capture.
While the full findings are not yet available in the abstract, the central claim is clear: for direct-decision models, accuracy is a necessary but not sufficient condition for reliability. The study specifically focuses on JEV-like models, hinting that this bias may be systemic within that family of architectures. The work adds to a growing conversation about evaluation metrics that go beyond simple correctness in structured-output tasks.