The abstract introduces a new problem: picking the best model from a set of large language models when each has a different cost per query. The authors argue that typical evaluation methods treat all queries as equally expensive, which is unrealistic in practice.

To address this, they formulate the task as a variant of the multi-armed bandit problem, specifically with dueling feedback. In this setting, the algorithm learns by comparing two models at a time rather than observing absolute scores, and it must balance exploration and exploitation under a budget.

The work is motivated by real-world LLM APIs, where costs vary by model and provider. By making cost explicit in the bandit framework, the approach could help practitioners choose a model more economically. The abstract is truncated, so details on the algorithm and theoretical guarantees are not yet available in the source.