Choosing which large language model to deploy is often a one-time decision based on offline benchmarks. But when models process streaming data, those benchmarks are only rough proxies for real performance. A new arXiv paper directly addresses this gap, framing the problem as one of online active model selection rather than a static choice.
The work appears to treat model selection as a sequential decision task, where the system must pick among candidate LLMs as data arrives. The abstract mentions an oracle, implying the authors compare their method against an ideal selector that knows the best model at each step. This suggests the goal is to minimize the gap between the chosen model and the oracle's choice over time.
Because the abstract is truncated, details of the proposed algorithm and experimental results are not available in the source. Still, the core motivation is clear: relying on benchmarks alone is insufficient for streaming applications, and active, online selection may offer a more realistic path to good performance.