The paper, posted on arXiv, examines a deployed agentic home-automation system that makes several structurally different kinds of large language model calls. These range from routing user intent and classifying actions to grounding language in a device registry, planning multi-agent pipelines, and writing the Python code that those pipelines execute. Rather than evaluating the system as a whole, the authors perform a per-call-site evaluation, testing small language models at each distinct call type.
The title states the paper's central claim: not every call needs a frontier model. By breaking down the pipeline into individual call sites, the study aims to identify where smaller, cheaper models might suffice without degrading performance. The abstract does not report specific results, but the framing suggests that a one-size-fits-all approach to model selection is wasteful for agentic systems.
This work is relevant to the growing field of LLM-based agents, where cost and latency are practical concerns. If small models can handle routine calls while frontier models are reserved for the most complex steps, deployed systems could become significantly more efficient. The paper's per-call-site methodology offers a template for such selective deployment.