Time-series foundation models are often judged on forecast accuracy, but the authors of a new arXiv preprint argue that accuracy alone does not determine an agent's value for operational decisions. To test that claim, they propose Forecast Workflow Bench, a benchmark for evaluating language-model decisions when forecast tools are budgeted.

The benchmark centers on two measures: decision quality and forecast cost. Evaluating an agent requires looking at how well its decisions perform in the operational setting and how many forecast queries or tools it consumes along the way.

The authors emphasize that accuracy alone is insufficient and that both dimensions need to be considered together. The paper is one source, so no independent comparison is offered here.