A new benchmark called TypedBench aims to evaluate a class of decision models that produce calibrated probabilities over typed answers—such as categorical choices, ordinal levels, or binary outcomes—rather than generating free-form text. These models use a non-generative interface, meaning they return structured probability estimates that software can act on directly through thresholds.
The benchmark focuses on three main dimensions: calibration, framing sensitivity, and cost. Calibration asks whether predicted probabilities match observed outcomes; framing sensitivity checks whether small changes in question wording alter results; and cost captures the practical expense of deploying these models. Together, these dimensions address whether a model can be trusted for decisions that depend on consistent, well-calibrated probability estimates.
Because the source is a single arXiv announcement, there are no competing findings to compare. The paper positions TypedBench as a tool for researchers and developers building systems that rely on System One models, particularly where threshold-based decisions make reliability and consistency essential. Further details on the benchmark's construction and results would require reading the full paper.