Long-horizon tool agents frequently make incremental progress even when they fail to complete a task, which has led to interest in partial-credit evaluation. But a new preprint argues that naive partial-credit schemes can be misleading. The paper, posted on arXiv, introduces PartHackBench, a benchmark designed as a certified equal-progress stress test for these evaluators.

The motivation is that evaluators may reward milestones that are temporary, later reversed, or not actually attributable to the agent's own actions. PartHackBench appears designed to expose such failures by providing controlled tests where progress can be compared fairly. The authors position this as a way to certify whether an evaluator gives credit only for real, lasting progress.

The abstract does not yet detail the benchmark's construction or results, so the full methodology remains to be examined. Still, the work points to a key challenge in agent evaluation: distinguishing meaningful progress from superficial or coincidental success. Further details are available in the preprint on arXiv.