Automated program repair (APR) systems are typically judged by one blunt metric: whether the generated patch passes the test suite. A new arXiv paper argues that this leaderboard-style evaluation misses important differences in patch quality, such as maintainability, security, and the computational resources required to produce the fix.
The authors propose a "Weighted…" evaluation framework, according to the abstract, though the full details of the weighting scheme are not included in the available text. The paper's central claim is that a patch that passes tests can still be poor in practice if it is insecure, hard to maintain, or expensive to generate.
Because only the abstract is available, the specific weights, datasets, and model comparisons are not yet known. The paper does, however, signal a growing concern in the APR community that test-suite success alone is an insufficient proxy for real-world usefulness.