According to an arXiv paper, current code benchmarks typically evaluate AI agents by whether they can produce a test that reproduces a known bug or a patch that fixes a described issue. The authors argue that neither task isolates the specific skill of property-based testing, where an agent must infer general properties and generate inputs to check them.

To fill that gap, the paper introduces PBT-Bench, a benchmark explicitly designed for property-based testing. The abstract does not include full details of the benchmark's tasks or evaluation methodology, but the stated goal is to separate this capability from other code-related skills.

As a single-source report, there are no conflicting findings to compare. The contribution is a proposed evaluation target rather than a set of experimental results.