Autonomous security agents are becoming increasingly capable at finding bugs, but evaluating their performance remains a challenge. When you point one at a realistic target, the output is essentially a report the agent wrote about itself: confident prose, a list of findings, and no clear way to tell which of those findings are genuine. That makes it hard for teams to trust or compare these tools.
XRanges is a new scoring system designed to fill that gap. Instead of relying on an agent's self-assessment, it provides a benchmark that scores how well the agent actually performs against realistic targets. This gives security teams a more objective way to measure and compare autonomous agents before deploying them.
According to The Hacker News, the benchmark was tested by 545 hackers before its release. That community testing phase suggests the scoring methodology has been stress-tested by a wide range of security practitioners, adding credibility to the results it produces.