Execution-based task verifiers play a crucial role in determining whether an agent has successfully completed a task in benchmark environments like AppWorld and WorkArena. A new preprint describes an audit of these shipped verifiers using source-informed mutation tests. The authors deliberately avoid modifying the checkers themselves, instead probing them with mutated versions of tasks to identify blind spots.

According to the abstract, the audit targets both AppWorld and WorkArena. The description of a specific mutation in AppWorld is cut off in the available text, but the method is clear: by comparing the verifier's decisions on original and mutated tasks, the researchers aim to reveal where the verifiers can be fooled or fail to generalise.

Because the source is a single preprint with a truncated abstract, this article can only report the stated approach and scope. The full findings are not yet available in the excerpt.