Software-engineering agents are changing who can write code, but they are also changing how much code there is to check. A new paper on arXiv, "A Case Study in Assuring AI-Written Software," highlights a two-sided problem: these agents let non-programmers build systems they could not create on their own, and they produce so much code that even experts cannot meaningfully inspect all of it. Both sides, the abstract suggests, complicate any attempt to assure that the resulting software is trustworthy.

The study focuses on assurance as a separate concern from ordinary code review. Rather than assuming that traditional inspection methods will scale, it examines how to evaluate AI-written software in a setting where human oversight is necessarily limited. The case-study format offers a concrete way to probe that tension, without claiming a one-size-fits-all solution.

For anyone relying on AI-generated code, the lesson is simple: trust cannot be inferred from the fact that an agent wrote working-looking software. The paper points to the need for deliberate assurance practices built around the unique characteristics of AI-written systems.