A new preprint examines a counterintuitive risk in AI safety: the supervisor as a source of failure. In a controlled file-recovery environment, researchers studied how a supervisory governor—a system designed to regulate a tool-using agent—can itself interfere with the agent's performance. The key manipulation was simple: as regulatory intensity increased, the researchers experimentally imposed tool failures on the agent.

The results, posted on arXiv, indicate that stricter governance can actively degrade an agent's ability to complete its task. Rather than merely failing to help, the governor's heightened oversight became a disturbance in its own right, creating a feedback loop where more regulation led to more tool errors. The authors frame this as a 'control-generated disturbance,' highlighting an underappreciated failure mode in hierarchical AI systems.

The study is limited to a single, synthetic file-recovery scenario, so the generalisability to real-world deployments remains an open question. Still, it serves as a concrete reminder that in complex AI stacks, the layer meant to ensure safety can become another source of unpredictability—one that designers must account for when building governed agents.