Chaos engineering, the practice of deliberately breaking systems to expose weaknesses, is getting an AI upgrade. Gremlin’s new Foresight AI add-on automates the labor-intensive parts of the process: setting up failure experiments, interpreting the resulting telemetry, and recommending fixes. The company says this turns what used to be manual prep and post-testing work into an agentic loop that runs until the induced failure stops occurring.

To keep the AI grounded, Gremlin pairs large language models with its Failure Atlas, a repository of millions of experiments run over the past decade on tens of thousands of systems. The company says this reduces hallucinations by anchoring the model’s answers in real-world failure data rather than general internet knowledge. Gremlin uses a mix of open and closed LLMs, depending on which performs best at the time.

Foresight AI works within the enterprise’s existing access controls to limit the blast radius of any test. Once a failure is induced, the AI identifies the root cause, proposes a fix, and either applies it or hands a report to a site reliability engineer. Gremlin co-founder Kolton Andrus notes that while many AIOps tools only diagnose problems after a crash, this approach deliberately breaks the system first to prevent future incidents.

Analyst Jason English of Intellyx cautions that any service promising to break systems for the greater good needs executive sign-off, and that effective use requires a cultural shift in how organizations view risk. As AI accelerates both application development and infrastructure deployment, Gremlin argues that automated chaos testing can keep pace with changes that human reviewers cannot manually verify in time.