Most LLM benchmarks test isolated tasks, but real deployments involve long conversations. This paper argues that agents can appear competent in single turns yet drift into inconsistency over time. To study this, the authors set up a controlled 20-step multi-agent environment and track when and why consistency breaks down.
Rather than averaging errors, they borrow survival analysis from biostatistics to estimate the "survival" of consistency over interaction steps. They also introduce a failure-rationale taxonomy to classify the reasons behind breakdowns. The abstract does not report specific results, so the paper's findings will need to be read in full.