A new paper on arXiv introduces InvestigationWorlds, an agentic environment for legal investigation. The authors propose using a rarely exploited artifact of U.S. civil litigation—the summary judgment motion—as the foundation for constructing realistic agent tasks.

Summary judgment motions rely on a record composed of real evidence, which gives the environment a grounded, factual basis. This stands in contrast to many AI benchmarks that depend on simplified or artificially generated scenarios. The environment appears intended to test how well AI agents can navigate legal evidence and perform investigation-style reasoning.

The abstract is brief, so details on task formats and evaluation metrics are not yet available. Still, the idea of grounding agentic AI in actual legal procedure could offer a useful stress test for systems meant to assist with legal work.