Urban diagnosis requires pulling together heterogeneous observations to spot problems, pinpoint affected areas, and trace their causes. That process is central to evidence-based planning, but it is complex and data-heavy. A new paper on arXiv introduces DUDA-Bench, a benchmark designed to test whether LLM agents can handle this kind of multimodal reasoning.

The benchmark frames urban diagnosis as a structured task: agents receive diverse data sources and must produce a diagnosis that covers the problem, its location, and its contributing factors. By standardizing this evaluation, the authors aim to measure how well current LLM agents integrate different modalities and reason about real-world urban conditions.

Because the abstract is the only available source, details on the benchmark's construction, datasets, and results are not yet described here. The significance lies in the attempt to create a shared evaluation target for AI-assisted urban analysis, potentially pushing toward more reliable tools for planners and city managers.