Frontier language models rarely enter clinical workflows because realistic, longitudinal benchmarks are scarce. Real electronic health record (EHR) data cannot be openly shared due to privacy constraints, leaving developers without the large-scale, realistic testbeds needed to train and evaluate models safely.

To address this, the authors introduce a synthetic hospital benchmark that is open, verifiable, and physician-validated. By generating synthetic longitudinal EHR data, the benchmark aims to mimic real patient trajectories while avoiding privacy restrictions, and physician validation helps ensure clinical plausibility.

The approach could lower a key barrier to clinical AI adoption: providing a shared, realistic evaluation environment that researchers can use without exposing sensitive patient information. The abstract does not detail the generation method or validation results, but the stated goals point toward a practical path for benchmarking models on longitudinal records.