A joint research effort between Germany and the UK has found that anonymized patient records included in medical datasets years ago can still affect AI systems trained on them. Even though names and direct identifiers were removed, the underlying physiological and clinical patterns in the data remain distinct enough that a model can later associate a current patient with a specific historical record.

The consequence, according to the researchers, is that an AI trained on such old data may treat a 2026 patient as if they were the person from decades earlier, leading to incorrect diagnoses or treatment recommendations. The effect is not hypothetical — the study demonstrates that these long-dormant traits can surface during inference, overriding more recent and relevant clinical signals.

Because this is a single source, there is no independent corroboration or dissent to report. The authors suggest that medical AI developers should audit the age and provenance of their training datasets, and likely exclude or re-weight very old anonymized records to avoid this hidden bias.