Tuesday, 22 September 2026

Search
Latent Digest

TECHNOLOGY, TRACKED ACROSS DISCIPLINES

Research Digest

Three New Papers Rethink AI Privacy Beyond Training-Time Defenses

Recent research shifts privacy focus to retrieval deletion, semantic sanitization metrics, and the protective effect of missing data.

· 1 min read · 3 sources

Three new arXiv papers approach AI privacy from different angles, and together they suggest the field is moving beyond the familiar focus on private training. One paper examines retrieval-augmented systems, warning that vector indexes may keep deleted items in their search graph. The authors note that existing deletion interfaces can stop deleted identifiers from appearing in returned results, but the underlying retention remains a concern.

A second paper tackles evaluation of sanitization. It argues that AI-based removal of concepts from images and text is often assessed in an ad hoc way, and proposes 'contrastive privacy' as a formal definition with a quantitative test and semantic interpretation. A third paper looks at missing data, arguing that absence of information can itself provide privacy amplification in sensitive fields like medicine and finance.

The papers differ in scope: one targets system-level deletion, another measurement of sanitized outputs, and the third a structural property of incomplete datasets. They agree, however, that privacy is not a single training-time property, and that the full lifecycle of AI systems—indexes, outputs, and data gaps—deserves scrutiny.

Sources · 3

  1. 01Beyond Private Training: The New Landscape of AI PrivacyarXiv
  2. 02Contrastive Privacy: A Semantic Approach to Measuring Privacy of AI-based SanitizationarXiv
  3. 03On the Inherent Privacy Amplification of Missing DataarXiv

More in Research Digest