AIBlindspot
← All case studies
DATDAT-002 — Data Privacy and Protection Failures

LLM Training Data Memorisation Enables Personal Data Extraction

4/5Sector: OtherGeography: GlobalStage: DevelopIngested: —

Executive Summary

Large language models can reproduce verbatim personal data, including names and contact details, when prompted with partial contextual strings from training corpora. Organisations deploying LLMs risk breaching data protection obligations and face regulatory liability if PII ingested during training is recoverable by users.

Domain

Data Management

Blindspots in data quality, privacy, bias, lineage, lifecycle, and third-party data dependencies.

Source

MIT AI Risk Repository — Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024) ↗

https://airisk.mit.edu/

Could this happen in your organisation?

A Velinor AI Audit maps your active AI portfolio against the 50+ blindspots and benchmarks against documented sector failures like this one. A board-ready foresight document in 5 weeks.