AIBlindspot
← All case studies
DATDAT-002 — Data Privacy and Protection Failures

LLMs Memorise and Leak Personally Identifiable Information from Training Data

4/5Sector: OtherGeography: GlobalStage: OperateIngested: —

Executive Summary

Large language models can memorise and reproduce personal data including names, addresses and telephone numbers, either inadvertently or through deliberate adversarial prompting. Organisations deploying such models face regulatory exposure under data protection law and reputational risk if PII surfaces in generated outputs.

Domain

Data Management

Blindspots in data quality, privacy, bias, lineage, lifecycle, and third-party data dependencies.

Source

MIT AI Risk Repository — The Ethics of Advanced AI Assistants (Gabriel2024) ↗

https://airisk.mit.edu/

Could this happen in your organisation?

A Velinor AI Audit maps your active AI portfolio against the 50+ blindspots and benchmarks against documented sector failures like this one. A board-ready foresight document in 5 weeks.