DATDAT-002 — Data Privacy and Protection Failures
LLMs Memorise and Leak Personally Identifiable Information from Training Data
4/5Sector: OtherGeography: GlobalStage: OperateIngested: —
Executive Summary
Large language models can memorise and reproduce personal data including names, addresses and telephone numbers, either inadvertently or through deliberate adversarial prompting. Organisations deploying such models face regulatory exposure under data protection law and reputational risk if PII surfaces in generated outputs.
Domain
Data Management
Blindspots in data quality, privacy, bias, lineage, lifecycle, and third-party data dependencies.
Source
MIT AI Risk Repository — The Ethics of Advanced AI Assistants (Gabriel2024) ↗https://airisk.mit.edu/
Could this happen in your organisation?
A Velinor AI Audit maps your active AI portfolio against the 50+ blindspots and benchmarks against documented sector failures like this one. A board-ready foresight document in 5 weeks.