DATDAT-002 — Data Privacy and Protection Failures
LLM Training Data Memorisation Leaks Personal Information
4/5Sector: OtherGeography: GlobalStage: OperateIngested: —
Executive Summary
Large language models memorise and reproduce personal data ingested during training, including information individuals never consented to share. Organisations deploying such models face regulatory exposure under data protection law and reputational liability for third-party privacy breaches.
Domain
Data Management
Blindspots in data quality, privacy, bias, lineage, lifecycle, and third-party data dependencies.
Source
MIT AI Risk Repository — Ethical and social risks of harm from language models (Weidinger2021) ↗https://airisk.mit.edu/
Could this happen in your organisation?
A Velinor AI Audit maps your active AI portfolio against the 50+ blindspots and benchmarks against documented sector failures like this one. A board-ready foresight document in 5 weeks.