DATDAT-002 — Data Privacy and Protection Failures
Pre-Trained Language Models Memorise and Expose Personal Data
4/5Sector: OtherGeography: GlobalStage: DevelopIngested: —
Executive Summary
Large language models trained on internet corpora retain and can reproduce personal data including phone numbers, email addresses, and home addresses. Organisations deploying such models face regulatory liability and reputational harm if memorised personal data is surfaced through user queries.
Domain
Data Management
Blindspots in data quality, privacy, bias, lineage, lifecycle, and third-party data dependencies.
Source
MIT AI Risk Repository — Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements (Deng2023) ↗https://airisk.mit.edu/
Could this happen in your organisation?
A Velinor AI Audit maps your active AI portfolio against the 50+ blindspots and benchmarks against documented sector failures like this one. A board-ready foresight document in 5 weeks.