AIBlindspot
← All case studies
DATDAT-002 — Data Privacy and Protection Failures

Generative AI Training Datasets Contain Personal and Identifiable Information

4/5Sector: OtherGeography: GlobalStage: DevelopIngested: —

Executive Summary

Generative AI developers routinely scrape web data containing personal information, and fine-tuning with proprietary datasets compounds PII exposure across the supply chain. Organisations deploying such models face regulatory liability under data protection law without adequate provenance controls.

Domain

Data Management

Blindspots in data quality, privacy, bias, lineage, lifecycle, and third-party data dependencies.

Source

MIT AI Risk Repository — Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024) ↗

https://airisk.mit.edu/

Could this happen in your organisation?

A Velinor AI Audit maps your active AI portfolio against the 50+ blindspots and benchmarks against documented sector failures like this one. A board-ready foresight document in 5 weeks.