DATDAT-001 — Data Quality and Completeness Issues
Toxic Training Data Corrupts LLM Output Quality and Safety
5/5Sector: OtherGeography: GlobalStage: DevelopIngested: —
Executive Summary
Large language models trained on data containing hate speech, threats, and offensive language reproduce those harmful patterns in deployment. Organisations face reputational, legal, and regulatory exposure when such outputs reach customers or staff.
Domain
Data Management
Blindspots in data quality, privacy, bias, lineage, lifecycle, and third-party data dependencies.
Source
MIT AI Risk Repository — Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024) ↗https://airisk.mit.edu/
Could this happen in your organisation?
A Velinor AI Audit maps your active AI portfolio against the 50+ blindspots and benchmarks against documented sector failures like this one. A board-ready foresight document in 5 weeks.