HUMHUM-004 — Trust and Acceptance Issues
LLM Capability Overstatement and Inconsistent Reliability Mislead Users
3/5Sector: TechnologyGeography: GlobalStage: OperateIngested: —
Executive Summary
Large language models exhibit unpredictable performance across domains due to benchmark contamination, prompt sensitivity, and developer exaggeration of capabilities. Organisations relying on these systems risk material decisions being made on unreliable outputs, exposing them to reputational and liability consequences.
Domain
Human Factors
Blindspots in change management, skills, human-AI collaboration, trust, workforce, and culture.
Source
MIT AI Risk Repository — Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024) ↗https://airisk.mit.edu/
Could this happen in your organisation?
A Velinor AI Audit maps your active AI portfolio against the 50+ blindspots and benchmarks against documented sector failures like this one. A board-ready foresight document in 5 weeks.