AIBlindspot
← All case studies
HUMHUM-004 — Trust and Acceptance Issues

LLM Capability Overstatement and Inconsistent Reliability Mislead Users

3/5Sector: TechnologyGeography: GlobalStage: OperateIngested: —

Executive Summary

Large language models exhibit unpredictable performance across domains due to benchmark contamination, prompt sensitivity, and developer exaggeration of capabilities. Organisations relying on these systems risk material decisions being made on unreliable outputs, exposing them to reputational and liability consequences.

Domain

Human Factors

Blindspots in change management, skills, human-AI collaboration, trust, workforce, and culture.

Source

MIT AI Risk Repository — Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024) ↗

https://airisk.mit.edu/

Could this happen in your organisation?

A Velinor AI Audit maps your active AI portfolio against the 50+ blindspots and benchmarks against documented sector failures like this one. A board-ready foresight document in 5 weeks.