GOVGOV-001 — Accountability Framework Gaps
AI Benchmarks Saturate and Fail to Detect Capability Advances
3/5Sector: OtherGeography: GlobalStage: DevelopIngested: —
Executive Summary
AI evaluation benchmarks are reaching performance ceilings, rendering them unable to detect meaningful capability improvements in new models. Regulators and procurement bodies relying on saturated benchmarks risk systematically underestimating the power of deployed general-purpose AI systems.
Domain
Governance & Compliance
Blindspots in accountability, regulatory compliance, ethics, risk management, data governance, and audit.
Source
MIT AI Risk Repository — Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) ↗https://airisk.mit.edu/
Could this happen in your organisation?
A Velinor AI Audit maps your active AI portfolio against the 50+ blindspots and benchmarks against documented sector failures like this one. A board-ready foresight document in 5 weeks.