GOVGOV-001 — Accountability Framework Gaps
LLM Evaluators Producing Biased or Incorrect Assessments of Other AI Models
4/5Sector: OtherGeography: GlobalStage: DevelopIngested: —
Executive Summary
General-purpose AI models used to evaluate other AI systems generate flawed ratings, favouring verbose or politically skewed outputs. When embedded in training pipelines, these errors compound, producing models optimised to exploit evaluator weaknesses rather than perform correctly.
Domain
Governance & Compliance
Blindspots in accountability, regulatory compliance, ethics, risk management, data governance, and audit.
Source
MIT AI Risk Repository — Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) ↗https://airisk.mit.edu/
Could this happen in your organisation?
A Velinor AI Audit maps your active AI portfolio against the 50+ blindspots and benchmarks against documented sector failures like this one. A board-ready foresight document in 5 weeks.