OPSOPS-001 — Monitoring and Alerting Inadequacies
LLM Inconsistency Across Users, Sessions and Conversations
4/5Sector: OtherGeography: GlobalStage: OperateIngested: —
Executive Summary
Large language models produce materially different answers to identical queries depending on user, session, or conversational context. Operational decisions based on such outputs carry unquantified variance risk, undermining audit trails and regulatory defensibility.
Domain
Operational Management
Blindspots in monitoring, incident response, performance, scalability, integration, and business continuity.
Source
MIT AI Risk Repository — Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024) ↗https://airisk.mit.edu/
Could this happen in your organisation?
A Velinor AI Audit maps your active AI portfolio against the 50+ blindspots and benchmarks against documented sector failures like this one. A board-ready foresight document in 5 weeks.