AIBlindspot
← All case studies
OPSOPS-001 — Monitoring and Alerting Inadequacies

LLM Inconsistency Across Users, Sessions and Conversations

4/5Sector: OtherGeography: GlobalStage: OperateIngested: —

Executive Summary

Large language models produce materially different answers to identical queries depending on user, session, or conversational context. Operational decisions based on such outputs carry unquantified variance risk, undermining audit trails and regulatory defensibility.

Domain

Operational Management

Blindspots in monitoring, incident response, performance, scalability, integration, and business continuity.

Source

MIT AI Risk Repository — Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024) ↗

https://airisk.mit.edu/

Could this happen in your organisation?

A Velinor AI Audit maps your active AI portfolio against the 50+ blindspots and benchmarks against documented sector failures like this one. A board-ready foresight document in 5 weeks.