AIBlindspot
← All case studies
OPSOPS-001 — Monitoring and Alerting Inadequacies

LLM Performance Shifts from Minor Prompt Formatting Changes

4/5Sector: OtherGeography: GlobalStage: OperateIngested: —

Executive Summary

Large language models produce significantly different outputs when prompt formatting varies in spacing, casing, or separators, undermining the reliability of performance benchmarks. Organisations cannot trust evaluation results or vendor comparisons without controlling for formatting variables across all tests.

Domain

Operational Management

Blindspots in monitoring, incident response, performance, scalability, integration, and business continuity.

Source

MIT AI Risk Repository — Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) ↗

https://airisk.mit.edu/

Could this happen in your organisation?

A Velinor AI Audit maps your active AI portfolio against the 50+ blindspots and benchmarks against documented sector failures like this one. A board-ready foresight document in 5 weeks.