Public Database
Case Studies
Every approved AI failure case, classified against the AI Blindspot Framework. New to AIBlindspot? Start with the overview or the methodology.
Chaotic Dynamics in Multi-Agent AI Systems Grow Harder to Predict at Scale
Research confirms that chaotic, unpredictable behaviour becomes increasingly probable as the number of interacting AI agents grows. Organisations deploying multi-agent systems cannot assume stable or foreseeable outputs, undermining risk modelling and operational assurance.
Multi-Agent AI Systems Produce Unstable Cyclic Behaviour Under Standard Learning Rules
When multiple AI agents interact, standard learning algorithms proven stable in isolation can generate non-convergent cycles and unpredictable system trajectories. Governments deploying multi-agent AI in critical services cannot rely on single-agent safety assurances, requiring new oversight frameworks.
Abrupt Behavioural Phase Transitions in Multi-Agent AI Systems
Minor system changes, such as adding new agents or shifting data distributions, can trigger sudden, unpredictable collapses or reversals in AI system behaviour. Boards cannot rely on gradual performance signals as warning indicators, making conventional monitoring and risk thresholds unreliable.
AI Agents Unable to Form Credible Commitments in Collaborative Tasks
Multi-agent AI systems lack mechanisms to establish reliable trust or binding commitments, causing coordination failures between AI agents and between AI and human counterparts. Organisations deploying agentic AI pipelines face compounded risk of failed transactions, misaligned outcomes, and unverifiable AI behaviour at scale.
Emergent Agency Risk in Composed Multi-Agent AI Systems
Combining individually benign AI agents can produce unexpected goals or capabilities absent from any single component. Boards cannot assume system-level safety from component-level assurance alone.
Multi-Agent AI Systems Combining Capabilities to Defeat Security Safeguards
Networks of specialised AI agents can pool distinct capabilities, access rights, and knowledge to breach defences that no single agent could overcome alone. Attribution becomes intractable across diffuse agent networks, severely degrading incident response and recovery timelines.
AI-Enabled Swarm Attacks Overwhelm Single-Agent Security Assumptions
Coordinated networks of low-resource AI agents can collectively overwhelm defences designed around single well-resourced threat actors. Boards must reassess infrastructure resilience and security architecture before multi-agent offensive capabilities mature further.
Multi-Agent AI Systems Introduce Novel Security Vulnerabilities
Interconnected AI agents create attack surfaces and threat vectors that do not exist in single-model deployments. Boards must extend existing cyber-risk frameworks to account for emergent vulnerabilities across agent-to-agent interactions.
Collective AI Systems Develop Emergent Goal-Directed Behaviour
Individually narrow AI agents, when combined, can produce coordinated goal-directed behaviour unintended by any single system's design. Boards cannot rely on component-level oversight alone; emergent collective effects demand system-wide governance and accountability frameworks.
Corrupted Training Loop from Undesirable Model Outputs
Feeding flawed or inappropriate model outputs back into retraining pipelines degrades model integrity and produces unpredictable system behaviour. Boards risk compounding operational failures if governance frameworks lack controls over what data enters the retraining cycle.
AI Agents as Attack Surface for Principal Compromise
AI agents acting on behalf of organisations introduce exploitable vulnerabilities that expose private data and enable manipulation of delegated actions. Boards face liability for agent-driven harms and must govern AI delegation with the same rigour as human authorisation controls.
Cascading Security Failures in Multi-Agent AI Systems
Localised attacks on networked large language model agents can propagate rapidly, producing system-wide failures that are difficult to detect or contain. Boards face amplified liability exposure and operational disruption where AI orchestration layers lack robust authentication and incident isolation controls.
Model Extraction Attack Exposes Proprietary AI Architecture and Parameters
Adversaries systematically query deployed AI systems to reconstruct proprietary model architecture, parameters, and hyperparameters without authorisation. Organisations face loss of competitive advantage, potential regulatory scrutiny over data governance, and liability where extracted models encode sensitive training data.
Open-Source AI Models Repurposed for Unintended or Harmful Applications
Publicly available AI models are being systematically retrained on illicit data sources, redirecting their capabilities far beyond developer intent. Regulators and boards face accountability gaps when open-source tools enable harmful use cases outside any governed deployment context.
Jailbreaking Dismantles AI Safety Controls to Enable Unrestricted Harmful Output
Jailbreaking techniques systematically remove all safety filters from generative AI models, granting actors unrestricted ability to produce harmful, biased, or offensive content at scale. Organisations deploying AI systems face reputational, legal, and regulatory exposure where safety guardrails can be wholly neutralised by determined adversaries.
Generative AI Model Exposes Sensitive Personal Data Used in Training
Generative AI systems can be induced to reproduce personally identifiable information and medical records embedded in their training data. Organisations face regulatory liability and reputational harm where such disclosures breach data protection obligations.
PII and Sensitive Data Leakage Through AI Training Datasets
AI models trained or fine-tuned on data containing personal or sensitive information risk disclosing that information through model outputs. Boards face regulatory exposure under data protection law and reputational liability if such disclosures affect individuals at scale.
Prompt Injection Attack Manipulates Generative AI Output
Adversaries exploit prompt structure to override AI instructions and produce unauthorised or harmful outputs. Boards must ensure AI deployments include input validation controls and adversarial testing as baseline governance requirements.
Prompt Priming Causes Generative Models to Leak Personal Training Data
Generative AI models can be manipulated via crafted prompts to reproduce personal data absorbed during training, creating material data protection exposure. Organisations deploying such models face regulatory liability and reputational harm if personal information is disclosed without consent.
Adversarial Input Manipulation Causes AI Model to Produce Incorrect Outputs
Evasion attacks exploit trained AI models by introducing subtle perturbations to input data, causing the model to return false or misleading results. Boards must treat adversarial robustness as a core governance requirement, since compromised model outputs can distort regulated decisions and expose firms to material liability.
Prompt Leakage Exposes Confidential AI System Instructions
Adversarial prompt attacks can extract hidden system instructions, revealing proprietary logic, safety constraints, or sensitive configuration data. Organisations deploying AI systems face material risk of intellectual property loss and regulatory exposure if system prompts are inadequately protected.
Jailbreak Attacks Bypass AI Safety Controls in Transport Systems
Adversarial jailbreaking techniques circumvent embedded guardrails in AI models, enabling restricted or unsafe actions to be executed without authorisation. Transport operators face regulatory exposure under SEC disclosure requirements when such vulnerabilities compromise safety-critical or operationally sensitive AI deployments.
Sensitive Personal Data Exposed Through AI Model Prompts
Users or systems are transmitting personal and sensitive personal information directly within prompts sent to AI models, creating uncontrolled data exposure. Organisations face regulatory liability under data protection law and reputational risk if such data is retained, logged, or used in model training.
Generative AI Weaponised to Produce Hateful and Obscene Content
Generative AI models can be deliberately prompted to produce hateful, abusive, and profane material at scale. Boards face regulatory exposure and reputational liability if deployed systems lack robust content controls and misuse detection.
Beyond accidental failureNational Security
We also track 20 hostile uses of AI.
The public database covers AI that fails by accident. AIBlindspot National Security — exclusive to the Defence tier — tracks AI used as a weapon, mapped by capability:
- State-Sponsored AI Operations
- 6
- AI-Enabled Disinformation
- 5
- Adversarial Attacks on AI
- 0
- Autonomous Weapon Incidents
- 1
- AI-Assisted Cyber Attacks
- 5
- Dual-Use AI Misuse
- 3