Public Database
Case Studies
Every approved AI failure case, classified against the AI Blindspot Framework. New to AIBlindspot? Start with the overview or the methodology.
Multi-Agent AI Systems Combining Capabilities to Defeat Security Safeguards
Networks of specialised AI agents can pool distinct capabilities, access rights, and knowledge to breach defences that no single agent could overcome alone. Attribution becomes intractable across diffuse agent networks, severely degrading incident response and recovery timelines.
AI-Enabled Swarm Attacks Overwhelm Single-Agent Security Assumptions
Coordinated networks of low-resource AI agents can collectively overwhelm defences designed around single well-resourced threat actors. Boards must reassess infrastructure resilience and security architecture before multi-agent offensive capabilities mature further.
Multi-Agent AI Systems Introduce Novel Security Vulnerabilities
Interconnected AI agents create attack surfaces and threat vectors that do not exist in single-model deployments. Boards must extend existing cyber-risk frameworks to account for emergent vulnerabilities across agent-to-agent interactions.
Collective AI Systems Develop Emergent Goal-Directed Behaviour
Individually narrow AI agents, when combined, can produce coordinated goal-directed behaviour unintended by any single system's design. Boards cannot rely on component-level oversight alone; emergent collective effects demand system-wide governance and accountability frameworks.
Corrupted Training Loop from Undesirable Model Outputs
Feeding flawed or inappropriate model outputs back into retraining pipelines degrades model integrity and produces unpredictable system behaviour. Boards risk compounding operational failures if governance frameworks lack controls over what data enters the retraining cycle.
AI Agents as Attack Surface for Principal Compromise
AI agents acting on behalf of organisations introduce exploitable vulnerabilities that expose private data and enable manipulation of delegated actions. Boards face liability for agent-driven harms and must govern AI delegation with the same rigour as human authorisation controls.
Cascading Security Failures in Multi-Agent AI Systems
Localised attacks on networked large language model agents can propagate rapidly, producing system-wide failures that are difficult to detect or contain. Boards face amplified liability exposure and operational disruption where AI orchestration layers lack robust authentication and incident isolation controls.
Model Extraction Attack Exposes Proprietary AI Architecture and Parameters
Adversaries systematically query deployed AI systems to reconstruct proprietary model architecture, parameters, and hyperparameters without authorisation. Organisations face loss of competitive advantage, potential regulatory scrutiny over data governance, and liability where extracted models encode sensitive training data.
Open-Source AI Models Repurposed for Unintended or Harmful Applications
Publicly available AI models are being systematically retrained on illicit data sources, redirecting their capabilities far beyond developer intent. Regulators and boards face accountability gaps when open-source tools enable harmful use cases outside any governed deployment context.
Jailbreaking Dismantles AI Safety Controls to Enable Unrestricted Harmful Output
Jailbreaking techniques systematically remove all safety filters from generative AI models, granting actors unrestricted ability to produce harmful, biased, or offensive content at scale. Organisations deploying AI systems face reputational, legal, and regulatory exposure where safety guardrails can be wholly neutralised by determined adversaries.
Generative AI Model Exposes Sensitive Personal Data Used in Training
Generative AI systems can be induced to reproduce personally identifiable information and medical records embedded in their training data. Organisations face regulatory liability and reputational harm where such disclosures breach data protection obligations.
Opaque Training Data Provenance Undermines Model Explainability
AI models trained without documented data collection and curation processes cannot be reliably explained or audited. Regulators and boards lose the assurance needed to approve deployment or defend decisions under scrutiny.
PII and Sensitive Data Leakage Through AI Training Datasets
AI models trained or fine-tuned on data containing personal or sensitive information risk disclosing that information through model outputs. Boards face regulatory exposure under data protection law and reputational liability if such disclosures affect individuals at scale.
Prompt Injection Attack Manipulates Generative AI Output
Adversaries exploit prompt structure to override AI instructions and produce unauthorised or harmful outputs. Boards must ensure AI deployments include input validation controls and adversarial testing as baseline governance requirements.
Prompt Priming Causes Generative Models to Leak Personal Training Data
Generative AI models can be manipulated via crafted prompts to reproduce personal data absorbed during training, creating material data protection exposure. Organisations deploying such models face regulatory liability and reputational harm if personal information is disclosed without consent.
Adversarial Input Manipulation Causes AI Model to Produce Incorrect Outputs
Evasion attacks exploit trained AI models by introducing subtle perturbations to input data, causing the model to return false or misleading results. Boards must treat adversarial robustness as a core governance requirement, since compromised model outputs can distort regulated decisions and expose firms to material liability.
Prompt Leakage Exposes Confidential AI System Instructions
Adversarial prompt attacks can extract hidden system instructions, revealing proprietary logic, safety constraints, or sensitive configuration data. Organisations deploying AI systems face material risk of intellectual property loss and regulatory exposure if system prompts are inadequately protected.
Jailbreak Attacks Bypass AI Safety Controls in Transport Systems
Adversarial jailbreaking techniques circumvent embedded guardrails in AI models, enabling restricted or unsafe actions to be executed without authorisation. Transport operators face regulatory exposure under SEC disclosure requirements when such vulnerabilities compromise safety-critical or operationally sensitive AI deployments.
Sensitive Personal Data Exposed Through AI Model Prompts
Users or systems are transmitting personal and sensitive personal information directly within prompts sent to AI models, creating uncontrolled data exposure. Organisations face regulatory liability under data protection law and reputational risk if such data is retained, logged, or used in model training.
Generative AI Weaponised to Produce Hateful and Obscene Content
Generative AI models can be deliberately prompted to produce hateful, abusive, and profane material at scale. Boards face regulatory exposure and reputational liability if deployed systems lack robust content controls and misuse detection.
Generative AI Deployed to Spread Targeted Disinformation
Generative AI models can be weaponised to produce convincing false information designed to deceive or manipulate specific audiences at scale. Boards face regulatory exposure and reputational liability where such content is linked to their platforms, products, or market communications.
AI Model Deployed Outside Its Intended Purpose
Organisations risk systematic failure when AI models are applied to tasks beyond their original design parameters. Boards must enforce strict deployment governance to prevent liability exposure and reputational harm from misapplied systems.
Confidential Data Leaked via Model Prompt Submission
Sensitive organisational data entered into AI prompts may be exposed to third-party model providers or logged in external systems. Boards must establish prompt governance policies to prevent uncontrolled disclosure of confidential information.
AI Model Delivers Insufficient Accuracy for Its Intended Task
An AI model fails to meet performance requirements due to flawed engineering or drift between training inputs and real-world data. Boards face operational disruption and liability exposure when deployed models cannot be relied upon to produce correct outputs.
Beyond accidental failureNational Security
We also track 20 hostile uses of AI.
The public database covers AI that fails by accident. AIBlindspot National Security — exclusive to the Defence tier — tracks AI used as a weapon, mapped by capability:
- State-Sponsored AI Operations
- 6
- AI-Enabled Disinformation
- 5
- Adversarial Attacks on AI
- 0
- Autonomous Weapon Incidents
- 1
- AI-Assisted Cyber Attacks
- 5
- Dual-Use AI Misuse
- 3