Threat Modeling for AI and LLMs: OWASP Framework and Operational Mitigations

Threat modeling per AI generativa e LLM rischi e mitigazioni

Threat modeling for generative AI systems and Large Language Models systematically identifies vulnerabilities and compromise methods for models, analyzing not only technical aspects but also the socio-cultural, regulatory, and ethical contexts in which they operate.

For an overview of red teaming practices for GenAI systems, consult the complete guide to GenAI Red Teaming.

Reference frameworks for AI threat modeling

The NIST AI Risk Management Framework (AI RMF) provides a solid foundation for defining risks, threat sources, and specific attack objectives for AI systems. MITRE ATLAS maps real-world adversarial attack scenarios against machine learning models, while the OWASP AI Security and Privacy Guide offers practical guidelines for identifying and mitigating threats in AI systems.

Unlike traditional software-oriented frameworks, these tools address specific AI challenges such as algorithmic bias, CBRN (Chemical, Biological, Radiological, Nuclear) risks, CSAM (Child Sexual Abuse Material), and NCII (Non-Consensual Intimate Images), which require dedicated assessment approaches.

Operational threat modeling process for AI systems

The threat modeling process for AI systems is divided into four phases:

  1. Architecture modeling: mapping system components, data flows, interfaces, and supply chain dependencies.
  2. Threat identification: listing technical and contextual threats using frameworks like MITRE ATLAS and the OWASP AI Top 10.
  3. Mitigation definition: establishing security controls proportional to the identified risk.
  4. Iterative validation: testing and updating the model based on new threats and architectural changes.

Mapping threats to architectural components

Every AI system component presents specific attack surfaces. The data collection phase can be compromised via data poisoning; training can be subject to backdoor attacks; inference APIs are exposed to prompt injection and model extraction. Mapping OWASP threats to architectural components allows for the identification of which controls to apply at each stage of the model lifecycle, from data collection to production deployment.

Responsible AI and Trustworthy AI threats

Beyond technical vulnerabilities, AI systems must address risks related to fairness, accountability, and transparency. A model can produce discriminatory output even without malicious intent, or generate harmful content that violates ethical or regulatory policies. Threat modeling must therefore include scenarios of systemic bias, lack of explainability, and potential misuse of the model, evaluating the impact on specific communities and different regulatory contexts.

Differences from traditional software

AI models are distinguished by the unpredictability of their behavior, especially under boundary conditions or adversarial attacks. Unlike deterministic software, an LLM can produce unexpected output even with seemingly innocuous input. Threat modeling must therefore consider the entire supply chain: data collection and storage, training, testing, deployment, monitoring, and continuous model updating.

Attack scenarios and operational mitigations

Prompt injection

An attacker constructs malicious inputs to bypass LLM safeguards and execute unintended commands. Effective mitigations: rigorous input validation, contextual filters, response sandboxing, and separation between system instructions and user content.

Deepfake manipulations

The use of GANs, Diffusion Models, and LLMs allows for the creation of fictitious audio or video to impersonate corporate figures and induce fund transfers or the disclosure of sensitive data. Countermeasures: multi-factor verification protocols for critical communications, staff training on deepfake recognition, and automated detection systems.

RAG (Retrieval-Augmented Generation) vulnerabilities

A malicious actor inserts content with phishing links or malware into external sources that the RAG system integrates into its responses. If the LLM returns this content without validation, users may be induced to visit malicious sites. Validation of retrieved content is required, along with careful moderation and sanitization of outputs before presentation to the user.

Malicious code generation

The LLM may suggest code containing backdoors or intentional vulnerabilities. Continuous verification of generated code, the use of static analysis tools, and awareness of LLM limitations are fundamental to preventing the introduction of risks into the development cycle.

Components and attack surfaces to analyze

Comprehensive threat modeling must cover all threat vectors relevant to the AI system:

  • Model architecture and data flows between components
  • Data collection, storage, training, and testing pipelines
  • Deployment channels, inference APIs, and monitoring systems
  • Interfaces between models, external data sources, and end users
  • Supply chain of pre-trained models and third-party dependencies

Multilevel approach and operational benefits

Every AI application operates with specific assets, architecture, and user bases. Integrating threat modeling with technical and social red teaming activities allows for balancing human oversight, bias mitigation, and systemic risk assessment. Security measures thus become more aligned with the organization’s real needs and intended use contexts.

An often underestimated element is the continuous monitoring of external threats: knowing which actors are developing attack techniques against AI systems, which vulnerabilities are being discussed in underground forums, and which indicators of compromise emerge over time is an integral part of a mature defensive posture. A structured threat intelligence and digital risk protection service allows the threat modeling process to be fueled with updated data on real threats, making mitigations more precise and timely.

Adopting a structured approach to threat modeling for AI systems allows for identifying vulnerabilities before deployment, reducing exposure to regulatory and reputational risks, and building stakeholder trust through transparent and verifiable security practices.

  • What are the main frameworks for AI threat modeling?
  • The most widely used frameworks are the NIST AI RMF for risk management, MITRE ATLAS for mapping adversarial attacks, and the OWASP AI Security Guide for practical security guidelines.
  • How does AI threat modeling differ from traditional threat modeling?
  • AI threat modeling must consider the unpredictability of model behavior, risks related to bias and fairness, and the entire supply chain of data and pre-trained models, in addition to classic technical vulnerabilities.
  • What are Responsible AI threats?
  • These are risks related to fairness, accountability, transparency, and the ethical use of AI models, which can produce discrimination or harmful content even without malicious intent from developers.
  • What are the most common attacks against LLM systems?
  • The most frequent attacks include prompt injection to bypass safeguards, deepfake manipulations to impersonate users, RAG vulnerabilities that introduce malicious content, and code generation with backdoors.
  • How are vulnerabilities in RAG systems mitigated?
  • Effective mitigations include rigorous validation of content retrieved from external sources, output moderation, link sanitization, and verification of the reliability of sources integrated into the system.

Further reading

To learn more about red teaming practices and mitigation strategies for generative AI systems, consult these articles:

Protect your organisation with Threat Intelligence and Digital Risk Protection.

Choose ISGroup for a practical, tailored engagement:

  • A focused assessment of your environment and requirements
  • Clear findings with a prioritised, actionable roadmap
  • Direct support from experienced specialists through remediation and implementation
Talk to an expert