The operational blueprint for GenAI Red Teaming defines a structured four-phase approach to assess the security of generative artificial intelligence systems: Model, Implementation, System, and Runtime. Each phase includes detailed checklists, assessment tools, and specific deliverables to identify vulnerabilities and test the defenses adopted throughout the model’s entire lifecycle.
For an overview of GenAI Red Teaming and its role in AI system security, consult the complete guide to GenAI Red Teaming.
The four phases of the blueprint
Phase 1: Model Evaluation
Model evaluation focuses on the intrinsic security and robustness of the AI model, verifying:
- Lifecycle security (MDLC): model provenance, malware injection risk, training data pipeline security
- Robustness: testing for toxicity, bias, alignment, and attempts to bypass intrinsic defenses
- Inference attacks: assessment of architecture, training, parameters, fingerprinting, and deployment
- Extractability: testing for knowledge extraction, training data, weights, embeddings, policies, and prompt templates
- Instruction tuning: retention manipulation, fine-tuning limits, collisions, and instruction priority
- Socio-technological risks: demographic bias, hate speech, harmful content, toxicity, stereotypes, discrimination
- Data risk: access violations, IP extraction, watermarking, recovery, and reconstruction of sensitive data
- Alignment control: jailbreak effectiveness, prompt injection, value limits, safety layer bypass
- Adversarial robustness: attack patterns, unknown vulnerabilities, edge cases, emergent capabilities
- Technical damage vectors: code generation capabilities, support for cyber attacks, exposure of scripts or infrastructure vectors
Model phase deliverables:
- Vulnerability Report
- Robustness Assessment
- Defensive Mechanism Evaluation
- Risk Assessment Report
- Ethics and Bias Analysis
Phase 2: Implementation Evaluation
Implementation evaluation verifies application controls and security measures integrated into the system:
- Prompt safety: evasion, context manipulation, multi-message attack chains, role and persona-based attacks
- Knowledge retrieval security: poisoning in vector databases, manipulation of embeddings, cache, or retrieval results
- System architecture: model isolation bypass, firewall/proxy evasion, rate limiting and filtering bypass, cross-request correlation
- Content filtering: policy enforcement, filter evasion, multilingual consistency, context-aware manipulation
- Access control: authentication/authorization, session management, roles, privilege escalation, token and service-to-service control
- Agent/tool/plugin security: tool access control, sandboxing, agent behavior, feedback loops, function call security
Phase 3: System Evaluation
System evaluation examines infrastructural components, interactions between the model and other elements, and the supply chain:
- Remote Code Execution: code execution from model output, command injection, template injection, path manipulation
- Sandbox escape: side channels, timing/power/cache/memory/network analysis, error leakage
- Supply chain: dependency integrity, repository security, pipelines, container images, third parties
- Risk propagation: error propagation, system interaction chains, cross-service impact, and data chain impact
- System integrity: output validation, input sanitization, version/config/backup/audit consistency
- Resource control: rate limiting bypass, exhaustion probes, quotas and capacity, DoS resilience
- Security measure efficacy: authentication, encryption, policy enforcement, incident response, monitoring, and alert coverage
- Control bypass: evasion of firewalls, proxies, WAFs, API gateways, monitoring, and enforcement gaps
Phase 4: Runtime / Human & Agentic Evaluation
Runtime evaluation analyzes vulnerabilities during real-world operations, human interaction, and agentic systems:
- Business process integration: AI-human hand-off, race conditions, privilege escalation, automated decision boundaries
- Multi-component AI: detection leakage between AIs, failover, cascade breakdown, cross-service authentication
- Over-reliance: over-trust, decisions without human oversight, fallback and degraded mechanisms
- Social engineering: prompt injection via operators, abuse of trust bonds, authority impersonation, manipulation of AI traits
- Downstream impact: manipulation propagation, integrity chaining, format-based injection, hallucinated content on dependent systems
- System boundary: API authentication/authorization, rate limit bypass, unauthorized access, input validation
- Monitoring evasion: detection blind spots, audit gaps, threshold manipulation, monitoring bypass
- Agent boundary: contextuality, decision limits, and agent capabilities
- Chain-of-custody: traceability of AI actions, audit of decision-making processes, intermediate accounting in workflows
- Agentic AI Red Teaming: control/hijacking of agent authorizations, checker-out-of-the-loop, chain impact, knowledge base poisoning, context manipulation, resource/service exhaustion, supply chain attacks
Benefits of the structured approach
Efficient risk identification
Early detection of issues at the model level allows for mitigating vulnerabilities before they propagate to subsequent phases, reducing remediation costs and risk exposure.
Multilayered defense
Combining model-level and system-level controls increases overall robustness. For example, Image Markdown vulnerabilities can be mitigated through both model controls and implementation-level filters.
Resource optimization
Distinguishing between model issues and system issues allows for targeted resource allocation, avoiding costly interventions on non-critical components and focusing efforts where they have the greatest impact.
Continuous improvement
Identifying root causes allows for effective improvement iterations. For example, when managing PII extraction errors, understanding whether the problem lies in the model or the implementation guides the choice of the most appropriate solution.
Comprehensive risk assessment
Comparing theoretical risks with real operational risks provides an accurate view of actual exposure and the effectiveness of adopted countermeasures.
Lifecycle view and assessment activities
Acquisition
During model acquisition, activities include:
- Model integrity verification
- Malware scanning
- Performance benchmarking
- Testing controls such as alignment and bias/toxicity prevention
Experimentation/Training
In the experimentation and training phase, the focus is on:
- Identifying vulnerabilities in base components
- Detecting abuse in data pipelines
- Verifying the security of fine-tuning processes
Serving/Inference
During service delivery, activities include:
- Runtime abuse detection
- Testing for RCE and SQL injection
- Attempts to bypass security and safety measures
- Monitoring interactions in production
Complete operational workflow
The GenAI Red Teaming process follows a structured workflow that includes:
- Scoping: defining the perimeter and objectives
- Resource discovery: mapping models, systems, and dependencies
- Scheduling: planning test activities
- Test execution: conducting checks according to checklists
- Reporting: documenting results
- Debrief: presenting and discussing findings
- Report updates: integrating feedback and insights
- Risk dispositioning: prioritizing and assigning remediation
- Postmortem review: analyzing lessons learned
- Retesting: verifying the effectiveness of corrections
Automated assessment tools
Automated tools for LLM assessment are particularly useful in the Model Evaluation phase but always require manual review of the results.
Advantages of automation
- Speed and coverage: a greater number of scenarios can be evaluated in less time
- Consistency: standardization of assessments via static datasets
- Advanced analysis: identifying patterns and behaviors difficult to detect manually
Limitations and considerations
The non-deterministic nature of generative models requires careful weighting of automated results. Tools can produce false positives and false negatives, making manual validation by experts essential.
Reusing results across phases
Information gathered during model evaluation can be reused in subsequent phases:
- Test cases: findings from the Model phase become scenarios to be verified in Implementation and System phases
- Prioritization: identified risks guide resource allocation in subsequent phases
- Model-independent tests: some controls (e.g., moderation filters) must be tested independently of the specific model
Useful resources
To effectively implement the blueprint and understand the broader context of GenAI Red Teaming, consult these resources:
- GenAI Red Teaming – overview of the framework and methodologies
- GenAI Red Teaming Techniques – in-depth look at operational techniques used in each phase
- GenAI Red Teaming Risks – detailed analysis of risks and threats to be assessed
- Red Teaming Tools and Datasets – overview of automated tools and reference datasets
- GenAI Red Teaming Metrics – KPIs and metrics to measure assessment effectiveness
- What is the difference between model evaluation and system evaluation?
- Model evaluation focuses on the intrinsic characteristics of the AI model (robustness, bias, alignment), while system evaluation examines the infrastructure, integrations, and components surrounding the model. This distinction allows for identifying whether a problem can be solved by improving the model or by intervening in the system architecture.
- Why do automated tools require manual validation?
- Generative models are non-deterministic, meaning they can produce different outputs for the same input. Automated tools can generate false positives (flagging non-existent issues) or false negatives (missing real vulnerabilities). Manual validation by experts is essential to correctly interpret results and contextualize them for the specific use case.
- How does the blueprint integrate with the model lifecycle?
- The blueprint aligns with the three main lifecycle phases: Acquisition (integrity verification and benchmarking), Experimentation/Training (testing pipelines and base components), and Serving/Inference (runtime abuse detection and operational security testing). Each lifecycle phase requires specific assessment activities that the blueprint organizes in a structured way.
- What are the main deliverables of a GenAI Red Teaming exercise?
- Deliverables include: Vulnerability Report (list of identified vulnerabilities), Robustness Assessment (evaluation of model resistance), Defensive Mechanism Evaluation (effectiveness of controls), Risk Assessment Report (risk analysis), and Ethics and Bias Analysis (ethical and bias evaluation). These documents guide remediation and continuous improvement activities.
- How is the evaluation of agentic systems managed?
- Agentic systems require specific tests in the Runtime/Agentic phase, including: control and hijacking of authorizations, chain impact (impact of action chains), knowledge base poisoning, context manipulation, resource exhaustion, and supply chain attacks. The complexity of agents requires particular attention to decision boundaries and action traceability.
Want to give your company the highest level of cyber security? ISGroup SRL is here to help with cyber security solutions tailored to your business.
Would you like us to take care of everything for you? Our Virtual CISO and vulnerability management services are a perfect fit for your organization.
Already know what you need? Explore our services:
- Vulnerability Assessment
- Network Penetration Testing
- Web Application Penetration Testing
- Mobile Application Security Testing
- Ethical Hacking
- Training
And much more. Protect your company with the best cybersecurity experts!
