Appendix B of the OWASP “LLM Red Teaming” project presents a list of tools and datasets, developed and selected based on the collective experience of the operators and authors involved. The catalog includes resources designed for Red Teaming on GenAI and LLMs. The list is not exhaustive and is updated with new selected solutions. Organizations wishing to include specific tools for GenAI Red Teaming in the catalog should contact the OWASP team to propose their inclusion. The use of tools from public repositories involves risks: it is the users’ responsibility to evaluate their security before adoption.
For a complete overview of the methodologies and the operational framework, consult the GenAI Red Teaming guide.
Tools for LLM and GenAI Red Teaming
-
ASCII Smuggler: A tool for hiding content within prompts.
https://embracethered.com/blog/ascii-smuggler.html (Open Source) -
Adversarial Attacks and Defences in Machine Learning (AAD) Framework: A Python framework for defending ML models against adversarial examples.
https://github.com/changx03/adversarial_attack_defence (Source available) -
Adversarial Robustness Toolbox (ART): A Python library for ML security.
https://github.com/Trusted-AI/adversarial-robustness-toolbox (MIT License) -
Advertorch: A Python toolbox for research on robustness and adversarial attacks in PyTorch.
https://github.com/BorealisAI/advertorch (GNU LGPL v3.0) -
CleverHans: A Python library for testing the vulnerability of ML systems to adversarial examples.
https://github.com/cleverhans-lab/cleverhans (MIT License) -
CyberSecEval: A benchmark for quantifying cybersecurity risks and capabilities in LLMs.
https://ai.meta.com/research/publications/cyberseceval-3-advancing-the-evaluation-of-cybersecurity-risks-and-capabilities-in-large-language-models/ (MIT License) -
DeepEval: LLM evaluation, unit testing, and multiple output metrics.
https://github.com/confident-ai/deepeval (Apache License 2.0) -
Deep-pwning: A lightweight framework for evaluating the robustness of ML models against motivated adversaries.
https://github.com/cchio/deep-pwning (MIT License) -
Dioptra: A platform for testing the reliability of AI systems.
https://pages.nist.gov/dioptra/index.html (CC BY 4.0) -
Foolbox: A tool for adversarial attacks and ML robustness benchmarking in PyTorch, TensorFlow, and JAX.
https://github.com/bethgelab/foolbox (MIT License) -
Garak: A kit for GenAI red-teaming and assessment.
https://garak.ai/ (Apache License 2.0)
https://github.com/NVIDIA/garak -
Giskard: A testing suite for ML and LLMs.
https://www.giskard.ai/ (Apache License 2.0) -
Generative Offensive Agent Tester (GOAT): An automated system that simulates adversarial conversations to identify vulnerabilities in LLMs.
https://arxiv.org/abs/2410.01606 -
Gymnasium: A Python library with standard APIs for reinforcement learning testing and development.
https://github.com/Farama-Foundation/Gymnasium (MIT License) -
Harmbench: A scalable open-source framework for evaluating automated Red Teaming methods and LLM attacks/defenses.
https://github.com/centerforaisafety/HarmBench (MIT License) -
HouYi: A framework for prompt injection attacks in LLM-integrated applications.
https://github.com/LLMSecurity/HouYi?tab=readme-ov-file (Apache License 2.0) -
JailbreakingLLMs – PAIR: Jailbreak tests for LLMs using Prompt Automatic Iterative Refinement.
https://github.com/patrickrchao/JailbreakingLLMs (MIT License) -
Llamator: Pentesting for RAG applications.
https://github.com/RomiconEZ/LLaMator (CC) -
LLM Attacks: Automation in constructing adversarial attacks on LLMs.
https://llm-attacks.org/ (MIT License) - LLM Canary: Benchmarking and scoring for LLMs. (Apache License 2.0)
-
Modelscan: Detection of Model Serialization attacks.
https://github.com/protectai/modelscan (Apache License 2.0) -
MoonShot: A modular tool for evaluating LLM applications.
https://github.com/aiverify-foundation/moonshot (Apache Software License 2) -
Prompt Fuzzer: A security testing tool for GenAI prompts against dynamic LLM attacks.
https://github.com/prompt-security/ps-fuzz (MIT License) -
Promptfoo: Red Teaming, penetration testing, and vulnerability scanning for LLMs.
https://github.com/promptfoo/promptfoo (MIT License) -
ps-fuzz: An interactive tool for GenAI prompt security.
https://github.com/prompt-security/ps-fuzz (MIT License) -
PromptInject: Quantitative analysis of LLM robustness against adversarial prompts.
https://github.com/agencyenterprise/PromptInject (MIT License) -
Promptmap: Prompt injection on ChatGPT instances.
https://github.com/utkusen/promptmap (MIT License) -
Python Risk Identification Toolkit (PyRIT): A Microsoft library for evaluating the robustness of LLM endpoints regarding content such as hallucinations, bias, and prohibited topics.
https://github.com/Azure/PyRIT (MIT License) -
SplxAI: Automated Red Teaming for Conversational AI.
https://splx.ai/ -
StrongREJECT: A jailbreak benchmark with an evaluation methodology.
https://github.com/alexandrasouly/strongreject,
https://arxiv.org/abs/2402.10260 (MIT License)
Datasets for GenAI Red Teaming
-
AdvBench: Universal and transferable adversarial attacks on aligned language models.
https://github.com/llm-attacks/llm-attacks (Open Source) -
BBQ Bias Benchmark for Question Answering: A bias benchmark for QA tasks.
https://github.com/nyu-mll/BBQ (Open Source) -
Bot Adversarial Dialogue Dataset: A dataset of adversarial dialogues for bots.
https://github.com/facebookresearch/ParlAI/tree/main/parlai/tasks/bot_adversarial_dialogue (Open Source) -
HarmBench: A standard framework for automated Red Teaming and robust refusal.
https://github.com/centerforaisafety/HarmBench (Open Source) -
JailbreakBench: An open benchmark for LLM robustness against jailbreaking.
https://github.com/JailbreakBench/jailbreakbench (Open Source) -
HAP: Efficient models for detecting hate, abuse, and profanity.
https://arxiv.org/abs/2402.05624 (Open Source)
Additional AI Security Resources
The OWASP project also highlights the AI Security Solutions Landscape, a resource that collects both traditional and emerging security controls to address LLM and Generative AI risks mapped in the OWASP Top 10.
Useful Insights
To learn more about operational methodologies and reference frameworks for Red Teaming on GenAI systems, consult these articles:
- GenAI Red Teaming: A complete guide to the security of generative artificial intelligence systems
- Operational Red Teaming techniques for LLMs and GenAI
- Metrics and KPIs to evaluate the effectiveness of GenAI Red Teaming
- Risks and threats in GenAI systems: mapping and prioritization
- Red Teaming for Agentic AI systems: challenges and approaches
Want to give your company the highest level of cyber security? ISGroup SRL is here to help with cyber security solutions tailored to your business.
Would you like us to take care of everything for you? Our Virtual CISO and vulnerability management services are a perfect fit for your organization.
Already know what you need? Explore our services:
- Vulnerability Assessment
- Network Penetration Testing
- Web Application Penetration Testing
- Mobile Application Security Testing
- Ethical Hacking
- Training
And much more. Protect your company with the best cybersecurity experts!
