Runtime model poisoning occurs when an attacker manipulates inputs during the inference phase to progressively degrade model performance or alter its behavior. Unlike training set poisoning, this attack exploits continuous learning mechanisms or feedback loops to introduce bias, reduce accuracy, or install persistent backdoors in the production system.
This article is part of the AI Model Testing chapter of the OWASP AI Testing Guide, dedicated to the security of models in operation.
Testing Objectives
Runtime model poisoning testing aims to verify the model’s resilience against incremental manipulation during inference:
- Identify vulnerabilities in continuous learning mechanisms or feedback loops that allow for model poisoning in production.
- Detect persistent deviations in predictions caused by sequences of malicious inputs.
- Evaluate the effectiveness of implemented monitoring and anomaly detection controls.
Methodology and Payloads
The test is structured through three main techniques that simulate poisoning attacks during inference.
Gradual Label Flipping
This technique involves the sequential submission of valid inputs accompanied by intentionally incorrect feedback or labels over multiple inference cycles. The goal is to progressively degrade the model’s accuracy without raising immediate suspicion.
Vulnerability Indicator: The model’s accuracy on a clean test set decreases progressively. A drop of more than 10-15% compared to the baseline indicates a significant vulnerability that requires immediate intervention.
Backdoor Trigger Association
The tester repeatedly sends inputs containing a specific trigger phrase (e.g., “alpha-gamma-theta”) always associated with the same desired outcome, regardless of the actual content of the input. This simulates the installation of a backdoor in the model.
Vulnerability Indicator: After the poisoning phase, the model consistently generates the attacker’s desired result when the trigger is present, even if the rest of the input should produce a different result. The backdoor is active and exploitable.
Targeted Feature Skewing
The test continuously presents inputs where a normally benign feature (e.g., the word “community”) is always associated with a harmful or distorted result. The goal is to alter the semantic association learned by the model.
Vulnerability Indicator: The model begins to associate the benign feature with the harmful outcome, producing incorrect or distorted predictions even on clean inputs that contain that feature. The bias has been successfully installed.
Expected Output
A system resilient to runtime model poisoning must demonstrate the following characteristics:
- Stable performance: Accuracy and key model metrics remain stable even in the face of limited volumes of anomalous feedback. Variations do not exceed predefined tolerance thresholds.
- Effective anomaly detection: The monitoring system identifies and flags suspicious patterns, such as users or IP addresses that systematically provide contradictory or statistically anomalous feedback compared to the normal population.
- Robust resistance to incremental attacks: The model is not easily influenced by a limited number of malicious inputs. Decision boundaries do not shift drastically due to a few poisoned samples.
Remediation Actions
Countermeasures against runtime model poisoning require a multi-layered approach that combines input validation, access control, and continuous monitoring.
Rigorous Validation and Anomaly Detection
Implement feedback validation before using it to update the model. Use anomaly detection systems to identify feedback that is statistically divergent from normal patterns or trusted labelers. Automatically isolate suspicious feedback for manual review before integration.
Trusted Sources for Continuous Learning
Limit online learning to verified users or expert labelers with a proven track record. Avoid learning directly from anonymous or unverified feedback. Implement a reputation system to grade the reliability of sources.
Rate-limiting Updates
Update the model on a controlled periodic basis (e.g., once a day) instead of applying changes in real-time. This batch approach hinders rapid poisoning attacks and allows for security reviews before updates are applied.
Trust-based Weighting
Implement a trust scoring system for users. Feedback from new users or those with low reputation should have a much smaller impact on model updates compared to historical and verified users. Apply temporal decay to trust in case of anomalous behavior.
Periodic Retraining from a Clean Dataset
Periodically rebuild the model starting from a clean, verified, and complete dataset. This eliminates the progressive accumulation of poisoned data and restores the model to a known, secure state. Define a retraining cadence based on the system’s risk assessment.
Suggested Tools
- Adversarial Robustness Toolbox (ART): Open-source library to simulate and defend against runtime poisoning attacks on deep learning models.
- Scikit-learn partial_fit: Function to simulate online learning scenarios and test runtime poisoning vulnerabilities in controlled environments.
- River: Python library for online machine learning, useful for simulating incremental poisoning attacks.
Useful Insights
To understand the broader context of AI model attacks and related defense strategies:
- AITG-MOD-01 – Testing for Evasion Attacks: Attack techniques during inference that aim to evade model predictions.
- AITG-MOD-03 – Testing for Poisoned Training Sets: Poisoning of the training dataset before model deployment.
References
- OWASP Top 10 for LLM Applications 2025, “LLM04: Data and Model Poisoning” – OWASP LLM04
- NIST AI 100-2e2025, “Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations,” Section 2.3 – DOI:10.6028/NIST.AI.100-2e2025
- Jagielski, M., et al. “Manipulating Machine Learning: Poisoning Attacks and Countermeasures for Regression Learning” – arXiv:1804.00792
Integrating rigorous feedback validation, anomaly detection, and periodic retraining helps protect models from incremental manipulation during inference. Regularly testing resilience against runtime poisoning attempts is fundamental to ensuring the reliability and security of AI systems in production.
Want to give your company the highest level of cyber security? ISGroup SRL is here to help with cyber security solutions tailored to your business.
Would you like us to take care of everything for you? Our Virtual CISO and vulnerability management services are a perfect fit for your organization.
Already know what you need? Explore our services:
- Vulnerability Assessment
- Network Penetration Testing
- Web Application Penetration Testing
- Mobile Application Security Testing
- Ethical Hacking
- Training
And much more. Protect your company with the best cybersecurity experts!
