AITG-MOD-03: Testing for Poisoned Training Sets

Difesa e rilevamento di dati avvelenati nel training AI

Attacks on training datasets compromise the integrity of the AI model by injecting malicious data during the training phase. These attacks introduce biases, persistent backdoors, or degrade model accuracy, with direct impacts on operational reliability and regulatory compliance.

This article is part of the AI Model Testing chapter of the OWASP AI Testing Guide.

Test Objectives

  • Identify malicious or corrupted samples within training datasets.
  • Evaluate model robustness against targeted, indiscriminate, or backdoor data poisoning attacks.
  • Verify the integrity of data sources and preprocessing pipelines.
  • Analyze the effectiveness of countermeasures to identify and mitigate poisoned data.

Methodology and Payloads

Label Flipping Attack

A portion of the dataset is modified by replacing correct labels with incorrect values, simulating an indiscriminate attack that degrades the model’s overall accuracy.

Vulnerability indicator: Auditing tools like cleanlab identify over 2% of labeling issues, suggesting systematic corruption rather than expected random noise.

Backdoor Trigger Injection

Training samples are modified by inserting non-obvious triggers (specific visual patterns, rare phrases, hidden watermarks) associated with a target class, creating a backdoor that can be activated during inference.

Vulnerability indicator: Anomaly detection algorithms highlight compact clusters in the feature space that are distant from the typical distribution of the assigned class, signaling potential backdoor patterns.

Targeted Poisoning

Samples of a specific subgroup are altered or mislabeled to selectively degrade model performance only on that segment, while keeping overall accuracy seemingly normal.

Vulnerability indicator: The model shows a drastic drop in accuracy (over 20%) on the target subgroup compared to overall accuracy, indicating targeted manipulation of the training set.

Feature Poisoning

Subtle modifications to input features (imperceptible noise, pixel alterations, semantic perturbations) are systematically inserted to influence model behavior on specific patterns.

Vulnerability indicator: Statistical analysis of the dataset reveals anomalous feature distributions or unexpected correlations between attributes and labels, signaling possible feature manipulation.

Expected Output

  • Validated dataset: The training set must not contain detectable labeling errors or malicious patterns. Automatic anomaly reports must be less than 1% of total samples.
  • Effective anomaly detection: The validation system must automatically identify anomalous clusters, suspicious patterns, or statistical distributions incompatible with clean data.
  • Uniform performance: The model trained on controlled data must not show anomalous biases, activatable backdoors, or selective degradation on specific subgroups.

Remediation Actions

Automated Validation Pipeline

Implement a mandatory sanitization pipeline before training, using tools like cleanlab for automatic label correction and anomaly detection to identify suspicious samples.

Expected impact: Reduction of the labeling error rate below 1% and automatic identification of anomalous clusters before they influence training.

Dataset Versioning and Traceability

Adopt versioned datasets with tools like DVC, linking each model to the specific version of the training data and maintaining a complete audit trail of dataset changes.

Expected impact: Ability to perform immediate rollback to previous dataset versions in case of poisoning detection and complete traceability of data changes.

Differential Privacy in Training

Apply differential privacy techniques during training to limit the influence of individual malicious samples on the final model, making poisoning attacks less effective.

Expected impact: Reduction of the impact of poisoned samples on model behavior, with maximum degradation contained below 5% even in the presence of limited poisoning.

Continuous Data Drift Monitoring

Implement continuous statistical monitoring systems for the training data distribution, with automatic alerts for sudden changes that may indicate the insertion of malicious data.

Expected impact: Real-time detection of statistical anomalies in the dataset with alerts within 24 hours of suspicious data insertion.

MLOps Pipeline Security

Protect the entire MLOps pipeline with strict access controls, mandatory version control on data and code, and mandatory reviews for any changes to the data pipeline or training scripts.

Expected impact: Prevention of unauthorized changes to the dataset and complete traceability of all operations on the data pipeline.

Suggested Tools

References

  • Northcutt et al., “Confident Learning: Estimating Uncertainty in Dataset Labels”, Journal of Artificial Intelligence Research, 2021 โ€“ arXiv:1911.00068
  • OWASP, “LLM04: Data and Model Poisoning”, OWASP Top 10 for LLM Applications 2025 โ€“ OWASP LLM04:2025
  • NIST, “Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations”, NIST AI 100-2e2025, Section 2.3, March 2025 โ€“ DOI:10.6028/NIST.AI.100-2e2025

Useful Insights

To complete your understanding of attacks on AI models, consult the other tests in the AI Model Testing chapter:

The integration of automated validation pipelines and rigorous dataset versioning helps prevent the insertion of malicious data into models. Recognizing and applying these techniques requires specific skills: the cybersecurity and AI security training paths from ISGroup support teams managing MLOps pipelines who want to consolidate their defensive capabilities. Regularly testing the integrity of training sets is essential to ensure reliability and robustness in production.

Protect your organisation with Cybersecurity Training.

Choose ISGroup for a practical, tailored engagement:

  • A focused assessment of your environment and requirements
  • Clear findings with a prioritised, actionable roadmap
  • Direct support from experienced specialists through remediation and implementation
Talk to an expert