AITG-MOD-06: Testing for Robustness to New Data

Test AITG-MOD-06 per Robustezza e Data Drift AI

The AITG-MOD-06 test identifies vulnerabilities related to the lack of robustness in AI models when exposed to new or out-of-distribution (OOD) data. These issues manifest as performance drops or unexpected behaviors when the model encounters distributions different from those used during training, compromising reliability and security.

This article is part of the AI Model Testing chapter of the OWASP AI Testing Guide.

Test Objectives

  • Evaluate the model’s resilience when facing new or previously unseen data distributions.
  • Identify vulnerabilities that cause significant performance degradation with OOD data.
  • Verify the effectiveness of defensive strategies in maintaining accuracy and stability in the event of distribution shifts.

Methodology and Payloads

Data Drift Simulation

Use tools like deepchecks or evidently to compare the statistical properties of training data with new production data. This approach allows for the detection of gradual or sudden changes in distributions that can compromise model performance.

Vulnerability Indicator: Significant drift in many features, with the mean shifting beyond 3 standard deviations or a Population Stability Index (PSI) greater than 0.25.

Out-of-Distribution (OOD) Inputs

Insert inputs that are semantically far from those known during training, such as providing an image of a car to a classifier trained only on dogs and cats. This test verifies whether the model is capable of recognizing when it is operating outside its domain of expertise.

Vulnerability Indicator: The model returns high-confidence predictions for known classes instead of flagging the input as unknown, such as classifying a car as a “dog” with 98% confidence.

Edge Case and Boundary Testing

Systematically generate inputs at the limits of expected ranges or rare but plausible scenarios, such as extreme values in numerical features or unusual combinations of attributes. This approach identifies zones of fragility where the model did not receive sufficient exposure during training.

Vulnerability Indicator: Erratic or highly uncertain predictions on edge cases, signaling a failure to generalize outside the core of the training distribution.

Expected Output

  • Stable performance on new data: Accuracy, precision, and recall should not drop beyond a predetermined threshold (5-10%) on data with moderate drift compared to training.
  • Correct handling of OOD inputs: A robust model provides low-confidence scores or explicitly classifies data as “unknown” when encountering out-of-distribution data, rather than generating erroneous high-confidence predictions.
  • Low data drift score: PSI lower than 0.1 and passing major validation checks between training data and new datasets.

Remediation Actions

Continuous Drift Monitoring

Integrate tools like deepchecks or evidently into MLOps pipelines to automatically detect data drift, concept drift, and performance decay, triggering alerts in case of anomalies.

Expected Impact: Timely detection of changes in distributions before they cause significant performance degradation in production.

Robust Training and Data Augmentation

Apply data augmentation to produce diversified datasets that expose the model to greater variations and encourage generalization. Include domain randomization and synthetic data generation techniques to broaden distributional coverage.

Expected Impact: Improved ability of the model to generalize on distributions other than training ones, reducing the risk of failure on new data.

Uncertainty Quantification

Design the model to express its degree of uncertainty using techniques such as ensemble methods, Bayesian neural networks, or probability calibration. Route cases with highly uncertain predictions to manual review.

Expected Impact: Automatic identification of OOD or ambiguous inputs, allowing for escalation to human operators instead of generating erroneous high-confidence predictions.

Periodic Retraining

Schedule regular retraining sessions on recent data inclusive of production data, keeping the model updated to changes in real-world distributions. Implement continuous learning strategies where appropriate.

Expected Impact: Maintenance of performance over time even in the presence of gradual drift, adapting the model to natural data evolution.

Domain Adaptation

In the presence of predictable drift, use targeted strategies to teach the model to remain invariant to expected changes. Apply transfer learning and fine-tuning techniques on specific target domains.

Expected Impact: Improved robustness on known or anticipatable distribution shifts, reducing the need for complete retraining.

Suggested Tools

  • DeepChecks: Python library for validating and testing ML models and data, with drift detection and other checks.
  • Evidently AI: Python library for evaluating, testing, and monitoring ML models in production with interactive reports on data drift and performance.
  • Alibi Detect: Python library for outlier, adversarial attack, and drift detection, with algorithms to identify OOD data.

Further Reading

To complete the evaluation of model robustness, consult the related tests that address other aspects of AI security:

References

  • Rabanser, Stephan, et al. “Failing Loudly: An Empirical Study of Methods for Detecting Dataset Shift.” NeurIPS 2019. arXiv:1810.11953
  • OWASP. “LLM05: Improper Output Handling.” OWASP Top 10 for LLM Applications 2025. OWASP LLM05
  • NIST. “Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations.” NIST AI 100-2e2025, Section 4.2, March 2025. DOI:10.6028/NIST.AI.100-2e2025

Integrating continuous monitoring and robust training strategies helps maintain model resilience in production. Regularly testing robustness against new data is fundamental to ensuring reliability and security in real-world scenarios.

Want to give your company the highest level of cyber security? ISGroup SRL is here to help with cyber security solutions tailored to your business.

Would you like us to take care of everything for you? Our Virtual CISO and vulnerability management services are a perfect fit for your organization.

Already know what you need? Explore our services:

And much more. Protect your company with the best cybersecurity experts!