AITG-INF-06: Testing for Dev-Time Model Theft

Dev-Time Model Theft protezione e strategie di remediation

During the development of AI models, intellectual property is exposed to concrete risks of theft. Proprietary models, training datasets, and strategic components can be stolen before they even reach production due to insecure environments, insufficient access controls, and unprotected storage practices. Dev-Time Model Theft represents a critical threat to organizations investing in artificial intelligence.

This article is part of the AI Infrastructure Testing chapter of the OWASP AI Testing Guide.

How model theft occurs during development

Attackers exploit three main areas to steal models during development phases:

  • Unauthorized access: theft through compromised credentials or excessive permissions in development and training environments
  • Weak access controls: insufficient isolation between development, testing, and production environments that facilitates lateral movement
  • Insecure storage and transfer: model artifacts and datasets stored without encryption or adequate protections during training phases

Methodology and payloads

Hardcoded credential scanning

The first vector to verify concerns credentials embedded in the source code. Automated scanning tools identify API keys, passwords, and access tokens in the project’s Git repositories.

Vulnerability indicator: valid credentials allow read access to storage containing models or training datasets, exposing the organization’s entire intellectual property.

Exfiltration via CI/CD pipelines

Continuous integration and deployment pipelines represent a prime target. The test verifies whether users with a developer role can modify the pipeline by adding steps that exfiltrate artifacts to external servers.

Vulnerability indicator: the pipeline allows unauthorized modifications without generating security alerts or applying outbound network policies, enabling the transfer of proprietary models externally.

Model extraction via insecure development APIs

Development environments often expose internal or staging APIs used for debugging and evaluating models. These interfaces, if accessible from external environments or lacking adequate authentication, allow for the direct download of model files and parameters.

Vulnerability indicator: exposed development APIs allow the download of proprietary artifacts without authentication or with easily obtainable credentials.

Expected output

A properly secured AI development infrastructure presents these verifiable characteristics:

  • Absence of secrets in code: no hardcoded credentials or API keys in source code repositories
  • Fortified CI/CD pipeline: modifications subject to security review, sandboxed runners, and strict control of outbound traffic
  • Granular access to models: model files and training datasets accessible exclusively to authorized services and personnel, with complete logging of all operations

Remediation actions

Access control and secret management

Implementing strict RBAC (Role-Based Access Control) on all development resources constitutes the first layer of defense. Passwords and API keys must be stored exclusively in secure vaults with periodic rotation and a complete audit trail.

Expected impact: elimination of hardcoded credentials and reduction of the risk of unauthorized access to model artifacts.

CI/CD pipeline security

Reinforcing pipeline security requires branch protection rules, mandatory reviews for critical changes, and the use of sandboxed runners with restrictive egress rules. Every change to the pipeline must be tracked and approved.

Expected impact: prevention of exfiltration through compromised pipelines and complete traceability of build infrastructure changes.

Digital signature and artifact integrity

All model artifacts must be digitally signed during the build process. Deployment verifies the signature to ensure integrity and the absence of tampering throughout the entire distribution chain.

Expected impact: guarantee of artifact integrity and immediate detection of attempts to replace or tamper with models.

Encrypted storage and audit logging

Model artifacts must be stored encrypted in private repositories with granular access control. Continuous audit logging records every access and operation on models, allowing for the timely detection of anomalous activity.

Expected impact: protection of the confidentiality of proprietary assets and complete visibility into model access.

Monitoring and Data Loss Prevention

Implement monitoring systems that control all access to models and automatically block unauthorized file transfer attempts. DLP policies must be configured to detect and prevent the exfiltration of proprietary assets.

Expected impact: real-time detection and blocking of exfiltration attempts, with a significant reduction in the risk of intellectual property theft.

Suggested tools

  • TruffleHog: automated scanning of exposed credentials in Git repositories
  • Gitleaks: detection of hardcoded secrets in source code
  • git-secrets: prevention of committing AWS credentials and other secrets
  • HashiCorp Vault: centralized management of secrets and credentials

Frequently Asked Questions

  • What are the signs of potential model theft in progress?
  • Anomalous access to model repositories, downloading large volumes of data by unauthorized users, unapproved changes to CI/CD pipelines, and connection attempts to unexpected external servers are all indicators of potential exfiltration in progress.
  • How can I verify if credentials are exposed in the code?
  • Use automated scanning tools like git-secrets, TruffleHog, or Gitleaks to identify hardcoded credentials in repositories. These tools analyze Git history and detect patterns of exposed API keys, passwords, and tokens.
  • What controls should be implemented on CI/CD pipelines?
  • Branch protection rules, mandatory review for pipeline changes, sandboxed runners with restrictive egress rules, complete logging of all operations, and automatic alerts for unauthorized changes are essential controls.
  • How can training datasets be protected from theft?
  • Store datasets encrypted in storage with granular access control, implement audit logging to track every access, use DLP policies to prevent exfiltration, and limit access to authorized users and services only.
  • What is the difference between Dev-Time Model Theft and Runtime Model Theft?
  • Dev-Time Model Theft occurs during development and training phases, when models are still being worked on. Runtime Model Theft occurs when the model is already in production and is extracted through repeated queries or direct access to the inference infrastructure.

Useful resources

To better understand the AI infrastructure security context and related threats:

How ISGroup supports you

ISGroup supports organizations in protecting AI assets throughout the entire development cycle through the Secure Architecture Review service. The team evaluates the development architecture, identifies vulnerabilities in model management processes, and provides concrete recommendations for implementing effective security controls.

For source code security verification and the identification of exposed credentials, ISGroup offers the Code Review service, which analyzes code to identify bad practices and vulnerabilities before they reach production.

References

  • OWASP GenAI โ€“ Generative AI Security
  • MITRE ATT&CK โ€“ Data Staged: Model Theft
  • NIST AI Security Guidelines โ€“ Protecting AI Artifacts and Intellectual Property

The integration of strict access controls, secure secret management, and continuous monitoring helps protect AI intellectual property during development. Regularly testing the security of CI/CD pipelines and development environments is fundamental to ensuring the protection of proprietary assets in production.

Want to give your company the highest level of cyber security? ISGroup SRL is here to help with cyber security solutions tailored to your business.

Would you like us to take care of everything for you? Our Virtual CISO and vulnerability management services are a perfect fit for your organization.

Already know what you need? Explore our services:

And much more. Protect your company with the best cybersecurity experts!