Threat modeling for AI code: from prompt to pull request

Threat modeling per codice AI da prompt a pull request

From prompt to pull request: threat modeling for AI-generated code

A coding agent can transform an issue into a credible pull request: it reads the repository, proposes a plan, modifies files, updates tests, and prepares a diff ready for review. The concrete risk is that the team evaluates only whether the feature works, without asking whether the initial prompt truly described data, roles, potential abuses, trust boundaries, and production impact.

Threat modeling is essential at this stage: not as a bureaucratic exercise, but as a lightweight method to understand what is changing, what can go wrong, what controls are needed, and whether the verification performed is sufficient before merging, deploying, or going live.

For CTOs, tech leads, and security champions, the question is not “can we use AI agents to write code?”. The useful question is: how do we transform prompts, agent plans, and pull requests into a traceable risk decision?

Why threat modeling changes with AI-generated code

In traditional application threat modeling, you start with features, assets, actors, data flows, and trust boundaries. With code written by AI agents, you must add a layer: the path that leads from the prompt to the diff.

A request like “add team invites” can generate data models, APIs, emails, temporary tokens, roles, admin pages, tests, environment variables, and a migration. Each piece introduces different threats: reusable invites, tokens that are too long in logs, roles applied only in the frontend, unprotected endpoints, incomplete tenant isolation, email spoofing, or dependencies added without review. The prompt rarely contains all of this, because the agent fills the gaps with plausible conventions. Threat modeling serves to make those gaps explicit before they become accepted code.

The four questions to apply to every AI-generated PR

OWASP summarizes threat modeling in four questions: what are we building, what can go wrong, what will we do about it, and have we done enough? In the AI coding context, these questions must be applied to prompts, plans, diffs, and runtime.

Question Meaning in an AI agent workflow
What are we building? Which task did the agent receive, which files did it modify, and which data and systems does it touch?
What can go wrong? Which abuses become possible regarding roles, data, APIs, tools, pipelines, and exposed surfaces?
What are we doing about it? Which controls are implemented in code, tests, configuration, reviews, pipelines, or processes?
Have we done enough? What evidence blocks the merge, which risks are accepted, and who governs them?

This structure avoids two frequent errors: treating threat modeling as an abstract meeting or reducing it to a generic vulnerability checklist.

Start with real assets, data, and actors

The first step is not to read the diff line by line, but to understand which assets enter the scope of the change: personal data, customer documents, payments, administrative roles, tokens, internal APIs, webhooks, repositories, pipelines, logs, storage, databases, and cloud services.

Then you need the actors. It is not enough to distinguish between “user” and “admin”: in a real feature, there may be invited users, suspended users, owners, partial admins, service accounts, external integrations, members of another tenant, deactivated customers, support operators, scheduled jobs, and anonymous attackers. A minimal actor-action-resource matrix helps much more than a generic review. For each role, you must ask: what resources can it read, create, modify, delete, or export? Does the backend really verify this, or did the agent implement the control only in the interface?

Map the data flow, even if lightly

A threat model does not necessarily have to produce a perfect diagram, but it must make data flows, data stores, processes, external entities, and trust boundaries visible. If a PR crosses the frontend, API routes, database, object storage, email provider, webhooks, and deployment pipelines, the risk is not in the same place for all components.

With AI-generated code, the data flow must also include elements often forgotten: agent prompts, chat history, tool outputs, build logs, environment variables, secret managers, test fixtures, seeds, snapshots, and deployment artifacts. A secret copied into a prompt or a log does not always appear in the final diff, but it may still require rotation.

The most important boundary is where an untrusted input enters a trusted component. A value coming from a form, ticket, document, webhook, uploaded file, LLM output, or external tool must be validated before becoming a query, command, authorization, email, decision, or downstream action.

Write abuse cases, not just user stories

User stories describe the intended use; abuse cases describe what a user, tenant, integration, or malicious content can do when the system is used out of sequence. Some useful examples for an AI-generated PR:

  • A user changes the ID in the URL and reads a record belonging to another tenant.
  • An expired invitation is reused to gain access.
  • A webhook is replayed or signed with the wrong key.
  • A read-only role calls the modification API directly.
  • An uploaded file contains a payload that ends up in HTML, a parser, or storage.
  • A prompt injection in a ticket induces the agent to modify a policy or use a tool.
  • A migration exposes test data or sensitive fields in production.

STRIDE remains useful as a stimulus—spoofing, tampering, repudiation, information disclosure, denial of service, elevation of privilege—but in AI-generated code, it must be linked to concrete cases, not used as a decorative acronym.

Evaluate the agent’s plan before the diff

Many teams start the review when the pull request is ready. With AI agents, it is better to anticipate: if the agent proposes a plan that touches auth, roles, data, cloud, CI/CD, secrets, or dependencies, the threat model must start before the diff becomes large.

The agent’s plan should be read with precise questions: is it creating new endpoints or exposing previously non-existent routes? Is it changing middleware, policies, permission checks, or roles? Is it adding dependencies, scripts, workflows, or deployment commands? Is it modifying tests to confirm its own implementation? Is it assuming that a client-side check is enough? Is it using real data, logs, or secrets to complete the task? If the answer is yes to even one of these questions, the PR should be broken down or marked as high-risk, because an effective review cannot treat a copy correction and a change that crosses authorizations, databases, and pipelines in the same way.

Connect threats, controls, and tests

A useful threat model produces verifiable controls. If the threat is “user from tenant A reads data from tenant B”, the control is not “manage authorizations”: it is a server-side policy, a tenant-filtered query, tests with two distinct tenants, access-denied logging, and middleware review.

If the threat is “reusable invitation token”, the controls can include expiration, single-use, binding to email or tenant, hashing the token at rest, rate limiting, and audit logs. If the threat is “prompt injection via ticket”, the controls are separation between data and instructions, human confirmation for sensitive tools, allowlists of actions, and tests with hostile content.

Scanners and automatic tests help, but they do not prove on their own that business logic, tenant isolation, authorizations, and trust boundaries are correct in the real context.

Define what blocks merge and go-live

Threat modeling should not end with “we talked about the risks”: it must produce decisions. Evidence that exposes data, permits privilege escalation, bypasses authorizations, reveals secrets, makes storage or databases public, weakens pipelines, disables security controls, or introduces unvetted critical dependencies must block the merge or go-live.

Other points can become planned remediation, but only with clear owners, dates, and residual risk. Improving logging, adding alerts, strengthening documentation, or making the role matrix more granular can be planned; accepting uncertain access control before go-live cannot.

Threat modeling checklist for AI-generated PRs

  • Identify the original prompt/task, the agent or tool used, and the parts generated or modified.
  • List assets, real data, roles, tenants, external systems, and secrets touched.
  • Draw a minimal data flow with trust boundaries, processes, data stores, and external entities.
  • Write abuse cases specific to the feature, not just generic vulnerabilities.
  • Verify object-level authorizations and tenant isolation with different users.
  • Check if the PR modifies middleware, policies, roles, callbacks, redirects, CORS, or sessions.
  • Separate code review, pipeline review, dependency review, and test review.
  • Look for secrets in prompts, logs, repositories, builds, workflows, and configurations.
  • Demand negative tests derived from threats, not just happy-path tests.
  • Document what blocks the merge, what blocks the go-live, and what remains as an accepted risk.

How to integrate the method into the development cycle

Threat modeling for AI coding must be lightweight and repeatable. A long session is not needed for every commit, but clear thresholds are: a PR that touches copy or layout can follow the normal flow, while a PR that touches auth, data, APIs, payments, cloud, CI/CD, secrets, roles, or agentic tools must trigger a more rigorous check.

In the practical process, the task should include security constraints; the agent’s plan should be approved when it touches sensitive areas; the PR should be small; tests should cover abuse cases; the review should have technical owners; the final decision should leave a trace. When the use of AI agents becomes continuous, this becomes a topic of Software Assurance Lifecycle: development rules, branch protection, security gates, ownership, evidence, remediation, and recurring checks.

When to involve an independent verification

An internal verification may be enough for isolated changes, without real data, without roles, without exposed APIs, and with expert reviewers. An independent verification is needed when the code generated or modified by AI enters a product used by customers, employees, or partners, or when the PR touches authorizations, data, pipelines, cloud, payments, secrets, or critical logic.

Scenario Main risk Recommended control
AI-generated PR on auth, roles, APIs, queries, secrets, or dependencies Vulnerabilities or regressions in the code Code Review
Continuous adoption of coding agents in the engineering cycle Non-repeatable checks on tasks, PRs, and releases Software Assurance Lifecycle
Decision on go-live, residual risk, business impacts, or compliance Risk not prioritized or accepted without an owner Risk Assessment
Architecture, data flows, trust boundaries, and complex integrations Weak design assumptions Secure Architecture Review
Apps or APIs already exposed to users and customers Abusable behavior from the outside Web Application Penetration Testing

The choice should not be a standard package. If the risk is in the diff, you need a Code Review. If the risk is in the agent adoption process, you need a Software Assurance Lifecycle. If the team must decide on residual risk, priorities, and impact, you need a Risk Assessment. If the change crosses architecture, data flows, and integrations, you need a design review.

Evidence to prepare

To make verification effective, you need the repository, pull request, task description, starting prompt or issue, list of agents or tools used, parts generated or modified, application roles, data schema, exposed APIs, integrations, CI/CD workflows, environment variables, main dependencies, and available environments.

Process evidence is also useful: who approved the agent’s plan, who reviewed the diff, which tests were added, which threats were considered, which risks were accepted, and which remediations were planned. This information avoids a blind review, because the code does not always tell why a decision was made, what assumptions the agent made, or what risk the team thought it was accepting.

Frequently asked questions

  • Is threat modeling also needed for a single AI-generated PR?
  • Yes, if the PR touches data, roles, APIs, secrets, payments, pipelines, cloud, or business logic. For small changes, a lightweight threat model is enough, but the questions about assets, actors, flows, and abuse remain useful.
  • Are threat modeling and Code Review the same thing?
  • No. Threat modeling defines what can go wrong and what controls are needed. Code Review verifies whether the code, diff, and tests implement those controls without regressions.
  • Is STRIDE enough for AI-generated code?
  • STRIDE is useful for not forgetting families of threats, but it must be adapted to prompts, agents, repositories, pipelines, tool calls, secrets, and runtime data. The important part is transforming every threat into a verifiable control.
  • Are automatic tests generated by AI enough?
  • No. AI-generated tests tend to cover the intended path. You also need negative tests, multi-user cases, role abuse, tenant isolation, manipulated inputs, and checks on pipelines and configurations.
  • When is a Risk Assessment needed?
  • When the team must decide whether to accept a risk, prioritize remediation, or link the project to business impacts, personal data, compliance, vendors, or operational continuity.

Protect your organisation with Software Assurance Lifecycle.

Choose ISGroup for a practical, tailored engagement:

  • A focused assessment of your environment and requirements
  • Clear findings with a prioritised, actionable roadmap
  • Direct support from experienced specialists through remediation and implementation
Talk to an expert

Sources and references