AI Coding and Security for Software Houses: Assurance Program

AI Coding e Sicurezza per Software House: Programma di Assurance

Software houses and AI coding: how to ensure security when code is generated by agents

For a software house, AI coding is not just a matter of internal productivity: it is a matter of reliability toward the client. If a team uses agents to generate code, refactoring, tests, pipelines, or configurations, they must be able to demonstrate that speed has not eliminated reviews, data segregation, security checks, and delivery accountability.

The question coming from clients, procurement, and audit teams is concrete: how do you know that code generated or modified with AI has been checked before delivery? A credible answer cannot be “we trust our developers” or “the scanners are green.” You need an assurance program: internal rules, evidence, reviews, tests, pipeline gates, reports, and remediation.

Why the risk is different for a software house

An internal team can accept a risk and govern it within their own development cycle. A software house, however, delivers code to someone else: clients, auditors, procurement, legal departments, and security teams can request proof at any time.

The use of AI coding adds new questions that go beyond technical quality. Which parts were generated or modified by agents? What client data was used in the prompt or context? Which dependencies were introduced, and who reviewed the diff? Which tests cover authorizations, roles, and abuse? Which findings were corrected before delivery and which residual risks were communicated? If these answers do not exist, the software house loses commercial control as well as technical control.

Internal policy for AI use in client projects

The policy must establish what is permitted in client projects, not just what is convenient for the team. It must cover authorized tools, corporate accounts, prohibited data, permitted repositories, agent modes, terminal access, cloud agents, external tools, reviews, and exception management. The minimum points to include are:

  • No personal accounts on client code or data.
  • No secrets, real logs, or client data in unauthorized prompts.
  • Separate workspaces per client.
  • Exclusion of sensitive files from the agent’s context.
  • Approval for tools that read client repositories.
  • Mandatory review on auth, APIs, data, pipelines, and dependencies.
  • Traceability of AI-generated parts when relevant to the contract.

The policy must be simple to apply: if it is too abstract, every team will interpret it differently, and controls will lose consistency across projects.

Client data segregation

The most delicate risk is not always a bug in the code, but the uncontrolled passage of information between clients, tools, and environments. A software house must separate repositories and branches, IDE workspaces and agents, cloud environments, credentials, test datasets, logs and tickets, client documentation, and prompts/transcripts when stored.

An agent working on multiple projects must not have unnecessary cross-project context. A client dataset must not become a fixture reused on other projects, and a production error must not be pasted into a generic chat with unredacted tokens, emails, or identifiers.

Traceability: from prompt to release

You don’t need to save every prompt indiscriminately, but you must know which decisions led to the release. Useful evidence to keep includes starting issues or tasks, PRs and commits, files modified by AI, reviewers and approvals, scanners and tests performed, new dependencies, CI/CD changes, IaC and deployments, findings and fixes, and accepted residual risks.

This traceability is also useful internally: when a vulnerability emerges, when a client asks for explanations, or when a regression reappears in a subsequent release, having a documented path significantly reduces response times.

Mandatory review of sensitive areas

Not all AI-generated code has the same risk profile. A copy change does not require the same level of control as a change to middleware, roles, queries, APIs, pipelines, or payments. A software house should define the areas that always require technical review:

  • Authentication, sessions, password reset, MFA, OAuth.
  • Authorizations, roles, tenant isolation, policies, and middleware.
  • APIs, queries, exports, uploads, payments, and webhooks.
  • Secrets, logs, error handling, and configurations.
  • Dependencies, lockfiles, licenses, Dockerfiles.
  • CI/CD, IaC, cloud, IAM, and deployments.
  • Runtime prompts, application agents, RAG, and tool calling.

Code Review thus becomes a control test: not just “we read the code,” but “we verified these areas, with this scope, and these results.”

Tests and scanners: necessary but not sufficient

SAST, SCA, secret scanning, unit tests, and integration tests are necessary, but they do not prove on their own that the application respects business logic, authorizations, tenant isolation, and role abuse. For client projects, the pipeline should produce at least:

  • SAST with triage of new findings.
  • SCA and license scanning on manifests and lockfiles.
  • Secret scanning on repositories, logs, and artifacts.
  • Negative tests on roles, tenants, inputs, and failure modes.
  • Controls on IaC, containers, and configurations.
  • Evidence of fixes and retests.

When the app or APIs are reachable by users, partners, or the internet, it is also necessary to verify runtime behavior. In this context, Web Application Penetration Testing complements the Code Review: one analyzes the cause in the code, the other verifies if the exposed behavior is exploitable on the real system.

Dependencies, licenses, and supply chain

Agents can add packages to quickly close a feature, but for a software house, a dependency is not just a technical risk: it can become a problem of licensing, maintenance, known vulnerabilities, client policy, or audit. Each new dependency should be evaluated regarding motivation, version and lockfile, compatible license, maintenance status, known vulnerabilities, installation scripts, and the presence of alternatives already in the project.

For enterprise clients, SBOM, provenance, and SCA reports can become contractual evidence. Even when not explicitly requested, knowing what enters the release reduces response times in case of CVEs.

Reporting to the client: what to show

The client should not receive a generic narrative about AI, but evidence proportionate to the project. A useful report can include the scope of the release or review, parts generated or modified with AI when relevant, controls performed with scanners and tests, corrected blocking findings, planned findings, residual risks, new dependencies, scope limitations, and recommendations for subsequent releases.

This does not mean exposing every prompt or every internal detail, but making it demonstrable that the delivered code was not accepted blindly.

From single project to assurance program

The real leap in quality comes when controls do not depend on the individual project manager. Every project with AI coding should have similar thresholds, similar reports, similar gates, and clear responsibilities. The Software Assurance Lifecycle serves exactly this purpose: transforming reviews, tests, pipelines, remediation, and reporting into a repeatable process across different teams and clients.

For a software house, this has direct commercial value: it helps respond to procurement, audits, security questionnaires, and requests for evidence without having to rebuild the process from scratch every time.

Checklist for software houses using AI coding

  • Internal policy on AI use in client projects.
  • Authorized tools, corporate accounts, and minimum configurations.
  • Separation of repositories, workspaces, data, prompts, and credentials per client.
  • Prohibition of entering client data, secrets, or sensitive logs into unauthorized tools.
  • Traceability of PRs, commits, reviews, scanners, tests, and findings.
  • Mandatory review on auth, APIs, data, pipelines, cloud, and dependencies.
  • SAST, SCA, secret scanning, and negative tests with defined thresholds.
  • License and new dependency control.
  • Verification of CI/CD, IaC, containers, deployments, and rollbacks.
  • Client report with scope, evidence, fixes, and residual risk.
  • Remediation managed with owner, severity, SLA, and retest.

When to involve ISGroup

A software house can start with a process assessment or a verification of a pilot project. The choice depends on the risk and the current level of maturity.

Scenario Main risk Recommended control
AI coding used across multiple teams or clients Non-repeatable controls and difficult audits Software Assurance Lifecycle
AI-generated PRs on sensitive code Vulnerabilities or regressions in code Code Review
Application or API exposed to users/clients Runtime abuse and application vulnerabilities Web Application Penetration Testing
Client requires evidence before delivery Non-demonstrable scope and residual risk Assurance program and technical report

The value is not just finding bugs: it is making the delivery process defensible, documenting what was checked, by whom, with what evidence, and with what declared limitations.

Frequently asked questions

  • Must a software house always declare the use of AI coding?
  • It depends on the contract, client requirements, and internal policies. In practice, it is better to be prepared to explain which tools are used, which data is excluded, and which controls verify the result.
  • Does the client need to see the prompts?
  • Not necessarily. Prompts can contain sensitive information. Often, policies, scope, review evidence, tests, scanners, corrected findings, and residual risk are more useful.
  • How to demonstrate that AI-generated code has been checked?
  • With concrete evidence: PRs, commits, reviews, tests, SAST/SCA reports, secret scanning, dependency reports, Code Review, WAPT, findings, fixes, and retests.
  • When is an independent Code Review needed?
  • When code generated or modified with AI touches auth, roles, APIs, data, queries, secrets, dependencies, pipelines, cloud, or critical business logic.
  • When is WAPT needed in addition to Code Review?
  • When the app is reachable by users or clients. Code Review analyzes the source code; WAPT verifies if the exposed behavior is exploitable under real conditions.

Protect your organisation with Software Assurance Lifecycle.

Choose ISGroup for a practical, tailored engagement:

  • A focused assessment of your environment and requirements
  • Clear findings with a prioritised, actionable roadmap
  • Direct support from experienced specialists through remediation and implementation
Talk to an expert

Sources and references