Security audits for AI-generated applications: what is included and what is not
Those requesting a security audit for an AI-generated app often face two concrete questions: what will actually be checked and what remains outside the scope. The answer depends on the type of app: a classic web app generated with AI requires WAPT, Code Review, and configuration verification; an app that integrates LLMs also requires testing on prompts, RAG, tool calls, and outputs.
The point is not to judge AI as a development tool. It is much more practical: understanding which controls are needed when code or workflows generated or accelerated by AI enter a product, a business process, or an environment with real data.
Why an app that works is not necessarily secure
AI tools reduce the time needed to create code, interfaces, workflows, and configurations. However, this speed can compress steps that normally make software reliable: threat modeling, review, secret management, role controls, input validation, dependency verification, and manual testing of critical paths.
A demo works with a single user, dummy data, and implicit permissions. The same logic can fail when real customers, multiple tenants, different roles, public APIs, integrations, personal data, payments, or automations with external effects arrive. Security must be evaluated based on the actual behavior of the app, not on the promise of the tool that generated it.
What is needed to start an audit
To conduct a serious audit, you need: repository or build, URLs and test environments, definition of user roles, critical flows, data schema, external integrations, CI/CD pipelines, relevant cloud configurations, and an indication of which parts were generated or modified with AI.
What a proportionate audit includes
A proportionate audit covers authentication, authorization, APIs, data management, secrets, dependencies, configurations, logging, error handling, pipelines, and exposed surfaces. If the app uses LLMs, it also includes prompt injection, output handling, retrieval, memory, tool calls, and rate limits.
What is not included
An audit is not a certification of the AI tool used for development, it does not guarantee the absence of future bugs, and it does not replace governance, monitoring, patching, and continuous remediation. Without access to code, roles, and an environment representative of production, the audit inevitably becomes partial.
Main risks to check
- Ambiguous perimeter between WAPT, Code Review, VA, and AI testing: define in advance which surfaces are included in the test and which are not.
- Test environment not representative of production: dummy data and simplified permissions hide real vulnerabilities.
- Roles and data not provided to the tester: without access to real roles, authorization checks remain incomplete.
- Generated code not tracked in diffs: parts generated with AI that are not manually reviewed are a blind spot.
- LLM risks ignored for apps using models: prompt injection, unchecked tool calls, and unvalidated outputs are concrete vectors.
- Classic risks ignored because the app is “AI”: authentication, authorization, and secret management remain critical regardless of how the code was written.
- Deliverables without remediation priorities: a report without a distinction between blocking findings and residual risk does not support operational decisions.
The correct combination of controls depends on the impact, not the name of the tool. An exposed app requires manual application testing; a critical code change requires review; an internal workflow requires permission and credential control; an agentic app requires testing on prompts, tools, and outputs.
Minimum controls before go-live
- Map users, roles, real data, integrations, environments, and service owners.
- Identify which parts were generated or modified with AI and who reviewed them.
- Verify server-side authorizations, tenant isolation, and administrative functions.
- Search for secrets in code, prompts, logs, environment variables, builds, and repository history.
- Check dependencies, licenses, packages, templates, plugins, and generated components.
- Test hostile inputs, error handling, logging, rate limits, and unexpected paths.
- Separate blocking fixes, planned remediation, and consciously accepted residual risk.
- Repeat the test or retest after corrections that affect critical flows.
When an independent verification is needed
An independent verification is needed when the app or workflow handles real data, external users, roles, APIs, business integrations, payments, storage, automatic workflows, or critical code generated with AI. It is also needed when the team cannot demonstrate which parts have been reviewed and which controls block regressions or abuse.
The perimeter recommended by ISGroup in this case includes: Web Application Penetration Testing, Code Review, Vulnerability Assessment and, for apps with LLMs, AI Application Testing. The best review produces reproducible findings, remediation priorities, an indication of residual risk, and, when necessary, retests after corrections.
Operational questions for founders, CTOs, and security teams
- What real data enters the system and where is it saved, logged, or sent?
- What roles exist and what actions are blocked server-side, not just in the interface?
- What secrets, tokens, webhooks, or credentials would allow access to critical systems?
- What parts were generated or modified by AI and which were reviewed by a competent person?
- What tests cover abuse, errors, different roles, and different tenants, not just the happy path?
- What evidence can be shown to clients, audits, procurement, or management?
Useful resources
- Penetration test for SaaS and AI apps: how to set up a test on SaaS products or applications that integrate AI components, with a focus on perimeter and methodology.
- Security controls for AI apps before go-live: operational list of controls to complete before taking an app developed with AI online.
- AI Application Testing: in-depth look at specific tests for applications that integrate LLMs, agents, and autonomous workflows.
FAQ
- What is the difference between an audit, WAPT, and Code Review?
- WAPT verifies the behavior of the exposed app by simulating an external attacker. Code Review analyzes the source code for logical vulnerabilities and bad practices. The audit combines controls proportionate to the perimeter and can include configurations, processes, and AI-specific risks.
- When is AI Application Testing needed?
- When the app integrates LLMs, RAG, agents, tool calling, memory, or autonomous workflows. If AI was used only to write code, usually the priority controls remain WAPT and Code Review.
- What materials do I need to prepare before the audit?
- URLs, role definitions, critical flows, repository or build, architecture, integrations, test data, CI/CD pipelines, and a list of parts generated or modified with AI.
- How broad should the perimeter be?
- Broad enough to cover what can cause real impact: data, users, roles, APIs, administrative functions, storage, payments, integrations, and deployments.
- Is the report enough to go online?
- Only if blocking findings have been corrected or consciously accepted. The final decision must include an assessment of the completed remediation and the residual risk.
Sources and references
- OWASP Top 10 2021
- OWASP ASVS
- OWASP Code Review Guide
- NIST SP 800-218 SSDF
- OWASP Top 10 for LLM Applications 2025
If you are about to take an app or workflow developed with AI online, ISGroup can help you choose the right control: application testing, Code Review, architectural assessment, or targeted verification of AI-specific risks.
Protect your organisation with Web Application Penetration Testing.
Choose ISGroup for a practical, tailored engagement:
- A focused assessment of your environment and requirements
- Clear findings with a prioritised, actionable roadmap
- Direct support from experienced specialists through remediation and implementation
