From AI prototype to commercial SaaS: when a professional penetration test is needed
The critical transition is not when the demo works, but when the SaaS starts creating real accounts, separating tenants, processing payments, signing contracts, and integrating with external systems. At that point, the risk is no longer just technical: it becomes a commercial, legal, and reputational risk.
The goal is not to determine whether AI is a good or bad tool for development. The point is much more practical: understanding which controls are necessary when a result generated or accelerated by AI enters a product, a business workflow, or an environment containing real data.
This article is aimed at founders, CTOs, and SaaS builders. The focus is on the moment when users, payments, data, and contracts transform the prototype into a product and make a professional test necessary.
Why an app that works is not necessarily secure
AI tools reduce the time required to create code, interfaces, workflows, tests, and configurations. This speed, however, can compress steps that normally make software reliable: threat modeling, review, secret management, role controls, input validation, dependency verification, and manual testing of critical paths.
A demo can work perfectly with a single user, dummy data, and implicit permissions. The same logic can fail when real customers, multiple tenants, different roles, public APIs, integrations, personal data, payments, or automations with external effects arrive. This is why security must be evaluated based on the actual behavior of the app, not the promise of the tool that generated it.
Signs that the prototype has become a product
A penetration test is necessary when the app handles external users, personal data, distinct roles, payment plans, public APIs, webhooks, integrations with corporate systems, or an admin panel. Even a closed beta can be critical if it uses real data or if an enterprise client requires security evidence before purchasing.
What to test in an AI-generated SaaS
The scope must include authentication, object-level authorizations, tenant isolation, invitation and password reset flows, payments, APIs, uploads, exports, webhooks, support panels, and cloud configurations. AI-generated code must be analyzed with particular attention to the points where it decides who can see or modify what, because that is where the risks of unauthorized access are concentrated.
Main risks to check
- Incomplete tenant isolation between clients or workspaces: verify evidence, configuration, runtime behavior, and impact on real data.
- Broken access control on APIs not visible from the interface: verify evidence, configuration, runtime behavior, and impact on real data.
- Payment flows and webhooks not verified server-side: verify evidence, configuration, runtime behavior, and impact on real data.
- Exposed admin panels or support tools: verify evidence, configuration, runtime behavior, and impact on real data.
- Secrets in repositories, logs, prompts, or pipelines: verify evidence, configuration, runtime behavior, and impact on real data.
- Uploads and exports usable for data leakage: verify evidence, configuration, runtime behavior, and impact on real data.
- Permissive CORS, callbacks, and redirects: verify evidence, configuration, runtime behavior, and impact on real data.
These risks must be linked to the concrete perimeter of the application. An exposed app requires manual application testing; a critical code change requires review; an internal workflow requires checking permissions and credentials; an agentic app requires testing prompts, tools, and outputs. The correct combination depends on the actual impact, not the name of the tool used.
Report, remediation, and retest
A useful WAPT does not just produce a list of vulnerabilities. It must indicate impact, reproducibility, fix priority, evidence, residual risk, and retest. For a SaaS selling to business clients, the report also becomes proof of maturity for the client’s procurement and security review, and is often explicitly requested during enterprise evaluation phases.
Minimum controls before go-live
- Map users, roles, real data, integrations, environments, and service owners.
- Identify which parts were generated or modified with AI and who reviewed them.
- Verify server-side authorizations, tenant isolation, and administrative functions.
- Search for secrets in code, prompts, logs, environment variables, builds, and repository history.
- Check dependencies, licenses, packages, templates, plugins, and generated components.
- Test hostile inputs, error handling, logging, rate limits, and unexpected paths.
- Separate blocking fixes, planned remediation, and accepted residual risk.
- Repeat the test or retest after corrections that affect critical flows.
When an independent verification is needed
An independent verification is needed when the app or workflow handles real data, external users, roles, APIs, corporate integrations, payments, storage, automatic workflows, or critical code generated with AI. It is also needed when the team cannot demonstrate which parts have been reviewed and which controls block regressions or abuse.
For this type of scenario, the perimeter recommended by ISGroup includes Web Application Penetration Testing, Code Review, and the Vulnerability Management Service. The most useful verification is not generic: it must produce reproducible findings, remediation priorities, an indication of residual risk, and, when necessary, retests after corrections.
Operational questions for founders, CTOs, and security teams
- What real data enters the system and where is it saved, logged, or sent?
- Which roles exist and which actions are blocked server-side, not just in the interface?
- Which secrets, tokens, webhooks, or credentials would allow access to critical systems?
- Which parts were generated or modified by AI and which were reviewed by a competent person?
- Which tests cover abuse, errors, different roles, and different tenants, not just the happy path?
- What evidence can be shown to clients, audits, procurement, or management?
Useful resources
- Security controls before go-live: explores the controls to perform before bringing an AI app online without overlapping with the focus of this article.
- Secure code review for AI code: guide to reviewing code generated with AI, with attention to the specific risks of this type of development.
- Security audit for AI apps: overview of security auditing applied to applications developed or accelerated with AI tools.
FAQ
- When does an AI-created SaaS MVP need to perform a penetration test?
- When it leaves the demo perimeter: real users, personal data, payments, APIs, roles, tenants, or client contracts. Before public launch and before an enterprise sale is the most useful time to plan it.
- Is an automatic scan enough?
- No. Automatic scanners help identify known vulnerabilities, but they are unable to demonstrate tenant isolation, role abuse, business logic, correctness of payment flows, or object-level authorizations.
- Is a Code Review also needed?
- Yes, when the risk is in the application logic: authorizations, queries, secrets, payments, webhooks, integrations, and administrative functions are areas where manual code testing adds significant value compared to black-box testing alone.
- Does the test block the launch?
- Not if planned correctly. It serves to distinguish blocking fixes from corrections manageable before go-live and from remediations that can be addressed after launch in a controlled manner.
- What do enterprise clients ask for?
- They often require detailed reports, remediation evidence, retests, secure development policies, vulnerability management, and proof that the app has been verified by independent third parties.
Protect your organisation with Web Application Penetration Testing.
Choose ISGroup for a practical, tailored engagement:
- A focused assessment of your environment and requirements
- Clear findings with a prioritised, actionable roadmap
- Direct support from experienced specialists through remediation and implementation
