Code written with ChatGPT: how to verify it before going into production

Codice generato con ChatGPT Γ¨ sicuro in produzione

Code written with ChatGPT: is it safe before using it in production?

Those who arrive at this question are usually no longer experimenting: they already have a prototype, a function, a web app, or a workflow built or modified with ChatGPT that seems to work. The critical point is to understand whether that code can be connected to real data, users, payments, corporate APIs, or production environments without introducing unseen risks.

Using ChatGPT to write code has become the norm in many teams: a login function, an API route, a SQL query, a React component, a migration script, or a CORS configuration. The risky step is not the request made to the AI, but pasting a plausible response into a real application without verifying whether that code respects the product’s context: users, roles, data, backend, libraries, deployment, logs, secrets, and integrations.

The operational question is not whether ChatGPT is “safe” as a vendor in the abstract. It is another one: is the code written or corrected with ChatGPT safe within your application, with your data, your roles, your APIs, and your production environment?

This article is a strategic guide for developers, founders, and IT teams who want to transform AI speed into reliable releases, preventing “vibe coding” from turning into a technical or reputational crisis at the first go-live.

πŸ”΄ Web Application Penetration Testing: identify hidden risks and strengthen your security with a focused assessment by ISGroup specialists.

When the chat’s “it works” becomes a real risk

The delicate moment arises when the prototype changes state. With AI coding tools, it is easy to quickly obtain forms, APIs, and dashboards, but this speed can hide a misunderstanding: a feature that returns the expected output is not automatically secure. In general-purpose chatbots like ChatGPT, the risk often arises from copy-pasting isolated snippets without reviewing the global application context.

The risk grows significantly in the transition from local demo to online-accessible service, from dummy data to personal data subject to GDPR, from private repository to shared deployment pipeline, from single test user to different roles and privileges, and finally from throwaway code to a core function used by customers or partners. In each of these steps, security does not coincide with the first correct output: a function may work for the intended user and remain vulnerable to a malicious user, a different tenant, manipulated input, or an overly permissive deployment configuration.

From prompt to repository: where the most insidious flaws are born

ChatGPT often works as an isolated assistant: the user describes a problem, receives code, and decides where to insert it. This human mediation is powerful, but it is also the point where the most difficult-to-detect security flaws are born. The snippet generated by the AI lives in an information vacuum: it does not know that the app uses a centralized authorization middleware, that IDs must be derived from the protected session and not the client, or that there is a multi-tenant policy to be strictly respected.

The typical risk is not visibly broken code, but reasonable, readable, and functional code that forgets an essential check. A classic example is an endpoint that reads an order by ID: the AI might generate a syntactically perfect query that, however, does not verify whether that order actually belongs to the user requesting it. Without this check, the app is functional but exposed to an IDOR/BOLA vulnerability.

Specific risks of chat-generated code

Isolated snippets and absence of a threat model

Code generated via chat does not know the overall architecture or the specific attack vectors of your infrastructure. Without a threat model, the AI favors the simplest and most direct solution, which is often also the least protected.

Absent or superficial input validation

ChatGPT tends to write code optimized for the “happy path.” The suggested validation is often limited to the frontend β€” a mandatory email field, a basic check β€” while on the server side, the app remains open to injection (SQL, NoSQL, Path) or parameter manipulation via direct API calls that bypass the UI.

Authentication, sessions, and role management

Authentication snippets requested from the AI may contain obsolete or simplified patterns: JWTs without signature verification, sessions with infinite expiration, or authorization logic based on parameters passed by the client that can be easily altered, such as a role=admin in the browser payload.

Secrets, tokens, and credentials in prompts

The risk of exposing API keys, tokens, or credentials during a chat session is real. Without realizing it, one might paste logs or sensitive configurations to resolve a bug, sending data to OpenAI that ends up in the shared history or, if not configured correctly, in the model training data.

Obsolete or hallucinated suggested dependencies

The AI may suggest libraries that are no longer maintained or, in rare cases, non-existent package names. An attacker could exploit this hallucination by registering a malicious package with that name β€” an attack known as “dependency confusion” β€” hitting anyone who uncritically copies the suggested snippet.

The risk of context leakage: when corporate data ends up in the prompt

Beyond the security of the generated code, there is a mirror risk: the security of the data sent to ChatGPT. In an effort to solve a complex bug, a developer might paste entire configuration files with secrets not yet removed, application logs containing user emails or real session tokens, database schemas that reveal the company’s multi-tenant logic, or code snippets protected by intellectual property. Without the right precautions, this information becomes part of the history and potentially the training dataset, so the security of applications created with AI starts with the prompt usage policy.

Intellectual property protection and compliance

When ChatGPT enters the corporate workflow, it is essential to configure the environment correctly to avoid context leaks. Enterprise and Team plans are the only professional choice for companies: they guarantee that the data entered is never used for model training and offer an administration console to manage who can access AI tools. The “Temporary Chat” feature is useful for ad-hoc debugging sessions where you do not want to leave a trace of the conversation on OpenAI servers. For those using personal plans, it is mandatory to explicitly disable “Chat History & Training” to prevent data persistence.

Where AI speed can betray the product

The most dangerous vulnerabilities introduced by ChatGPT are not syntax errors, but plausible logical regressions. Here is where to pay maximum attention before go-live.

Authorization and isolation (IDOR/BOLA)

ChatGPT tends to generate code focused on making the request work. If you ask to write an endpoint to retrieve user data, the AI will likely write a query filtered by id, but it might forget to verify if the user in the session has the right to access that specific id. Before deploying, try changing an ID in the URL or API request payload β€” for example, from /api/user/101 to /api/user/102. If you can see or modify another user’s data without being them, the code has omitted a fundamental ownership check. This type of vulnerability, known as Insecure Direct Object Reference (IDOR), is among the most common and devastating in AI-created apps.

Secret management (hardcoded secrets)

During the generation of ready-to-use examples, ChatGPT often inserts API keys, database passwords, or test tokens directly into the code. The developer, in the rush to test the snippet, might forget to move these values to a secret manager. Look for strings that resemble private keys or tokens in the repository and move them immediately to protected environment variables: once a secret has ended up in a Git commit, it must be considered compromised and rotated immediately.

Business logic and misleading tests

AI-generated tests are often designed to confirm that the code works for the main use case (happy path), not to challenge it. An AI-generated test that passes is not proof of security: it is only proof of functionality. A test might confirm that a user can upload a file, but not verify if the same user can upload an executable file or a malicious script that compromises the server.

The problem of shadow IT: ChatGPT as a shadow developer

In companies, the risk does not only concern those who develop official applications, but also those who use ChatGPT to create small internal tools, automation scripts, or quick workflows. This phenomenon, known as AI-generated shadow IT, brings tools into production that escape the control of security teams. An internal tool created in an afternoon can save sensitive information in local databases or unprotected files, make it impossible to reconstruct who did what in the event of an incident, use unapproved and potentially vulnerable external libraries, and remain active as an orphaned service without a technical owner to ensure its maintenance.

Defining a corporate AI policy

To mitigate this risk, the company must provide clear guidelines that cover at least four areas: a whitelist of tools that specifies which versions of ChatGPT are authorized (e.g., only Enterprise), minimum validation criteria for every script that touches corporate data, rules on secret management with an absolute prohibition on inserting real keys into prompts or generated code, and clear ownership so that every AI-generated tool must have an identified human responsible.

Minimum checks before go-live

  • Every snippet that touches sensitive data passes through a centralized authorization middleware (trust boundary mapping).
  • Endpoints verify the user’s identity and permissions on the server for every single request (API route audit).
  • Database queries use parameters and not string concatenations suggested by the AI (input sanitization).
  • Every new library added to the project has been verified for reputation and date of the last release (dependency review).
  • An automatic scan (e.g., with truffleHog or git-secrets) has been performed to ensure no secrets have ended up in the repository (secret scanning).
  • Tests have been performed to try to bypass the login or access administrative functions without permissions (negative testing).
  • Errors returned to the client are generic and do not expose technical details such as stack traces or SQL queries (error handling).

Operational scenario: the refactoring that breaks authorizations

Imagine a team asking ChatGPT to refactor an authentication middleware to support roles. The AI produces clean, modern, and performant code. The developer integrates it, the app compiles, and functional tests are green. However, during the refactoring, the AI unintentionally omitted a check on a specific administrative route or introduced a permissive fallback logic β€” for example, if the role is not recognized, the user is treated as a guest but with access to certain records. This type of error is invisible to the naked eye in a diff of hundreds of lines, but it is a vulnerability ready to emerge at the first go-live.

Security cannot be delegated to the chat: it must be verified on the finished product.

When an independent verification is needed

The chat’s “it works” is not evidence of security. Professional verification is needed to bridge the gap between a plausible snippet and a production-ready function, especially when the app handles real data, payments, or critical processes.

If the risk concerns… The typical problem is… Recommended ISGroup service
Source, logic, auth, API Broken authorizations, secrets, dependencies Code Review
Exposed web app, sessions, input Abuse from outside, injection, BOLA Web Application Penetration Testing
Architecture, data flows, integrations Weak security assumptions Secure Architecture Review
Cloud, deploy, IAM, storage Misconfiguration or excessive privileges Cloud Security Assessment
Continuous use of AI in teams Lack of process and governance Software Assurance Lifecycle

Business impact and ROI of security in vibe coding

For a founder or a CTO, investing in a security review for AI-generated code is not just a technical precaution, but a strategic financial decision. The cost of a vulnerability discovered after go-live is estimated to be up to 30 times higher than that of an intervention during the development phase. A data leak or privilege escalation on a newly launched app can destroy a new brand’s reputation in a few hours.

Verifying AI code before publication allows you to accelerate time-to-market securely β€” using AI to write most of the code, but maintaining professional control over security β€” to avoid emergency remediation costs by blocking authorization bugs when they are still drafts in a chat, and to comply with GDPR by ensuring that user data is isolated and protected from day one.

Evidence to prepare before the professional review

To maximize the effectiveness of an external review β€” Code Review or WAPT β€” the development team should prepare some key information in advance:

  • List of AI perimeters: which modules have been generated or massively refactored with ChatGPT.
  • Sensitive data map: what information (GDPR, financial, industrial secrets) the app handles and where it is stored.
  • Access to test environments: staging URLs, credentials for different user roles (admin, user, guest), and API documentation.
  • Cloud and CI/CD configurations: deployment files, IAM policies, and automation scripts touched by the agent or suggested by the chat.

The final question is simple: has the code written with ChatGPT been verified as a corporate product, or has it just been accepted as a working diff? Security serves to ensure that the speed gained with AI is not lost in emergency remediation after an incident.

FAQ

  • Can ChatGPT write secure code?
  • Yes, but only if guided by prompts that include rigorous security constraints β€” for example, “use parameterized queries” or “implement server-side RBAC checks” β€” and if the result is integrated with competence into the context of the corporate architecture. AI tends to optimize for immediate functionality, not for defense-in-depth.
  • How can I prevent OpenAI from using my code for training?
  • The safest solution is to adopt ChatGPT Team or Enterprise plans, which exclude user data from training by default. For personal accounts, it is possible to disable “Chat History & Training” in the settings or use the “Temporary Chat” feature for sensitive sessions.
  • Is an automatic scan (SAST) sufficient for AI-generated code?
  • No. Automatic tools like SonarQube or Snyk are excellent for finding known vulnerability patterns (syntactic SQLi or XSS), but they rarely catch complex authorization logic errors or architectural hallucinations where the AI assumes a component is protected when it is actually exposed.
  • What is the risk of pasting logs or configurations into ChatGPT?
  • The main risk is context leakage: sensitive data such as API keys, real session tokens, or user emails can be saved in the AI vendor’s history. If an account is compromised or if the data is used for training, this information could become accessible to third parties.
  • What should I do if I discover a vulnerability in AI code after deployment?
  • The first step is to isolate the vulnerable component and rotate every secret β€” API keys, passwords β€” that might have been exposed. Subsequently, it is fundamental to perform a retrospective Code Review to understand if the error is systemic in the way the team integrates AI suggestions.
  • Does using ChatGPT affect GDPR or SOC2 compliance?
  • Yes. If ChatGPT is used to process or write code that handles personal data (PII), the company must ensure that the use of the tool complies with the vendor’s Data Protection Agreements (DPA). Enterprise plans are designed precisely to meet these compliance requirements.

Protect your organisation with Web Application Penetration Testing.

Choose ISGroup for a practical, tailored engagement:

  • A focused assessment of your environment and requirements
  • Clear findings with a prioritised, actionable roadmap
  • Direct support from experienced specialists through remediation and implementation
Talk to an expert