What a human reviewer really looks at in AI-generated code
An AI-generated pull request can be tidy, readable, and full of green tests. It might follow the project’s style, update documentation, fix visible bugs, and complete a feature in a few hours. The problem is that security cannot be evaluated by checking if the code compiles or if the demo works: you need to understand what has actually changed in the application’s behavior.
A Secure Code Review on AI-generated or modified code does not look for “AI errors” in the abstract. It reads the diff within the product: which data is touched, which routes are exposed, which roles can perform actions, which queries change, which secrets enter the perimeter, which dependencies have been added, which tests have been modified, and which configuration will go into production.
For developers, CTOs, and software houses, the critical point is this: AI-generated code can be plausible even when it introduces a security regression. A human review is needed to connect code, business logic, trust boundaries, deployment, and real risk before the merge becomes a vulnerability in production.
Why a functional review is not enough
A functional review asks if the feature does what was requested. A Secure Code Review also asks what can be abused, which assumptions have changed, and which controls have been weakened. This difference matters significantly when the code is generated by AI, because the assistant tends to optimize for completing the task, not for preserving security boundaries that are not explicitly stated in the prompt.
A human reviewer reads the code with different questions than those of the demo: what happens if the user changes an ID, if the token is expired, if the role is low-privileged, if the tenant is different, if the upload contains an unexpected file, if a dependency executes scripts, if a private variable ends up in the frontend, or if the pipeline skips a check.
Scanners help, but they do not replace this reading. SAST, dependency scanning, secret scanning, and linting intercept useful patterns, but they do not always understand business logic, tenant isolation, workflow states, abuse of legitimate functions, role semantics, or the impact of a multi-file refactoring.
The first check: what the PR actually changed
AI-generated PRs can be broad: a prompt asks to “add this feature” and the agent modifies routes, components, middleware, tests, data schemas, dependencies, configurations, and documentation. The diff is coherent, but the reviewer must separate features, refactoring, tests, and security changes.
The first task is to classify the touched files. Controllers, API routes, server actions, middleware, queries, migrations, policies, .env.example files, CI/CD workflows, Dockerfiles, package managers, lockfiles, storage policies, auth configs, and tests do not carry the same weight. A small change to middleware or CORS can be riskier than a hundred lines of UI.
When the diff is too large, it must be reduced or broken down. An effective review means being able to say: this commit changes application logic, this changes presentation, this changes dependencies, this changes deployment. If everything is mixed together, the risk of accepting a regression increases.
Trust boundary: where code can no longer be trusted
A human reviewer looks for trust boundaries. The browser is not trusted, the request body is not trusted, and the same applies to headers, query strings, local storage, roles sent by the client, uploaded files, and model outputs. Even an internal service can be less trusted than it seems if it receives input from users or external integrations.
AI-generated code can cross these boundaries without marking them: it might use user_id from the client instead of the session user, build a query with unvalidated parameters, treat a token claim as a final permission without application-level checks, or pass LLM output to HTML, SQL, shells, tickets, emails, or downstream tools.
In review, every trust boundary must produce a precise question: who checks this data, where is it validated, where is it authorized, where is it logged, and what happens if the value is manipulated? If the answer is not in the code, the behavior is not verifiable.
Auth, roles, and server-side authorization
One of the most important areas is access control. The reviewer checks if new routes require login, if they verify roles, ownership, and tenants, if they apply the correct middleware, and if they do not delegate security decisions to the UI. An app might show the correct buttons but have APIs that can be called directly.
The review must follow sensitive routes: data reading, modification, deletion, export, upload, invitations, role changes, admin functions, billing, refunds, approvals, and impersonation. For each, the code must answer who can call it, on which objects, in what state, and with what limits.
In AI-generated code, the most common suspicious patterns are: if (user) without role checks, queries by ID without ownership checks, tenant_id taken from the body, isAdmin read from the client, middleware not applied to new routes, and tests that only verify the correct user. These problems often do not emerge from the demo.
Business logic and out-of-sequence workflows
Business logic is one of the main reasons why a human review is needed. An agent might implement individual steps well but fail to protect the sequence: an expired invitation might be reused, an order might be modified after payment, a refund might be called by the wrong role, a request might be approved twice, or an export might include data outside the perimeter.
The reviewer looks for invariants: what must always be true? A paid order does not change price. A document from one tenant does not appear in another. A suspended user does not generate API keys. A single-use link cannot be reused. A payment cannot be confirmed just because the frontend showed success.
The code review must also read states, not just functions. If the code modifies statuses, roles, payments, quotas, limits, invitations, or authorizations, state checks, idempotency, and negative cases are required.
Queries, ORMs, and the access layer
Many vulnerabilities in AI-generated code are not in the syntax, but in the query. An ORM makes code readable, but it does not prevent queries without tenant filters, overly broad joins, missing conditions, or filters applied after retrieving the data. A function might load all records and filter client-side or application-side incompletely.
The reviewer checks where ownership, roles, and tenants are applied: in the query, in the service layer, in the database policy, or in the controller. If the filter is duplicated across many routes, the risk increases that a new AI-generated route will forget it. If the service key bypasses policies and the code does not filter, the database becomes exposed through the backend.
Migrations and seeds must also be read: the AI might add sensitive columns, change defaults, create indexes on personal data, make a critical field nullable, or introduce temporary tables that then remain in production. The data model is part of the security surface.
Input validation, output handling, and uploads
AI-generated code often handles expected form input well. The review must look at unexpected inputs: manipulated JSON bodies, query strings, headers, path parameters, file uploads, falsified MIME types, filenames, markdown, HTML, numeric values out of range, arrays that are too large, nested objects, and partial payloads.
Validation must be server-side and consistent with domain logic. If a quantity must be positive, if a role can only have certain values, if a file must be a PDF, or if a date cannot be in the past, these constraints cannot live only in the frontend form.
Output must also be read. Unescaped HTML, verbose error messages, stack traces, queries, internal paths, tokens, database details, and other users’ data can leak from responses, logs, exports, or email templates. A secure review checks both what enters and what leaves.
Secrets, configurations, and logs
A Code Review on AI-generated code looks for secrets in ordinary and non-ordinary places: source code, tests, .env.example, READMEs, Dockerfiles, workflows, scripts, logs, seeds, fixtures, cloud configs, source maps, and frontend bundles. The AI might copy a “temporary” key to a place that then reaches deployment.
The review does not just say “move it to env.” It verifies if that secret has already leaked, if it needs to be rotated, if it is public or private, if it has minimal scope, and if it can end up on the client. A service role key, a database URL, or a cloud token in the wrong place can have an immediate impact.
Logging and error handling are other sensitive areas. During debugging, the AI might add very verbose logs and leave them in the code: the reviewer checks that payloads, tokens, queries, PII, internal paths, full stack traces, or customer data are not printed in production.
Dependencies and supply chain in the diff
When the AI encounters an error, it often proposes a library. This might be correct, but every new dependency enters the product’s surface. Packages with similar names, poorly maintained libraries, incompatible licenses, postinstall scripts, vulnerable transitive dependencies, and modified lockfiles must be read carefully.
The reviewer compares the real need with the cost of introducing the package: does a library already exist in the project, is the dependency maintained, does it have known CVEs, does it execute scripts, does it change parsing, auth, crypto, upload, sanitization, templates, or networking, and does the lockfile show more changes than expected?
For AI-generated code, it is also important to read what is removed. Replacing a library can eliminate implicit checks, updating a framework can change security defaults, and moving a function to a “simpler” package can lose existing validations.
AI-generated tests: what they confirm and what they don’t challenge
Tests written by AI tend to confirm the implemented behavior. They are useful, but they often cover the happy path: correct input, correct role, correct state, expected response. A Secure Code Review also looks at the test diff to see if they have been weakened.
Signals to check: removed tests, less precise assertions, overly permissive mocks, missing negative cases, updated snapshots without explanation, absent authorization tests, errors handled as success, and coverage that rises but does not cover trust boundaries.
Negative tests are needed for every sensitive area: wrong user, wrong tenant, expired token, low role, malicious input, invalid file, out-of-sequence state, duplicate request, exceeded limit. If a security bug is found, it must become a regression in the pipeline.
Configurations, CI/CD, and deployment
An AI-generated PR can also modify what sends code to production: Dockerfiles, CI/CD workflows, environment variables, CORS, security headers, debug mode, logging, IaC, cloud permissions, deploy scripts, branch protection, and test gates. These files are not “peripheral.”
The reviewer looks for dangerous simplifications: disabled tests, removed SAST, bypassed secret scanning, active debug, CORS *, weakened CSP, tokens with broader scope, automatic deployment from unprotected branches, more verbose logs, or staging environments connected to production.
If a config change is necessary, it must be explicit and proportionate. A build that passes because checks were turned off is not a more stable release: it is a less observable release.
Prompts, rules files, and persistent instructions
When the project uses Cursor rules, AGENTS.md, Claude instructions, project prompts, or similar files, the review must include them if they influence the generated code. These instructions can guide the agent’s style, libraries, tests, security, deployment, and decisions.
A rules file can contain useful guidance, but also dangerous shortcuts: “disable checks in dev,” “use this key,” “ignore flaky tests,” “do not ask for confirmation,” or “use permissive policies.” If the agent follows those instructions, the risk enters the PR even if the file looks like documentation.
The reviewer checks if persistent instructions are consistent with the project’s level of responsibility. For teams that use AI coding continuously, these rules become part of the development process and deserve explicit ownership.
How to read an AI PR that is too large
When an AI-generated PR is too broad, the reviewer must reconstruct a sequence that the diff often does not tell well. First, identify the new behavior: what feature was added, what bug was fixed, what flow changed. Then, separate the supporting changes: refactoring, file moves, renames, test updates, dependencies, and configurations.
The next step is to isolate high-responsibility areas. Even in a huge PR, some files weigh more than others: middleware, API routes, auth providers, queries, migrations, secret handling, deploy files, CI/CD workflows, database policies, server-side functions, and payment logic. The review must start there, not from the visual components that are easier to read.
If the PR modifies UI, API, database, and deployment together, asking for a separation can be a security decision. It is not bureaucracy: it is needed to prevent a change to CORS, a new package, a more permissive policy, or a query without a tenant filter from remaining hidden inside an apparently innocent feature.
For software houses, this discipline is even more important. An AI PR delivered to a client must be explainable: what changed, which files were generated, which assumptions were accepted, and which tests prove that boundaries remained correct. If you cannot explain the diff, it is difficult to defend its security.
What must come out of the review: findings and remediation
A useful Code Review does not just produce a list of problems. It must turn the diff into decisions: what blocks the merge, what must be corrected before deployment, what can be planned, which tests must be added, and which risks remain accepted with an owner. This is particularly important with AI-generated code, because the same pattern can repeat in subsequent PRs.
Findings should be described with application context. “Missing authorization check” is less useful than “the export endpoint uses the tenant ID received from the client and allows an authenticated user to export data from another workspace.” The second formulation allows developers, CTOs, and PMs to understand impact, priority, and remediation.
Remediation should prefer structural corrections. If an endpoint lacks tenant isolation, correcting just that route might not be enough: a centralized helper, a database policy, a negative test, and a review rule for new routes might be needed. If a dependency was added without reason, the fix might be to remove it, replace it with an already approved library, or block automatic installations without review.
The final report should indicate at least: reviewed perimeter, files and commits considered, excluded areas, blocking findings, plannable findings, suggested tests, evidence of fixes, residual risks, and recommendations for future AI PRs. This makes the review reusable, not just corrective.
When an independent Code Review is needed
An internal review might be enough for small, isolated changes without real data, roles, exposed APIs, and with reviewers who are experts in the touched areas. An independent Code Review is needed when the AI-generated PR modifies auth, authorizations, queries, databases, secrets, dependencies, deployments, CI/CD, cloud, payments, uploads, workflows, or critical business logic.
It is also needed when the diff is too broad to understand, when the team has accepted many “Accept All” proposals, when scanners are clean but the risk is in the application logic, or when a software house must deliver code to a client and wants evidence of control before release.
Code Review does not slow down AI coding speed: it prevents speed from shifting the cost to after go-live. A finding corrected before the merge costs less than a regression on data, roles, or production deployments.
How ISGroup can verify AI-generated or modified code
ISGroup can perform a targeted Code Review on high-risk areas of AI-generated code: diff, auth, access control, business logic, queries, secrets, logging, dependencies, configurations, tests, and pipelines. The perimeter can be a single PR, a feature, a module, a repository, or a pre-go-live flow.
| If AI code modified… | Main risk | Recommended check |
|---|---|---|
| Auth, roles, middleware, queries, API, business logic, secrets, or dependencies | Vulnerabilities in code and application regressions | Code Review |
| Web app, API, uploads, exports, or flows exposed online | Behaviors abusable from the outside | Web Application Penetration Testing |
| CI/CD, review policies, branch protection, test gates, and continuous agent use | Non-repeatable controls on releases | Software Assurance Lifecycle |
| Cloud, IaC, IAM, buckets, databases, envs, and deployments | Misconfiguration or excessive privileges | Cloud Security Assessment |
| Services, hosts, or components exposed with known versions and configs | Known vulnerabilities and technical surfaces | Vulnerability Assessment |
The choice of control depends on what was touched. For AI-generated code, Code Review and WAPT are often complementary: the review finds the cause in the code, the application test demonstrates the real behavior.
Do you have a PR or a feature generated with AI that touches data, APIs, roles, dependencies, or deployments? ISGroup can perform a targeted review before the merge or go-live.
Evidence to prepare
Prepare the repository, branch, PR, feature description, relevant prompts or instructions, list of modified files, executed tests, parts generated by AI, added dependencies, changed configurations, involved environments, and already known risks.
For an effective review, contextual data is also needed: user roles, trust boundaries, exposed APIs, data schemas, auth providers, databases, storage, pipelines, environment variables, cloud services, and critical flows. Code does not speak for itself when the vulnerability depends on the domain.
If the PR is broad, it is advisable to indicate what should be considered refactoring and what changes behavior. If it is not clear, the first finding may be the lack of separation between functional changes and security changes.
Decision before the merge
Block the merge if the review finds incomplete access control, exposed secrets, unfiltered queries, service keys on the client, removed security tests, suspicious dependencies, debug in production, permissive CORS without reason, error handling that exposes data, or weakened pipelines.
You can plan improvements with low residual risk and a clear owner after the merge: documentation, non-urgent refactoring, progressive increase of tests, reduction of duplication, or naming improvements. Vulnerabilities that expose data, roles, payments, cloud, or deployments must be corrected first.
The decision must produce evidence: what was reviewed, which findings were corrected, which negative tests were added, which risks remain, and who owns them.
Checklist for AI-generated code review
- Classify files and areas touched by the diff.
- Separate features, refactors, tests, dependencies, and configurations.
- Verify auth, roles, tenants, middleware, and server-side controls.
- Read queries, migrations, policies, and the access layer.
- Check business logic, states, idempotency, and out-of-sequence cases.
- Verify input validation, output encoding, uploads, and error handling.
- Look for secrets, verbose logs, .env files, connection strings, and tokens.
- Review dependencies, lockfiles, scripts, and new packages.
- Check CI/CD, CORS, headers, Dockerfiles, envs, IaC, and deployments.
- Demand negative tests for every sensitive area.
FAQ
- If SAST and tests are green, is a Code Review still needed?
- Yes. Scanners and tests intercept patterns and expected paths, but not always business logic, authorizations, trust boundaries, workflows, and the diff’s impact on the product.
- What does a human reviewer look at that the AI does not see?
- Application context: who can do what, which data is sensitive, which boundaries must not be crossed, which business assumptions must remain true, and what risk reaches production.
- Which AI PRs are riskiest?
- Those that touch auth, roles, APIs, databases, secrets, dependencies, CI/CD, cloud, payments, uploads, exports, logging, or deployment configurations.
- Does Code Review replace WAPT?
- No. Code Review finds problems in the code and logic. WAPT verifies real behavior from the outside. On exposed apps, both are often needed.
- When is the Software Assurance Lifecycle needed?
- When AI coding is used steadily by the team and it is necessary to make reviews, tests, policies, gates, and controls repeatable on every release.
- Can I have only one feature reviewed?
- Yes, if the perimeter is clear: PR, module, flow, route, or component. For features that touch data and roles, it is necessary to also include connected tests, configurations, and dependencies.
Protect your organisation with Code Review.
Choose ISGroup for a practical, tailored engagement:
- A focused assessment of your environment and requirements
- Clear findings with a prioritised, actionable roadmap
- Direct support from experienced specialists through remediation and implementation
Do not miss the best of cybersecurity.
Weekly expert analysis, real attacks and practical solutions in one newsletter.
Subscribe to Cyber WeeklySources and useful references
- OWASP Code Review Guide: https://owasp.org/www-project-code-review-guide/
- OWASP ASVS: https://owasp.org/www-project-application-security-verification-standard/
- OWASP Top 10 2021: https://owasp.org/Top10/2021/
- OWASP Broken Access Control: https://owasp.org/Top10/en/A01_2021-Broken_Access_Control/
- OWASP Authorization Cheat Sheet: https://cheatsheetseries.owasp.org/cheatsheets/Authorization_Cheat_Sheet.html
- OWASP Secrets Management Cheat Sheet: https://cheatsheetseries.owasp.org/cheatsheets/Secrets_Management_Cheat_Sheet.html
- OWASP SAMM: https://owasp.org/www-project-samm/
