Devin by Cognition and autonomous software engineering: how to verify code produced by end-to-end agents
Devin shifts the AI coding paradigm to a different level compared to classic assistants. It doesn’t just complete code or suggest changes in a file: it can take charge of complete engineering tasks, work in a dedicated cloud environment, use terminals, browsers, and filesystems, install dependencies, run tests, fix bugs, and prepare pull requests.
For CTOs, engineering managers, and software houses, the relevant question is not just whether Devin produces working code. It is understanding how to verify an output generated by an agent that has traversed the entire task cycle: tickets, repositories, environments, commands, tests, logs, integrations, and pull requests. Before merging or deploying, the core issue is concrete: what has Devin changed in the code, tests, dependencies, secrets, configurations, APIs, and application behavior?
Why autonomous software engineering changes the review process
With a traditional assistant, the reviewer often sees a limited suggestion. With Devin, the result can be a complete PR: the agent has planned, explored the codebase, executed commands, modified files, updated tests, installed packages, and verified its own work. This makes the review structurally more difficult because the final diff doesn’t always tell the whole story: it doesn’t show the commands executed, the errors encountered, the logs read, the data seen in the browser, the packages tried and discarded, nor the assumptions made to pass the test suite.
A PR might look ready because the tests are green, but it could contain an authorization regression, an unapproved dependency, or a permissive fallback. The main risk is not that Devin gets the code wrong: it is accepting an autonomous cycle as if it were equivalent to a full human review. The team must be able to reconstruct what happened between the ticket and the PR.
Cloud workspace, terminal, and browser
Devin works as an autonomous agent in a dedicated environment. Official documentation describes the ability to write, execute, and test code; enterprise documentation discusses workspaces with shells, browsers, and code editors, as well as deployment models with cloud components and devboxes. This architecture is useful because the agent can continue working even when the developer is not present, but it is also a surface that must be governed with care.
A shell command can install packages, read environment variables, modify files, run migrations, execute tests, start services, or interact with cloud CLIs. The browser can see dashboards, internal screens, test data, monitoring tools, and application flows. Before using Devin on corporate repositories, you must explicitly decide which environments it can reach. Test accounts, synthetic data, allowlisted repositories, and limited secrets reduce risk; production secrets, dashboards with real data, live APIs, and sensitive cloud environments should not be available by default.
Agentic PRs: green tests are not enough
Devin can generate PRs that resolve tickets, bugs, or backlog tasks. The PR is the point where the team regains control, and if that control is limited to verifying that the code compiles and tests pass, the value of autonomy becomes its weak point.
Areas to verify manually include auth, authorizations, tenant isolation, APIs, middleware, input validation, error handling, logging, secrets, dependencies, pipelines, and deployment configurations. A change that fixes a login bug might shift a check in the frontend; an API fix might accept a client-side parameter that was previously derived from the server; a test generated alongside the code might confirm the new behavior instead of challenging it. Every PR generated by Devin on sensitive areas should have explicit reviewers, branch protection, negative tests, and a diff review by change set. If the modification exposes a web app or an API, the actual behavior must be tested with a WAPT approach, not just unit tests.
Authorizations and business logic
Many vulnerabilities introduced by autonomous agents do not break functionality: the app continues to do what the ticket requested, but it loses a constraint. A user might be able to read another tenant’s data, a role might call an unintended route, a workflow might skip a state, or server-side validation might be considered redundant and removed.
For this reason, the review of a Devin PR must include abuse cases. An account with low privileges must try to call administrative endpoints; a user from another tenant must try to read or modify records that are not theirs; an expired invitation must be reused; an out-of-sequence request must try to alter the state of an order, a case, or a workflow. Business logic is not always visible to automatic scanners: domain knowledge is required. Which operations require approval? Which data is separated by client, role, group, or country? Which conditions must not be controllable by the client?
Secrets in the autonomous workflow
When an agent works end-to-end, secrets can emerge in multiple places: repositories, .env files, shell history, test logs, build outputs, prompts, code comments, PR descriptions, CI/CD artifacts, browsers, and monitoring tools. Devin’s enterprise documentation indicates the use of a Secrets function to share credentials when necessary, but this does not eliminate application risk: the generated code can still print a value, pass it to the frontend, include it in a test, or copy it into an example file.
Before merging, secret scanning is required on the diff and relevant history. If a key has ended up in a prompt, log, artifact, or PR, it must be rotated. Access granted to Devin must be limited to the task: dev tokens, minimal scopes, separate repositories and environments, and no production secrets unless strictly necessary and approved.
Dependencies, scripts, and supply chain
An autonomous agent that wants to close a ticket might add libraries, update lockfiles, change build scripts, or install tools to reproduce a bug. This can be correct from a functional perspective and risky from a supply chain perspective. Every change to package.json, requirements.txt, pyproject.toml, go.mod, Cargo.toml, lockfiles, or postinstall scripts must be reviewed: you need to understand why the dependency was introduced, whether it is maintained, what license it has, what transitive dependencies it brings, whether there are known vulnerabilities, and whether the scripts execute unintended code.
Dependency review is particularly important when Devin resolves build or test errors, because in those cases the agent might choose the fastest path: replacing a library, downgrading a version, disabling a check, or adding a helper package. The team must decide if that choice is acceptable for the product, not just if it solves the immediate problem.
Integrations: GitHub, Slack, Linear, Datadog, AWS
Devin can fit into the team’s workflow through ticketing tools, repositories, chat, and monitoring, making it natural to delegate tasks from Linear or Jira, comment in Slack, open PRs on GitHub or GitLab, consult logs, and intervene on real issues. Every integration, however, expands the context and permissions available to the agent.
An overly broad GitHub integration can provide access to unnecessary repositories; a Slack thread can contain customer data or tokens; an observability tool can show sensitive payloads, stack traces, headers, queries, or internal endpoints; AWS access can allow modifications to resources unrelated to the ticket. Integration configuration must follow the principle of least privilege. Allowlisted repositories, dedicated channels, read-only permissions where sufficient, separate tokens per environment, and periodic access reviews significantly reduce risk.
Deploy, pipelines, and configurations
Devin can be used to fix CI, update Dockerfiles, correct configurations, modify pipelines, or prepare deployments. These changes are not neutral: they change how code reaches production. A PR that “fixes the pipeline” might widen permissions, expose secrets in logs, disable security steps, change branch protection, make a Dockerfile more permissive, modify environment variables, or skip tests. A deployment fix might introduce debug mode, open CORS, detailed errors, or missing rate limits.
Changes to CI/CD, containers, IaC, cloud config, env, secrets, feature flags, and deployment scripts must have a separate review. If the task touches cloud, network, IAM, buckets, databases, or pipelines, the perimeter extends beyond just Code Review and may require a Cloud Security Assessment or a Secure Architecture Review.
Audit trail: reconstructing the “why,” not just the “what”
In autonomous software engineering, the final diff is not enough. An audit trail is needed that links tickets, prompts, sessions, commands, tests, logs, errors, PRs, and approvals: without this chain, it is difficult to understand why the agent chose a solution and what alternatives it discarded. The audit trail is also useful for incident response: if a vulnerability reaches production, the team must be able to understand which task introduced it, which checks were passed, who approved the PR, which tests were missing, and which data or secrets were available to the agent.
Keeping logs and trajectories does not mean accumulating sensitive data. It means defining what to record, how to redact secrets, who can access audit logs, how long to keep them, and how to use them to improve policies over time.
Checklist before merging
PR, diff, and tests
- Review the PR for small, verifiable change sets.
- Check the added or modified tests and verify that they include negative cases, not just the happy path.
- If the agent updated snapshots, mocks, or tests to pass the suite, read those changes carefully.
Commands and environment
- Check the shell commands executed, packages installed, scripts launched, migrations, builds, and tests.
- Verify that the environment did not have production secrets or access to unnecessary real systems.
- If browsers or dashboards were used, ensure that no sensitive data was exposed.
Auth, data, and APIs
- Test server-side authorizations, roles, tenant isolation, IDOR/BOLA, endpoints not visible in the UI, expired tokens, malicious input, and business logic abuse.
- Areas modified by Devin should not be accepted just because the ticket is marked as closed.
Secrets and dependencies
- Run secret scanning on diffs, logs, artifacts, and relevant history.
- Rotate what has been exposed.
- Review new dependencies, lockfiles, install/build/test scripts, licenses, and known vulnerabilities.
Pipelines and deploy
- Check Dockerfiles, CI/CD, branch protection, secrets in pipelines, env, feature flags, IaC, cloud config, and deployment scripts.
- Every change that impacts the path to production must have an owner and dedicated approval.
When is an internal review enough and when is an independent verification needed?
An internal review might be enough if Devin worked on non-exposed tasks, without real data, without auth, without new dependencies, without pipelines, and without operational integrations. Even then, the team should keep the diffs, tests, and main commands.
Independent verification is needed when Devin has modified authentication, authorizations, APIs, data, dependencies, secrets, pipelines, cloud, deployments, or critical application logic. It is also needed when agentic PRs become a stable part of the development cycle and the team must define policies, branch protection, reviewers, logging, secret handling, and risk thresholds. The point is not to slow down Devin: it is to separate what can be delegated from what must be verified.
How ISGroup can verify code and workflows produced by Devin
The control changes based on what Devin has modified. If the risk is in the PR, application code, authorization logic, secrets, or dependencies, Code Review helps identify vulnerabilities and regressions before merging. If the result is an exposed web app or API, Web Application Penetration Testing verifies the actual behavior from the outside.
| If Devin touched… | Main risk | Recommended control |
|---|---|---|
| PR, application code, middleware, auth, roles, dependencies | Vulnerabilities or regressions in the code | Code Review |
| Web app, API, public routes, exposed user flows | Abusable behaviors from the outside | Web Application Penetration Testing |
| Pipelines, deploy, cloud config, IaC, secrets in CI/CD | Misconfiguration or excessive privileges | Cloud Security Assessment |
| Architecture, trust boundary, integrations, sensitive data | Weak architectural assumptions | Secure Architecture Review |
| Continuous use of Devin in the engineering cycle | Non-repeatable controls on agentic PRs | Software Assurance Lifecycle |
The choice of control depends on what has actually changed: code, exposed behavior, pipelines, integrations, or the development process. Before go-live, it is advisable to define that perimeter and verify the actual risk to the application.
Evidence to prepare before the review
Before involving an external team, it is advisable to collect tickets, PRs, branches, diffs, tests, executed commands, available logs, added dependencies, modified configurations, a list of accessible repositories, active integrations, and a description of the environments used by Devin. Information on secrets granted to the agent, branch protection policies, involved reviewers, CI/CD, deployment scripts, cloud config, processed data, consulted monitoring systems, and decisions already made on accepted risks or planned remediation are also useful.
FAQ
- Devin tests its own code: is that enough to go to production?
- No. The agent’s tests can confirm that the ticket has been resolved, but they do not prove on their own that authorizations, business logic, secrets, dependencies, and configurations are secure.
- What is the main risk of PRs generated by Devin?
- Accepting a PR because it is complete and passes tests, without reconstructing commands, dependencies, decisions, and impact on auth, APIs, data, and pipelines.
- Can Devin handle secrets securely?
- It can offer dedicated mechanisms to share credentials, but the generated code can still log, copy, or expose secrets. You need to check usage, scope, rotation, and artifacts.
- When is the Software Assurance Lifecycle needed?
- When Devin becomes a stable part of the engineering cycle: tickets, PRs, reviews, CI/CD, deployments, and remediation must have repeatable rules, not case-by-case decisions.
- When is Web Application Penetration Testing needed?
- When the code produced or modified by Devin exposes web apps, APIs, or flows reachable by external users, customers, partners, or employees.
Protect your organisation with Web Application Penetration Testing.
Choose ISGroup for a practical, tailored engagement:
- A focused assessment of your environment and requirements
- Clear findings with a prioritised, actionable roadmap
- Direct support from experienced specialists through remediation and implementation
Useful sources and references
- Devin official site: https://devin.ai/
- Devin docs index: https://docs.devin.ai/llms.txt
- Cognition official site: https://cognition.ai/
- Windsurf docs: https://docs.windsurf.com/
- Devin GitHub integration: https://cognitionai-enterprise.mintlify.app/Integrations/gh
- OWASP Top 10 for LLM Applications 2025: https://owasp.org/www-project-top-10-for-large-language-model-applications/
- OWASP Agentic AI Threats and Mitigations: https://genai.owasp.org/resource/agentic-ai-threats-and-mitigations/
- OWASP Top 10: https://owasp.org/Top10/
