Developers working on the Google stack today can use a wide variety of tools: Gemini Code Assist in the IDE, Gemini CLI in the terminal, Jules as an asynchronous agent on repositories, Google Antigravity as an agent-first environment, and Firebase and Google Cloud as backend and infrastructure. The advantage is clear: code, tests, UI, fixes, and configurations can advance much faster.
The risk arises when that speed crosses multiple surfaces simultaneously. A change can start from a prompt, pass through the repository, execute commands in the terminal, be verified in the browser, and end up in Firebase, Cloud Run, Cloud Functions, Cloud Storage, or IAM. Before go-live, the question is not whether Google tools are useful, but what they have changed in the product: code, APIs, permissions, Firebase Rules, service accounts, secrets, storage, deployments, and real data.
For a Google-native team, security does not end with Code Review or a Cloud Security Assessment. It requires looking at the app’s behavior, the generated code, the Firebase and Google Cloud configurations, and the process by which the agent produced or validated the changes.
Why the Google AI coding ecosystem must be treated as a single perimeter
Gemini Code Assist, Antigravity, and Jules do not have the same operating model. Gemini Code Assist works within the developer’s context and the IDE. Jules can take on asynchronous tasks, clone the repository into a cloud environment, install dependencies, modify files, and prepare changes for review. Antigravity, according to Google’s communications on Gemini 3, moves agents toward an experience where the editor, terminal, and browser are coordinated operational surfaces.
This difference matters. An inline suggestion that modifies a function has a limited and visible impact. An agent that modifies files, runs tests, installs packages, and verifies the UI in the browser produces a broader output: code, dependencies, configurations, logs, screenshots, tests, and assumptions about app behavior. If the backend is Firebase or Google Cloud, the transition from “it works” to “it is publishable” requires specific checks: it is not enough for the artifact to be functional; you must verify what has changed within authorization boundaries, data, endpoints, access rules, and cloud privileges.
Gemini Code Assist: agent mode, context, and sensitive areas
Gemini Code Assist can help with explanations, code generation, refactoring, agent mode, and repository tasks. When working on sensitive application areas, the review must focus on points the assistant cannot fully know: business rules, tenant isolation, internal roles, real data, compliance, corporate conventions, and undocumented architectural decisions.
The most concrete risks emerge when a suggestion touches authentication, authorization, middleware, routes, input validation, error handling, logging, dependencies, or deployment configurations. A change can be consistent with the prompt but wrong for the product: a check might be moved to the frontend, a route might be created without server-side auth, or a test might confirm only the happy path.
Before merging, changes generated or corrected with Gemini Code Assist must be read as production diffs. If they involve APIs, data, or roles, negative tests are required: unauthorized users, other users’ objects, different tenants, expired tokens, out-of-sequence requests, manipulated payloads, and direct access to routes not exposed in the UI.
Antigravity: editor, terminal, and browser in the same workflow
Antigravity shifts the focus from individual assistance in the IDE to a more agentic flow. Google describes it as a platform where agents can have direct access to the editor, terminal, and browser. This model is powerful because it allows for planning, modifying, executing, and verifying in a more integrated way, and that is precisely why control must be more rigorous.
When an agent works between the terminal and the browser, an error does not remain confined to the modified file: it can become a shell command, a package installation, a UI test, a configuration change, a local service startup, output containing secrets, screenshots with sensitive data, or a deployment started too early. The review must include what was executed, not just what was written.
Install, migration, deploy, delete, cloud CLI, modification of environment variables, and Firebase or Google Cloud configurations should require explicit approval. Commands executed by the agent must be tracked when the repository contains data, keys, or access to real services. If the browser records sessions, screenshots, or flows, it is necessary to prevent personal data, tokens, administrative dashboards, and unnecessary customer information from entering them.
Jules: autonomous PRs, cloud VMs, and functional tests
Jules is designed to delegate development tasks: bug fixes, tests, documentation, features, and updates. It works by cloning code into a virtual machine, installing dependencies, and modifying files. This reduces the impact on the local environment, but it does not prove that the result is secure.
A PR generated by Jules may be well-organized, compile, and pass tests, yet still contain dependency drift, tests that confirm the new behavior without challenging it, simplified authorization logic, overly permissive error handling, packages added to solve a problem, or environment configurations changed to make the build start.
Before accepting an asynchronously generated PR, the team should review diffs, lockfiles, install/build/test scripts, configuration files, new routes, middleware, updated tests, and data used in the environment. The VM should not receive production secrets or access to real systems not necessary for the task. If Jules has corrected or added tests, it is necessary to understand if they cover abuse and security regressions or just the expected path.
Firebase Rules generated or modified with AI
Firebase is often the point where a Google-native prototype becomes a product: Authentication, Firestore, Realtime Database, Cloud Storage, Functions, Hosting, and integrations. The most common risk is thinking that a check in the UI is equivalent to a security rule.
Firestore, Realtime Database, and Cloud Storage must have rules that start from “default deny” and open only for explicit cases. If the assistant generates rules to “make” read and write operations work, you must verify ownership, roles, tenants, document status, custom claims, and the separation between public and private data. Rules like generic authenticated access or overly broad paths may be sufficient for a demo but dangerous in production.
The review must include tests with different users, anonymous users, expired tokens, other tenants’ documents, private files, uploads, deletes, and out-of-sequence operations. Firebase Security Rules must be tested with emulators and negative cases, not just tried from the UI: if a rule allows read or write because the frontend does not show the button, the control is in the wrong place.
Google Cloud IAM and service accounts
AI tools can help write code that uses Google Cloud, but service accounts and IAM remain a critical surface. The typical risk is using overly broad roles to make Cloud Functions, Cloud Run, storage, databases, schedulers, or integrations work.
A service account should not have Owner, Editor, or primitive roles for convenience. Every workload should have a dedicated identity, minimal permissions, static keys avoided whenever possible, and audits on impersonation and access. If the agent suggests creating or downloading a JSON key, the team must ask if a more secure alternative exists and if that key could end up in repositories, logs, prompts, or artifacts.
For Firebase and Google Cloud apps generated with AI, Cloud Security Assessment and Code Review must meet: the code shows how service accounts and APIs are used, while the cloud configuration shows what those roles can actually do. An application vulnerability becomes more severe if the connected service account can read storage, databases, or secrets beyond what is necessary.
Tokens, API keys, and secrets in the frontend
One of the most frequent points of confusion in Firebase apps is the difference between intended client-side configurations and server-side secrets. Some Firebase values are intended to be used by the client alongside robust security rules. Other tokens, private keys, webhook secrets, service account keys, and external provider credentials must remain server-side.
AI assistants can generate code that seems to work but puts sensitive values into frontend bundles, exposed .env files, logs, tests, or examples. Before deploying, you must search for keys in the repository, history, prompts, build logs, artifacts, and code served to the browser. Public API keys should be limited with appropriate restrictions; private secrets must be managed in Secret Manager, read only by server-side runtimes, and never printed. If a key has ended up in an unintended context, removing it from the file is not enough: it must be rotated.
Cloud Functions, Cloud Run, and exposed APIs
Many apps created with Google AI tools use Cloud Functions, Cloud Run, or API endpoints to connect the frontend, database, storage, payments, CRM, ticketing, or AI models. These are exposed surfaces and must be tested as such.
The most common problems are public endpoints without auth, overly permissive IAM invokers, open CORS, detailed errors, non-allowlisted callbacks and redirects, missing rate limits, insufficient input validation, and logs containing PII or tokens. An agent can generate an endpoint to quickly complete a feature without applying the same security discipline present in the rest of the backend.
Before go-live, every endpoint must be called directly outside of the UI flow. You must test anonymous access, users with low roles, different tenants, manipulated payloads, repeated requests, unexpected callbacks, malicious files, and error conditions. If the app is reachable online, the WAPT must verify the real behavior, not just the code.
Dependencies, SDKs, and installation scripts
Gemini Code Assist, Jules, or Antigravity can add packages, update SDKs, modify lockfiles, or change build scripts. These changes seem operational, but they affect the supply chain and must be reviewed with the same attention as application code.
Every change to package.json, requirements.txt, pyproject.toml, lockfiles, postinstall scripts, build processes, or tests must be examined to understand why the package was introduced, whether it is maintained, what license it has, what transitive dependencies it brings, whether known vulnerabilities exist, and whether the script executes unexpected code.
In the Google and Firebase context, it is appropriate to pay attention to obsolete SDKs, libraries that handle tokens and sessions, unofficial wrappers for cloud services, and packages with similar names. A PR generated by an agent might be accepted because it resolves a build error, but it could introduce a dependency that the team would not have manually approved.
Real data in prompts, logs, screenshots, and VMs
Agents that work between the editor, terminal, browser, and VM can see more context than a simple chatbot: files, test outputs, screenshots, errors, logs, payloads, tickets, internal documentation, and browser sessions. This increases utility, but also the risk of data disclosure.
For debugging and validation, it is better to use synthetic or minimized data. Screenshots of dashboards, browser recordings, Cloud Functions logs, Firestore dumps, webhook payloads, and customer tickets should not enter the agent’s context unless strictly necessary. When they do enter, you need to know where they are stored, for how long, and who can access them.
Control is not just a matter of privacy. Real data in tests can make a change pass because the dataset is favorable, while it fails on edge cases. For security, you need datasets that include different roles, tenants, unauthorized records, private files, disabled users, and malicious input.
Checklist before go-live
Tools and workflow
- Identify which tools contributed to the code: Gemini Code Assist, Antigravity, Jules, Gemini CLI, or other Google tools.
- Collect PRs, diffs, branches, terminal commands, executed tests, browser recordings, screenshots, and artifacts.
- If an agent worked in a VM or asynchronously, verify which dependencies it installed and which files it modified.
Firebase and data
- Review Firestore, Realtime Database, and Cloud Storage Rules with a “default deny” approach.
- Test access with anonymous users, authenticated users, different roles, different tenants, and unauthorized files.
- Verify Authentication, custom claims, invitations, password resets, premium content, uploads, and deletes.
- If real data is used, check logs, retention, and backups.
Google Cloud and IAM
- Check service accounts, IAM roles, static keys, impersonation, Cloud Run/Functions invokers, storage, Secret Manager, Cloud Logging, and environments.
- Avoid primitive roles for convenience.
- Every workload must have an identity and permissions consistent with what it needs to do.
APIs, frontend, and secrets
- Verify public routes, CORS, callbacks, redirects, API keys, tokens in the frontend, server-side secrets, error handling, and rate limits.
- Perform secret scanning on repositories, history, logs, prompts, and artifacts.
- Rotate any key that ended up in the wrong place.
Dependencies and tests
- Review lockfiles, SDKs, added packages, install/build/test scripts, and generated tests.
- Integrate negative tests for IDOR/BOLA, tenant isolation, role abuse, malicious uploads, business logic, and expired sessions.
- If the app or APIs are online, perform WAPT on the exposed environment.
When an internal review is enough and when independent verification is needed
An internal review may suffice if Google AI tools have produced isolated, non-exposed changes, without real data, without Firebase Rules, without IAM, without service accounts, and without public APIs. Even then, it is useful to know which parts were generated and which were checked by a competent person.
Independent verification is needed when the app uses Firebase or Google Cloud with real data, when access rules, service accounts, Cloud Functions, Cloud Run, APIs, storage, callbacks, redirects, dependencies, or deployments have been modified. It is also needed when Jules or Antigravity have produced large PRs, executed commands, or validated behavior only with functional tests.
The point is not to slow down Google tools. It is to separate what can be accelerated from what must be verified: authorizations, real data, public surfaces, secrets, service accounts, rules, APIs, and cloud configurations.
How ISGroup verifies a Google, Firebase, and Cloud app created with AI
The control changes based on what Gemini Code Assist, Antigravity, or Jules have modified. If the risk is in the code, PRs, middleware, input validation, use of secrets, or dependencies, Code Review helps identify vulnerabilities and regressions before merging. If the app or APIs are reachable online, Web Application Penetration Testing verifies real behavior from the outside. If the risk concerns Firebase, Google Cloud, IAM, service accounts, storage, Cloud Run, or Cloud Functions, the Cloud Security Assessment verifies configurations and privileges in the real context.
| If the Google AI tool touched… | Main risk | Recommended control |
|---|---|---|
| Application code, Jules PRs, middleware, validation, error handling, dependencies | Vulnerabilities or code regressions | Code Review |
| Web app, APIs, Cloud Functions, Cloud Run, or exposed routes | Behaviors exploitable from the outside | Web Application Penetration Testing |
| Firebase Rules, Cloud Storage, IAM, service accounts, Secret Manager, Google Cloud config | Cloud misconfiguration or excessive privileges | Cloud Security Assessment |
| Trust boundary, multi-tenant, data flows, integrations, payments, integrated LLMs | Weak architectural assumptions | Secure Architecture Review |
| Continuous use of Gemini, Antigravity, or Jules in the release cycle | Non-repeatable controls on releases and pipelines | Software Assurance Lifecycle |
The choice of control depends on what has actually changed: code, exposed behavior, Firebase and Google Cloud, architecture, or the development process. Before go-live, it is advisable to define that perimeter and verify the actual risk to the application.
Have you used Gemini Code Assist, Antigravity, or Jules on an app that uses Firebase or Google Cloud? ISGroup can help you verify code, APIs, authorizations, Firebase Rules, service accounts, secrets, storage, and cloud configurations before real data and users enter production.
Evidence to prepare before the review
Before involving an external team, it is advisable to prepare repositories, PRs, branches, diffs, a list of tools used, environment URLs, role descriptions, Firebase projects, Google Cloud projects, service accounts, rules, Cloud Functions and Cloud Run, APIs, storage buckets, callbacks, redirects, and main dependencies.
Logs and artifacts cleared of sensitive data, a list of executed commands, generated tests, browser recordings (if available), decisions already made regarding accepted risks, and planned remediations are also useful. This evidence allows for distinguishing code problems from cloud misconfigurations, application vulnerabilities from process gaps, and immediate risks from improvable areas.
The question to ask before publication
The decision should not be “do we accept or not accept the agent’s PR” in the abstract. It should be: what data does it expose, what rules does it change, what service accounts does it use, what endpoints does it publish, what secrets does it handle, and what residual risk remains after remediation?
Gemini Code Assist, Antigravity, and Jules can greatly accelerate development on the Google stack. Security is needed to prevent that speed from bringing overly broad Firebase Rules, excessive service accounts, exploitable APIs, tokens in the frontend, accessible storage, or PRs that are not truly understood into production. The final question is simple: has the Google, Firebase, and Cloud app created or modified with AI been verified as an exposed product, or just accepted because it works in tests? If the answer is not clear, the next step is not to slow down development: it is to define the risk before real data, users, service accounts, and public surfaces enter production.
FAQ
- Does Gemini Code Assist make the code it generates secure?
- No. It can help produce and review code, but the team must verify application logic, authorizations, dependencies, secrets, Firebase Rules, and real behavior before production.
- Jules works in a VM: is that enough to trust the PR?
- No. The VM isolates task execution, but the PR can still introduce authorization bugs, dependency drift, weak tests, risky error handling, or configurations inconsistent with the product.
- Is Antigravity riskier than a traditional IDE?
- The risk changes because the agent can operate between the editor, terminal, and browser. This increases productivity but requires control over commands, diffs, tests, visible data, and touched configurations.
- Are Firebase Security Rules generated with AI reliable?
- They must be verified. A rule that makes the demo work may be too open for production. Tests with different users, tenants, roles, private files, and unauthorized operations are required.
- When is a Cloud Security Assessment needed?
- When AI-assisted development has touched Firebase, Google Cloud, IAM, service accounts, Cloud Run, Cloud Functions, Cloud Storage, Secret Manager, callbacks, redirects, or deployment configurations.
Protect your organisation with Cloud Security Assessment.
Choose ISGroup for a practical, tailored engagement:
- A focused assessment of your environment and requirements
- Clear findings with a prioritised, actionable roadmap
- Direct support from experienced specialists through remediation and implementation
Useful sources and references
- Gemini Code Assist agent mode: https://developers.google.com/gemini-code-assist/docs/agent-mode
- Gemini Code Assist agentic chat: https://developers.google.com/gemini-code-assist/docs/use-agentic-chat-pair-programmer
- Jules documentation: https://jules.google/docs/
- Jules product page: https://jules.google/
- Google Gemini 3 and Antigravity announcement: https://blog.google/products-and-platforms/products/gemini/gemini-3/
- Gemini 3 for developers: https://blog.google/technology/developers/gemini-3-developers/
- Firebase Security Rules basics: https://firebase.google.com/docs/rules/basics
- Firebase service accounts: https://firebase.google.com/support/guides/service-accounts
- Google Cloud service account best practices: https://cloud.google.com/iam/docs/best-practices-service-accounts
- Google Cloud IAM best practices: https://cloud.google.com/iam/docs/using-iam-securely
- OWASP Top 10: https://owasp.org/Top10/
- OWASP Top 10 for LLM Applications 2025: https://owasp.org/www-project-top-10-for-large-language-model-applications/
