Exposed database in vibe coding: how to avoid it before go-live

Database esposto nel vibe coding come prevenirlo

Exposed database in vibe coding: how to avoid it before go-live

A database can be exposed even when the app seems to be working perfectly. The UI shows only the correct data, login works, APIs respond, the deploy passes, and the founder can give a convincing demo. The risk arises elsewhere: a connection string in the repository, a database reachable from the internet, overly permissive rules, a public backup, a user with excessive privileges, or an AI-generated endpoint that returns more data than intended.

In vibe coding, the database comes early. Tools like Lovable, Bolt.new, Replit Agent, Cursor, Copilot, Codex, and Claude Code help connect tables, storage, authentication, APIs, dashboards, and deployments in very short timeframes. This speed is useful, but it often pushes towards configurations that make the MVP work immediately while postponing hardening. For founders, PMs, and junior developers, the question before beta is not just “does the database respond?”, but who can reach it, with what credentials, from which environments, through which APIs, with what rules, what backups, what exports, and what real data.

🔴 Web Application Penetration Testing: identify hidden risks and strengthen your security with a focused assessment by ISGroup specialists.

Exposed database doesn’t just mean SQL injection

When talking about an exposed database, many immediately think of SQL injection. It is a significant risk, but in AI-created apps, exposure often arises earlier: configuration, credentials, permissions, network, storage, backups, and APIs. A database can be protected against injection and still expose data because a policy is open, a key is in the client, or a dump ended up in a bucket.

The most common errors are practical: committed .env files, connection strings in logs, disabled RLS, permissive Firebase Security Rules, MongoDB accessible from overly broad IPs, Postgres with unnecessary public access, staging databases with real data, service keys used by the frontend, downloadable backups, unprotected CSV exports, and debug routes left online. The problem is not the chosen database — PostgreSQL, Supabase, Firebase, MongoDB Atlas, RDS, Cloud SQL, PlanetScale, Neon, Pinecone, pgvector, or a self-hosted database can all be used securely — but publishing an app before verifying how data enters, is read, is copied, and can exit.

Connection strings, .env files, and credentials in the wrong place

The first exposure is often a connection string. During AI development, the developer pastes a database URL into a chat, the agent generates an .env file, a deployment example ends up in the README, a log prints the configuration, or a private variable is placed in the frontend to resolve an error. A connection string can contain host, database, user, password, TLS parameters, and privileges: if it ends up in Git, prompts, tickets, logs, artifacts, or bundles, it is not enough to remove it from the file; it must be rotated. Even a staging credential can be critical if that staging environment contains real data or has access to shared resources.

Before go-live, it is necessary to search for database URLs, passwords, service keys, .env files, .pem files, dumps, backups, and CLI credentials in repositories, Git history, branches, tags, CI/CD logs, prompts, issues, and artifacts. The next step is to separate development, staging, and production keys: a key used by the agent or in a demo should never be able to read the real database.

Database reachable from the internet

Many cloud databases come with a reachable endpoint by default. Sometimes it is necessary; often it is just convenient. If the database accepts connections from broad IPs, from the entire internet, or from uncontrolled networks, passwords and secrets become the only barrier, which is insufficient for a production environment. For an MVP, this happens when the team needs to quickly connect apps, dashboards, local tools, and deployments.

It is important to check network access, IP allowlists, security groups, firewalls, VPCs, private networking, and connection rules. In self-hosted PostgreSQL, parameters like listen_addresses and pg_hba.conf decide where connections can come from; in MongoDB Atlas, the IP access list is a key control; in cloud environments, security groups and network routes must be consistent with the actual architecture. A production database should not be accessible from every temporary environment, preview deployment, or personal laptop without necessity: if administrative access is needed, it is preferable to use controlled channels like bastions, VPNs, private endpoints, or temporary access.

Permissive rules in Supabase, Firebase, and BaaS

BaaS platforms greatly accelerate development, but security depends on the rules. Supabase requires Row Level Security (RLS) and consistent policies to isolate users and tenants; Firebase uses Security Rules for Firestore, Realtime Database, and Storage. If the AI generates an open rule to make the demo work, the provider remains robust, but the data can leak.

In Supabase, a table with private data should not have generic policies or missing RLS. The anon key can stay in the client only if RLS and policies truly limit what the client can read and write, while the service role key must stay out of the frontend, logs, and repository. In Firebase, rules like allow read, write: if true or checks based only on request.auth != null can be too broad. Before go-live, it is useful to test reading and writing with different users, different tenant data, missing tokens, expired tokens, and direct requests, without trusting that the UI only shows the correct records.

Overly privileged database users

An app should not connect to the database with an administrative user if it is not needed. In vibe coding, however, the AI may suggest the credential that “solves” the permission error — superuser, database owner, service role, account with access to all tables, key with global privileges — making every bug more expensive. If the app has SQL injection, SSRF, RCE, log leaks, or an exploitable endpoint, the attacker inherits overly broad privileges. Even without an external attack, an AI agent with terminal or database tool access can perform operations that should not be within its scope.

The correct control is the principle of least privilege. This means using separate users for the application, migrations, maintenance jobs, analytics reading, and administration, limiting permissions to necessary tables, schemas, and operations, and avoiding runtime users with DROP, ALTER, global privileges, or access to data outside the scope.

Exposed backups, dumps, and exports

Often the operational database is more protected than its copies. SQL dumps, CSV exports, snapshots, automatic backups, temporary files, reports, pipeline artifacts, and support folders can contain the entire data heritage. In an MVP, these files are created for debugging, migration, demos, or testing and then forgotten in cloud buckets, BaaS storage, repositories, staging environments, laptops, ticketing systems, emails, CI/CD, container images, or temporary folders.

Backups and dumps must have limited access, encryption where appropriate, defined retention, access logging, and separation from public assets. A real dump in staging can be more exposed than production, and a CSV export generated by an admin route might be downloadable without proper controls. If a copy contains real data, it must be treated as production.

Generated APIs that expose the database

Even when the database is not directly public, APIs can expose it. An agent can generate CRUD endpoints, admin routes, search functions, exports, debug tools, filters, reports, or dashboards that read data without sufficient authorization, effectively making the endpoint the exposed database.

Frequent examples: /api/users returns all users, /api/orders?user_id=... accepts IDs from the client without verification, /api/export produces a full CSV, /api/debug/db remains active in production, an admin route only checks for login, a global search ignores tenants and roles, a serverless function uses the service key and does not filter results. Before go-live, it is necessary to inventory the routes that read from the database and try direct calls, manipulated parameters, missing tokens, low-privileged roles, different tenants, and unexpected HTTP methods. If an endpoint returns data that the UI would not show, the database is exposed through the app.

Storage, buckets, and metadata linked to the database

Many apps associate database records with files: documents, images, attachments, invoices, avatars, contracts, reports, exports, datasets. If the database contains only the path but the bucket is public, protecting the record is not enough; if the URL is predictable, a user can download others’ files even without a database query.

It is necessary to verify buckets, paths, signed URLs, link expiration, policies per user or tenant, previews, thumbnails, extracted text, and metadata. A file can be private but have a public preview, a filename can reveal customer information, and an export can be saved in storage and remain accessible after the intended expiration. The practical test is simple: upload files with two different users, then try cross-downloading and deleting by changing IDs, paths, or filenames. If the app uses Supabase Storage, Firebase Storage, S3, or similar, the check must include both storage policies and application logic.

Staging, preview, and real data

One of the most common errors is using real data outside of production. The team copies a dump for debugging, connects a preview deployment to the real database, uses the same keys in staging, imports customers for a demo, or tests an agent on real data. These environments usually have fewer controls, more people with access, and more verbose logs.

The dev/staging/prod separation must cover databases, storage, keys, users, callbacks, backups, and observability tools. If realistic data is needed, it is preferable to use synthetic datasets, masking, or minimized subsets, avoiding having preview URLs, temporary branches, or AI-generated environments point to the real database. When a staging database contains real data, it must be treated as production: limited access, protected backups, logging, retention, patching, and credential control.

Vector databases, RAG, and document leakage

AI apps that use RAG add another form of database: vector stores, embedding stores, document indexes. ChromaDB, Pinecone, pgvector, and other systems can contain fragments of documents, tickets, emails, contracts, manuals, knowledge bases, and customer data. If retrieval does not apply authorization, a user can receive context that belongs to others.

The security of the vector database is not limited to the database engine but requires checking what is indexed, with what metadata, in which namespace or tenant, who can retrieve it, what filters are applied before retrieval, and which documents end up in logs or the model’s prompt. For multi-tenant apps, it is necessary to separate namespaces or apply robust authorization filters, log which documents are retrieved, and plan for deletion and re-indexing when a user deletes data or changes permissions. If the model receives unauthorized documents, the response may have already exposed the data.

AI agents with database access

When an agent can use a terminal, MCP, SQL tools, cloud dashboards, or runtime credentials, the database becomes part of its operational perimeter. The risk is not just that the agent generates bad code: it can execute queries, modify data, run migrations, read tables, or print results in chats and logs.

Do not give an agent production credentials if it is not strictly necessary. It is preferable to use sandbox databases, synthetic data, read-only users, temporary permissions, and approval for destructive commands. If the agent must generate migrations, it is appropriate to review them before applying; if it must query data, it is necessary to limit available tables and rows. For agents connected via MCP or external tools, it is necessary to map what actions are possible — read, write, delete, export, create users, modify schema, access secret managers — and apply the principle of least privilege: the agent must have only what is needed for the task, not everything that is convenient.

What to verify before beta

Before opening a beta or importing real data, it is useful to prepare a data inventory: which databases exist, which environments use them, which users and keys access them, which APIs read or write, which backups are active, which buckets contain files linked to records, which agents or pipelines have credentials.

Then it is necessary to perform practical tests: secret scanning on repositories and history, network control, policy verification, access with minimum-privilege users, API testing on exports and debug routes, file downloads with different users, searching for dumps and backups in artifacts, and checking staging and preview environments. These steps find problems that a demo does not show.

Signs that the database is already too exposed

Some signs deserve immediate attention even if a breach has not yet been found: a committed .env file, a connection string in a log, a database that accepts connections from unrecognized IPs, a backup in a shared bucket, an admin dashboard reachable outside the intended network, a preview deployment with real data, a CSV export with more columns than necessary.

Other signs emerge from app behavior: a list query returns records for all users and then the UI filters client-side, an endpoint accepts limit=100000 and produces full exports, a search function returns cross-tenant results, a database user used by the app can read tables that no feature uses, an AI agent can execute queries on production to help with debugging. These signs are not all incidents, but they indicate that the data perimeter is not governed. Before connecting real users, it is worth stopping and rebuilding the data path: entry, query, API, storage, logs, backups, exports, and deletion.

What to fix first and what to plan

Priority must be given to everything that allows direct access or massive copying of data. Exposed production connection strings, databases reachable from unintended networks, public backups, service keys in the frontend, administrative database users used by the app, unauthorized exports, and open rules on private data must be corrected before go-live.

Remediation must close the risk, not just hide it. If a connection string has leaked, you must rotate the credential and check access logs. If a database was reachable from the internet, you must restrict the network and verify who connected. If a backup was public, you must remove access, rotate any included keys, and evaluate what it contained. If a policy was open, you must add negative tests with different users and tenants.

Other improvements can be planned with owners and deadlines: strengthening variable naming, automating secret scanning, documenting DB users, reducing retention, improving alerts, better separating staging and production. However, the protection of the database and its copies must be treated as a release requirement, not as a subsequent refinement.

How ISGroup can verify databases, storage, and APIs

ISGroup can verify if an app created with AI exposes the database through configurations, credentials, online surfaces, APIs, backups, storage, or cloud. The check starts from the actual way data is reached: network, app, BaaS, storage, pipelines, and agents.

If the project has… Main risk Recommended control
Reachable databases, hosts, services, panels, or endpoints Technical exposures and known configurations Vulnerability Assessment
Web apps, APIs, exports, uploads, or routes that read data Application abuse and unauthorized access Web Application Penetration Testing
Cloud databases, buckets, IAM, security groups, VPCs, BaaS, or backups Cloud misconfiguration or excessive privileges Cloud Security Assessment
Connection strings, queries, policies, DB users, or server-side logic in code Implementation errors and secrets in code Code Review
AI coding used continuously on databases, migrations, and deploys Non-repeatable release controls Software Assurance Lifecycle

The choice of control depends on where the data can exit: reachable database, exploitable API, exposed backup, public bucket, key in code, or agent with excessive permissions. Before beta, it is advisable to delineate these paths and close the most critical exposures.

Have you created an app with AI and are about to connect it to real data? ISGroup can verify databases, storage, APIs, credentials, backups, rules, and configurations before the MVP becomes an operational risk.

Evidence to prepare

To start a verification, it is useful to prepare a list of databases, providers, environments, managed connection strings, DB users, roles, policies, backups, storage, APIs, repositories, pipelines, agents used, and parts generated with AI. If using a BaaS, you must include the Supabase or Firebase project with rules, buckets, and keys; if using a cloud database, include network, security groups, IP allowlists, VPCs, and service accounts.

Test data is also needed: two users, two tenants if present, similar records, uploaded files, exports, backups, and accounts with different privileges. Without these elements, it is difficult to prove if the app truly isolates data and copies. If you have already received alerts about secrets, permissive rules, public databases, open buckets, or shared dumps, you should include them: often a single alert reveals a process that needs to be corrected before go-live.

Decision before go-live

Block the go-live if you find exposed production connection strings, databases reachable from unintended networks, RLS or Security Rules open on private data, overly privileged database users, downloadable backups or dumps, public buckets with private documents, APIs that export data without authorization, or agents with unchecked production credentials.

You can plan improvements with clear residual risk only after release: documentation, additional alerts, progressive reduction of non-critical privileges, key naming, automation of controls. The protection of operational data must be finalized before using customers, payments, documents, or corporate information.

Exposed database checklist

  • Search for connection strings and credentials in code, .env files, history, prompts, logs, and artifacts.
  • Verify IP allowlists, firewalls, security groups, VPCs, and administrative access.
  • Check RLS, Security Rules, ACLs, and policies on tables and storage.
  • Use minimum-privilege DB users, separated by environment.
  • Protect backups, dumps, snapshots, exports, and reports.
  • Test APIs, debug routes, exports, and dashboards that read from the database.
  • Verify buckets, signed URLs, previews, and metadata linked to records.
  • Avoid real data in staging, preview, and agentic environments.
  • Isolate vector databases and retrieval by user or tenant.
  • Limit AI agents and MCP tools with database access.

Frequently Asked Questions

  • Does an exposed database necessarily mean direct public access?
  • No. It can also mean an exposed connection string, downloadable backup, overly permissive API, public bucket, disabled RLS, staging with real data, or vector database without isolation.
  • If the database is not public, am I safe?
  • Not necessarily. Data can leak via APIs, exports, logs, backups, storage, credentials in code, or agents with database access.
  • What is the difference between Vulnerability Assessment and Web Application Penetration Testing in this case?
  • Vulnerability Assessment helps map exposed services and configurations. WAPT tests whether apps and APIs allow reading or modifying data abusively.
  • Do Supabase or Firebase avoid the problem?
  • They offer useful controls, but RLS, Security Rules, buckets, keys, users, and environment separation must be configured and tested on the actual project.
  • Should backups be protected like the database?
  • Yes. They often contain more data than the operational database and are less monitored. Dumps, snapshots, and exports with real data must have controlled access, retention, and storage.
  • What changes with a vector database?
  • Data can be retrieved as model context. If filters, namespaces, or tenant isolation do not work, a user can receive information indexed from others’ documents.

Protect your organisation with Web Application Penetration Testing.

Choose ISGroup for a practical, tailored engagement:

  • A focused assessment of your environment and requirements
  • Clear findings with a prioritised, actionable roadmap
  • Direct support from experienced specialists through remediation and implementation
Talk to an expert

Useful sources and references