Agentic Application Security: LangChain, LlamaIndex, AutoGen, and Risks to Monitor

Sicurezza applicazioni agentiche LangChain LlamaIndex AutoGen

LangChain, LangGraph, LlamaIndex, CrewAI, and AutoGen: Security for Custom Agentic Applications

With LangChain, LangGraph, LlamaIndex, CrewAI, and AutoGen, teams are not just using a coding assistant: they are building applications where an LLM reasons, retrieves context, uses tools, calls APIs, and coordinates steps. The security surface is no longer just that of a web app or an isolated model; it encompasses the agentic runtime, tools, memory, retrieval, and the permissions that govern them.

The point is not to decide whether AI is a good or bad choice for development. The point is much more practical: understanding what controls are needed when a result generated or accelerated by AI enters a product, a business workflow, or an environment with real data. This article is aimed at founders, CTOs, developers, and IT/security teams who are building custom agentic applications and want to understand where to focus their controls: tool abuse, RAG poisoning, memory poisoning, permissions, and unvalidated output.

Why an app that works is not necessarily secure

AI tools reduce the time required to create code, interfaces, workflows, tests, and configurations. However, this speed can compress the steps that normally make software reliable: threat modeling, review, secret management, role controls, input validation, dependency verification, and manual testing of critical paths.

A demo works with a single user, dummy data, and implicit permissions. The same logic can fail when real customers, multiple tenants, different roles, public APIs, integrations, personal data, payments, or automations with external effects are introduced. For this reason, security must be evaluated based on the actual behavior of the app, not on the promise of the tool that generated it.

Tool calling and excessive permissions

Every tool granted to the agent is an application permission in every sense. Read, write, email, tickets, databases, shell, or CRM access must have minimal scope, policy external to the model, audit logs, and explicit confirmation for sensitive actions. Delegating permission management to the prompt is one of the most common and dangerous errors in agentic applications: critical rules must reside in code, policies, and server-side controls, not in the system prompt.

RAG, memory, and untrusted context

Documents, tickets, web pages, emails, and knowledge bases can contain hostile instructions. Retrieval must filter by authorization, and retrieved content must not be able to override system policies or user permissions. The risk of RAG poisoning and memory poisoning between sessions or tenants is real and often underestimated: a malicious document inserted into the knowledge base can alter the agent’s behavior in a way that is undetectable without specific testing.

Output handling and automated decisions

Model output should not be used directly to generate HTML, build queries, execute commands, call APIs, manage authorizations, or make irreversible decisions. Structured validation, schemas, allowlists, sandboxes, and testing with malicious inputs are required before any output reaches a real system.

Key risks to monitor

The risks below are not theoretical: they concern the concrete behavior of the application on real data and must be verified based on evidence, configuration, runtime behavior, and actual impact.

  • Direct and indirect prompt injection: Hostile instructions injected into the prompt or documents retrieved by the system.
  • Tool abuse with excessive permissions: Agents that can perform unauthorized actions because tools do not have limited scope.
  • RAG poisoning and hostile documents: Malicious content in the knowledge base that alters the agent’s behavior.
  • Memory poisoning between sessions or tenants: Contamination of persistent memory that influences subsequent sessions or different users.
  • Sensitive information disclosure: Sensitive data exposed through output, logs, or model responses.
  • Unvalidated output to APIs or shells: Model results used directly as input for external systems without sanitization.
  • Authorizations delegated to the prompt: Access control logic entrusted to the system prompt instead of server-side policies.

These risks must be linked to the concrete perimeter of the application. An exposed app requires manual application testing; a critical code change requires review; an internal workflow requires permission and credential control; an agentic app requires testing on prompts, tools, and output. The correct combination depends on the real impact, not the name of the tool used to build it.

Minimum controls before go-live

  • Map users, roles, real data, integrations, environments, and service owners.
  • Identify which parts were generated or modified with AI and who reviewed them.
  • Verify server-side authorizations, tenant isolation, and administrative functions.
  • Search for secrets in code, prompts, logs, environment variables, builds, and repository history.
  • Check dependencies, licenses, packages, templates, plugins, and generated components.
  • Test for hostile inputs, error handling, logging, rate limits, and unexpected paths.
  • Separate blocking fixes, planned remediation, and accepted residual risk.
  • Repeat testing or retesting after corrections that affect critical flows.

When an independent verification is needed

An independent verification is needed when the app or workflow handles real data, external users, roles, APIs, business integrations, payments, storage, automated workflows, or critical code generated with AI. It is also necessary when the team cannot demonstrate which parts have been reviewed and which controls block regressions or abuse.

For agentic applications, the recommended perimeter includes: AI Application Testing, Code Review, and Secure Architecture Review. The most useful review is not generic: it must produce reproducible findings, remediation priorities, indications of residual risk, and, when necessary, retesting after corrections.

Operational questions for founders, CTOs, and security teams

  • What real data enters the system and where is it saved, logged, or sent?
  • What roles exist and which actions are blocked server-side, not just in the interface?
  • Which secrets, tokens, webhooks, or credentials would allow access to critical systems?
  • Which parts were generated or modified by AI and which were reviewed by a competent person?
  • Which tests cover abuse, errors, different roles, and different tenants, not just the happy path?
  • What evidence can be shown to customers, auditors, procurement, or management?

Useful resources

FAQ

  • What is the difference between an LLM app and an agentic app?
  • An agentic app does not just respond to a prompt: it retrieves context, plans, uses tools, calls APIs, or coordinates steps with external effects on real systems.
  • Can the prompt be a security barrier?
  • No. Critical rules must reside in code, policies, permissions, and server-side controls. The system prompt can guide the model’s behavior, but it is not a reliable security mechanism.
  • How do you test for indirect prompt injection?
  • By inserting hostile instructions into documents, pages, tickets, or records retrieved by the system and verifying tool calls, output, and data access in the agent’s responses.
  • When is AI Application Testing needed?
  • When the app integrates LLMs, RAG, memory, tool calling, agents, or workflows with actions on real systems, especially if it handles sensitive data or external users.
  • Is a Code Review also needed?
  • Yes. You need to read authorizations, tool wrappers, output validation, retrieval, logging, secret management, and operational limits: aspects that an application test alone does not cover.

Protect your organisation with Code Review.

Choose ISGroup for a practical, tailored engagement:

  • A focused assessment of your environment and requirements
  • Clear findings with a prioritised, actionable roadmap
  • Direct support from experienced specialists through remediation and implementation
Talk to an expert

Do not miss the best of cybersecurity.

Weekly expert analysis, real attacks and practical solutions in one newsletter.

Subscribe to Cyber Weekly

Sources and references