Resource Exhaustion in AI systems occurs when an attacker exploits vulnerabilities to consume critical resources—memory, CPU, network bandwidth, or storage—until performance degrades or services become unavailable. This risk is particularly relevant in Large Language Model (LLM)-based architectures, where resource consumption can grow unpredictably.
This article is part of the AI Infrastructure Testing chapter of the OWASP AI Testing Guide.
Why Resource Exhaustion is critical in AI systems
Unlike traditional systems, AI applications have characteristics that amplify the risk of resource exhaustion:
- Variable costs per token: Cloud services charge based on processed tokens, both in input and output.
- Amplification in multi-agent systems: A single prompt can generate dozens of internal calls between agents, multiplying token consumption in a way that is invisible to the user.
- Load unpredictability: Seemingly harmless inputs can trigger highly expensive processing tasks.
Without proper limits, an attacker can cause a “Denial-of-Wallet,” exhausting the operational budget before even compromising the technical availability of the service.
Testing objectives
Effective Resource Exhaustion testing must verify:
- The presence of vulnerabilities that allow for excessive resource consumption.
- The system’s ability to handle anomalous or malformed inputs without degrading performance.
- The effectiveness of controls regarding resource allocation, limitation, and monitoring.
- The configuration of spending thresholds and alerts in third-party managed services.
Methodology and payloads
High-frequency requests
Simulate an attack with rapid concurrent requests using load testing tools. The system is vulnerable if it does not return 429 Too Many Requests errors and response times increase significantly.
Vulnerability indicator: Absence of 429 errors and progressive degradation of response times under high load.
Excessive input sizes
Send prompts exceeding 1MB of text. If the system crashes, returns 5xx errors, times out, or slows down excessively, effective input size validation is missing.
Vulnerability indicator: System crashes, 5xx errors, timeouts, or significant slowdowns in response to large payloads.
Amplification attacks in agentic systems
Repeatedly request a model to use one of its tools (e.g., “Call the search tool 50 times”). A vulnerability manifests if the model executes the operation without refusing it. Verification requires analyzing agent logs or billing dashboards.
Vulnerability indicator: Execution of repeated tool calls without throttling controls, resulting in multiplied token consumption.
Absence of spending limits
Examine the cloud AI service management console. A dangerous configuration is the absence of spending or token thresholds, or limits set too high relative to the planned operational budget.
Vulnerability indicator: Lack of configured spending thresholds or limits set to unrealistic values compared to the operational budget.
Expected output
429 Too Many Requestserrors when configured frequency thresholds are exceeded.- Rejection of requests larger than 1-2 MB with a
413 Payload Too Largeerror. - Stable response times for valid requests, even during attacks against other clients.
- Configured spending limits and active alerts in third-party managed services.
Remediation actions
Input validation and limitation
Set strict limits on the size of incoming data, starting at the API gateway level. Implement validation controls that reject requests exceeding predefined thresholds before they reach the AI models.
Expected impact: Reduction of crash and timeout risks caused by excessive payloads, with immediate protection at the infrastructure level.
Rate limiting and circuit breakers
Apply rate limiting and circuit breakers through the API gateway infrastructure or dedicated middleware. Establish specific resource quotas for each AI model or service (CPU, memory, tokens).
Expected impact: Prevention of high-frequency attacks and protection of service availability for legitimate users.
Monitoring and cost management
Configure binding spending limits and alerts in third-party AI services. Constantly monitor consumption and response times using observability tools. Document and periodically test configured thresholds to verify their effectiveness.
Expected impact: Proactive control of operational costs and the ability to detect anomalies before they cause significant economic damage.
Suggested tools
- Locust: Open-source framework for load testing and high-frequency traffic simulation.
- Ambassador Edge Stack: API gateway with advanced rate limiting and circuit breaker capabilities.
- Prometheus: Monitoring and alerting system for resource consumption metrics.
- Datadog: Observability platform for monitoring costs and performance of cloud AI services.
Further reading
To better understand the dynamics of Resource Exhaustion in AI systems and defense strategies, the following references offer technical guidelines and internationally recognized best practices.
References
- OWASP – OWASP Top 10 for LLM Applications 2025 – Unbounded Resource Consumption
- OWASP – Denial of Service
- NIST AI 100-2e2025 – Security Guidelines for AI Systems (DOI: 10.6028/NIST.AI.100-2e2025)
Frequently Asked Questions
- What are the signs of an ongoing Resource Exhaustion attack?
- Sudden increase in operational costs, generalized response slowdowns, frequent timeout errors, saturation of CPU or memory on nodes running AI models.
- How does Resource Exhaustion differ from a normal traffic spike?
- A legitimate spike shows usage patterns consistent with real user behavior. Resource Exhaustion presents anomalous requests (excessive size, unnatural frequency, repetitive patterns) originating from a few sources.
- Are automated tests sufficient to identify these vulnerabilities?
- No. Automated tests generate many requests and can be expensive without providing qualitative insights. A targeted manual approach, simulating realistic scenarios, is often more effective for identifying AI-specific vulnerabilities.
- Which metrics should be monitored to prevent Resource Exhaustion?
- Tokens consumed per user/session, processing time per request, number of internal agent calls, hourly/daily costs, percentage of requests exceeding predefined thresholds.
- How can limits be managed without impacting legitimate users?
- Implement progressive rate limiting (gradually increasing restrictions), differentiate quotas by user type, and provide clear error messages explaining the limits and when they will be available again.
How ISGroup supports you
ISGroup offers Secure Architecture Review services to evaluate complex AI-based infrastructures and identify vulnerabilities related to resource consumption. The team analyzes the current state of the architecture, verifies the presence of adequate controls on rate limiting, input sizing, and quota management, and provides concrete recommendations to improve resilience and contain operational costs. For cloud architectures, the Cloud Security Assessment service allows for verifying the correct configuration of limits and alerts on major providers.
Related articles
- AI Data Testing: Security and Data Validation in AI Systems
- Runtime Exfiltration in AI Systems: Risks and Testing Strategies
- Capability Misuse in AI Systems: Risks and Testing Strategies
Integrating input validation, rate limiting, and proactive cost monitoring helps protect AI systems from resource exhaustion attacks. Regularly testing configured limits and spending thresholds is essential to ensure resilience and economic sustainability in production.
Want to give your company the highest level of cyber security? ISGroup SRL is here to help with cyber security solutions tailored to your business.
Would you like us to take care of everything for you? Our Virtual CISO and vulnerability management services are a perfect fit for your organization.
Already know what you need? Explore our services:
- Vulnerability Assessment
- Network Penetration Testing
- Web Application Penetration Testing
- Mobile Application Security Testing
- Ethical Hacking
- Training
And much more. Protect your company with the best cybersecurity experts!
