CVE-2026-71486: CWE-400: Uncontrolled Resource Consumption in vllm-project vllm
CVE-2026-71486 is a medium severity vulnerability in vllm-project vllm versions prior to 0.26.0. It affects the /v1/completions/derender and /v1/chat/completions/derender API endpoints, which accept client-supplied generated output data without enforcing limits on token counts or response sizes. This allows an authenticated API client to cause uncontrolled CPU and memory consumption proportional to the size of the attacker-supplied token data, bypassing normal generation output bounds.
AI Analysis
Technical Summary
The vulnerability exists because the derender endpoints in vllm accept GenerateResponse objects containing nested token ID lists from clients and postprocess them without enforcing model context length, max tokens, choice counts, or response size limits. Unlike normal generation paths that enforce these bounds, derender directly decodes and returns the supplied token IDs, causing CPU and memory usage to scale with attacker-controlled input size. The API routes are protected by API key middleware but do not validate or limit the size of the nested token ID arrays. This can lead to uncontrolled resource consumption (CWE-400) and potentially denial of service conditions.
Potential Impact
An authenticated API client can cause excessive CPU and memory usage on the server by submitting large or deeply nested generated output payloads to the derender endpoints. This can degrade service availability or cause denial of service due to resource exhaustion. There is no impact on confidentiality or integrity. The vulnerability requires authentication and does not affect normal generation endpoints that enforce output size limits.
Mitigation Recommendations
No official patch or fix is currently documented. Patch status is not yet confirmed — check the vendor advisory for current remediation guidance. Until a fix is available, restrict access to the derender endpoints to trusted clients only and monitor usage patterns for unusually large or complex derender requests. Implementing rate limiting or input validation on the size of token ID lists before processing could mitigate exploitation risk.
CVE-2026-71486: CWE-400: Uncontrolled Resource Consumption in vllm-project vllm
Description
CVE-2026-71486 is a medium severity vulnerability in vllm-project vllm versions prior to 0.26.0. It affects the /v1/completions/derender and /v1/chat/completions/derender API endpoints, which accept client-supplied generated output data without enforcing limits on token counts or response sizes. This allows an authenticated API client to cause uncontrolled CPU and memory consumption proportional to the size of the attacker-supplied token data, bypassing normal generation output bounds.
CVSS v3.1
Score 4.3medium
Affected software
vllm-project
vllm
Run on your own infrastructure? Check whether these packages are installed with threat-finder — our free open-source scanner.
AI-Powered Analysis
Machine-generated threat intelligence
Technical Analysis
The vulnerability exists because the derender endpoints in vllm accept GenerateResponse objects containing nested token ID lists from clients and postprocess them without enforcing model context length, max tokens, choice counts, or response size limits. Unlike normal generation paths that enforce these bounds, derender directly decodes and returns the supplied token IDs, causing CPU and memory usage to scale with attacker-controlled input size. The API routes are protected by API key middleware but do not validate or limit the size of the nested token ID arrays. This can lead to uncontrolled resource consumption (CWE-400) and potentially denial of service conditions.
Potential Impact
An authenticated API client can cause excessive CPU and memory usage on the server by submitting large or deeply nested generated output payloads to the derender endpoints. This can degrade service availability or cause denial of service due to resource exhaustion. There is no impact on confidentiality or integrity. The vulnerability requires authentication and does not affect normal generation endpoints that enforce output size limits.
Mitigation Recommendations
No official patch or fix is currently documented. Patch status is not yet confirmed — check the vendor advisory for current remediation guidance. Until a fix is available, restrict access to the derender endpoints to trusted clients only and monitor usage patterns for unusually large or complex derender requests. Implementing rate limiting or input validation on the size of token ID lists before processing could mitigate exploitation risk.
Technical Details
- Data Version
- 5.2
- Assigner Short Name
- GitHub_M
- Date Reserved
- 2026-08-06T19:56:23.725Z
- Cvss Version
- 3.1
- State
- PUBLISHED
Threat ID: 6a836e8fbf8831d5398572cf
Added to database: 08/17/2026, 20:26:55 UTC
Last enriched: 09/12/2026, 00:33:10 UTC
Last updated: 10/02/2026, 02:46:06 UTC
Views: 87
Community Reviews
0 reviewsCrowdsource mitigation strategies, share intel context, and vote on the most helpful responses. Sign in to add your voice and help keep defenders ahead.
Want to contribute mitigation steps or threat intel context? Sign in or create an account to join the community discussion.
Actions
Updates to AI analysis require Pro Console access. Upgrade inside Console → Billing.
Need more coverage?
Upgrade to Pro Console for AI refresh and higher limits.
For incident response and remediation, OffSeq services can help resolve threats faster.
Latest Threats
Check if your credentials are on the dark web
Instant breach scanning across billions of leaked records. Free tier available.