CVE-2026-73559: CWE-400: Uncontrolled Resource Consumption in vllm-project vllm
vLLM is an inference and serving engine for large language models. From 0.19.0 until 0.26.0, the /v1/completions CompletionRequest.prompt field in vllm/entrypoints/openai/completion/protocol.py accepts an unbounded list[str] or list[list[int]], prompt_to_seq() in vllm/renderers/inputs/preprocess.py and OnlineRenderer.preprocess_completion() in vllm/renderers/online_renderer.py expand every element, and vllm/entrypoints/openai/completion/serving.py creates one engine generator and response slot per prompt, allowing an authenticated API client to exhaust CPU, memory, async scheduling capacity, engine request slots, and response buffering with one request. This issue is fixed in version 0.26.0.
AI Analysis
Technical Summary
The vulnerability arises because the /v1/completions CompletionRequest.prompt field accepts an unbounded list of strings or lists of integers. Functions prompt_to_seq() and OnlineRenderer.preprocess_completion() expand every element in the prompt list, and the serving code creates one engine generator and response slot per prompt. This design flaw allows an authenticated API client to exhaust critical system resources with a single request, leading to denial of service. The flaw is identified as CWE-400: Uncontrolled Resource Consumption. The vulnerability affects vllm versions from 0.19.0 until 0.26.0 and is resolved in version 0.26.0.
Potential Impact
An authenticated API client can cause denial of service by exhausting CPU, memory, asynchronous scheduling capacity, engine request slots, and response buffering through a single crafted request with a large number of prompts. There is no impact on confidentiality or integrity reported. The impact is availability degradation due to resource exhaustion.
Mitigation Recommendations
Upgrade to vllm version 0.26.0 or later where this vulnerability is fixed. Patch status is confirmed by the vendor advisory indicating the fix in 0.26.0. Until upgrading, restrict or monitor authenticated API client usage to prevent abuse of the /v1/completions endpoint with large prompt lists.
CVE-2026-73559: CWE-400: Uncontrolled Resource Consumption in vllm-project vllm
Description
vLLM is an inference and serving engine for large language models. From 0.19.0 until 0.26.0, the /v1/completions CompletionRequest.prompt field in vllm/entrypoints/openai/completion/protocol.py accepts an unbounded list[str] or list[list[int]], prompt_to_seq() in vllm/renderers/inputs/preprocess.py and OnlineRenderer.preprocess_completion() in vllm/renderers/online_renderer.py expand every element, and vllm/entrypoints/openai/completion/serving.py creates one engine generator and response slot per prompt, allowing an authenticated API client to exhaust CPU, memory, async scheduling capacity, engine request slots, and response buffering with one request. This issue is fixed in version 0.26.0.
CVSS v3.1
Score 6.5medium
Affected software
Weaknesses
AI-Powered Analysis
Machine-generated threat intelligence
Technical Analysis
The vulnerability arises because the /v1/completions CompletionRequest.prompt field accepts an unbounded list of strings or lists of integers. Functions prompt_to_seq() and OnlineRenderer.preprocess_completion() expand every element in the prompt list, and the serving code creates one engine generator and response slot per prompt. This design flaw allows an authenticated API client to exhaust critical system resources with a single request, leading to denial of service. The flaw is identified as CWE-400: Uncontrolled Resource Consumption. The vulnerability affects vllm versions from 0.19.0 until 0.26.0 and is resolved in version 0.26.0.
Potential Impact
An authenticated API client can cause denial of service by exhausting CPU, memory, asynchronous scheduling capacity, engine request slots, and response buffering through a single crafted request with a large number of prompts. There is no impact on confidentiality or integrity reported. The impact is availability degradation due to resource exhaustion.
Mitigation Recommendations
Upgrade to vllm version 0.26.0 or later where this vulnerability is fixed. Patch status is confirmed by the vendor advisory indicating the fix in 0.26.0. Until upgrading, restrict or monitor authenticated API client usage to prevent abuse of the /v1/completions endpoint with large prompt lists.
Technical Details
- Data Version
- 5.2
- Assigner Short Name
- GitHub_M
- Date Reserved
- 2026-08-12T20:53:46.380Z
- Cvss Version
- 3.1
- State
- PUBLISHED
- Remediation Level
- null
Threat ID: 6a7de5d0bf8831d5396651b2
Added to database: 08/13/2026, 15:42:08 UTC
Last enriched: 08/13/2026, 15:58:49 UTC
Last updated: 08/13/2026, 22:50:00 UTC
Views: 7
Community Reviews
0 reviewsCrowdsource mitigation strategies, share intel context, and vote on the most helpful responses. Sign in to add your voice and help keep defenders ahead.
Want to contribute mitigation steps or threat intel context? Sign in or create an account to join the community discussion.
Actions
Updates to AI analysis require Pro Console access. Upgrade inside Console → Billing.
Need more coverage?
Upgrade to Pro Console for AI refresh and higher limits.
For incident response and remediation, OffSeq services can help resolve threats faster.
Latest Threats
Check if your credentials are on the dark web
Instant breach scanning across billions of leaked records. Free tier available.