CVE-2026-100650: Uncontrolled Resource Consumption in vllm-project vllm
vLLM through 0.29.0 fetches and fully materializes remote or inline media before enforcing its documented media controls (the VLLM_MAX_AUDIO_CLIP_FILESIZE_MB compressed-audio size cap, default 25 MB, and the per-modality --limit-mm-per-prompt item limits). Across four ingress paths — the shared media-acquisition layer (HTTPConnection.get_bytes()/async_get_bytes()), the chat completions audio_url/base64 path, the batch speech runner, and the Rust frontend POST /tokenize route — the server reads the entire HTTP response body, base64-decodes the inline payload, or spawns one fetch/decode task per media part, and only then applies the limit (or, on some paths, never applies it). A remote attacker can therefore cause the API server or batch-runner process to allocate memory and consume outbound bandwidth proportional to an attacker-chosen body size or media item count before the request is rejected, resulting in pre-inference memory and bandwidth exhaustion (denial of service). The chat and batch surfaces require an API key when one is configured; the Rust frontend /tokenize route is unauthenticated by design. There is no code execution or data disclosure impact.
AI Analysis
Technical Summary
vLLM versions prior to 0.29.0 have a flaw in media handling where the server fetches and fully materializes remote or inline media before applying documented media size and count limits. This occurs across four ingress paths: the shared media-acquisition layer (HTTPConnection.get_bytes()/async_get_bytes()), the chat completions audio_url/base64 path, the batch speech runner, and the Rust frontend POST /tokenize route. Because limits are enforced only after full data retrieval or sometimes not at all, a remote attacker can cause the server to allocate memory and consume outbound bandwidth proportional to attacker-controlled media size or count. This results in denial of service through pre-inference memory and bandwidth exhaustion. The chat and batch interfaces require API keys if configured, but the Rust frontend /tokenize route is unauthenticated by design. There is no impact on code execution or data confidentiality.
Potential Impact
An attacker can cause denial of service by exhausting memory and bandwidth resources on the vllm server or batch-runner process before media size limits are enforced or the request is rejected. This can degrade or disrupt service availability. There is no impact on code execution or data disclosure.
Mitigation Recommendations
Patch status is not yet confirmed — check the vendor advisory for current remediation guidance. Until a fix is available, restrict access to the unauthenticated Rust frontend /tokenize route if possible and monitor for abnormal resource usage. The chat and batch surfaces require API keys when configured, which limits exposure. No code execution or data disclosure risk reduces urgency but resource exhaustion risk remains.
CVE-2026-100650: Uncontrolled Resource Consumption in vllm-project vllm
Description
vLLM through 0.29.0 fetches and fully materializes remote or inline media before enforcing its documented media controls (the VLLM_MAX_AUDIO_CLIP_FILESIZE_MB compressed-audio size cap, default 25 MB, and the per-modality --limit-mm-per-prompt item limits). Across four ingress paths — the shared media-acquisition layer (HTTPConnection.get_bytes()/async_get_bytes()), the chat completions audio_url/base64 path, the batch speech runner, and the Rust frontend POST /tokenize route — the server reads the entire HTTP response body, base64-decodes the inline payload, or spawns one fetch/decode task per media part, and only then applies the limit (or, on some paths, never applies it). A remote attacker can therefore cause the API server or batch-runner process to allocate memory and consume outbound bandwidth proportional to an attacker-chosen body size or media item count before the request is rejected, resulting in pre-inference memory and bandwidth exhaustion (denial of service). The chat and batch surfaces require an API key when one is configured; the Rust frontend /tokenize route is unauthenticated by design. There is no code execution or data disclosure impact.
CVSS v4.0
Score 7.1high
Affected software
vllm-project
vllm
pkg:cargo/github/vllm-project/vllmRun on your own infrastructure? Check whether these packages are installed with threat-finder — our free open-source scanner.
AI-Powered Analysis
Machine-generated threat intelligence
Technical Analysis
vLLM versions prior to 0.29.0 have a flaw in media handling where the server fetches and fully materializes remote or inline media before applying documented media size and count limits. This occurs across four ingress paths: the shared media-acquisition layer (HTTPConnection.get_bytes()/async_get_bytes()), the chat completions audio_url/base64 path, the batch speech runner, and the Rust frontend POST /tokenize route. Because limits are enforced only after full data retrieval or sometimes not at all, a remote attacker can cause the server to allocate memory and consume outbound bandwidth proportional to attacker-controlled media size or count. This results in denial of service through pre-inference memory and bandwidth exhaustion. The chat and batch interfaces require API keys if configured, but the Rust frontend /tokenize route is unauthenticated by design. There is no impact on code execution or data confidentiality.
Potential Impact
An attacker can cause denial of service by exhausting memory and bandwidth resources on the vllm server or batch-runner process before media size limits are enforced or the request is rejected. This can degrade or disrupt service availability. There is no impact on code execution or data disclosure.
Mitigation Recommendations
Patch status is not yet confirmed — check the vendor advisory for current remediation guidance. Until a fix is available, restrict access to the unauthenticated Rust frontend /tokenize route if possible and monitor for abnormal resource usage. The chat and batch surfaces require API keys when configured, which limits exposure. No code execution or data disclosure risk reduces urgency but resource exhaustion risk remains.
Technical Details
- Data Version
- 5.2
- Assigner Short Name
- VulnCheck
- Date Reserved
- 2026-09-26T02:33:07.898Z
- Cvss Version
- 4.0
- State
- PUBLISHED
Threat ID: 6ab7c9a7f7a7c5410652fd32
Added to database: 09/26/2026, 13:33:27 UTC
Last enriched: 09/26/2026, 14:03:01 UTC
Last updated: 09/27/2026, 01:57:11 UTC
Views: 11
Community Reviews
0 reviewsCrowdsource mitigation strategies, share intel context, and vote on the most helpful responses. Sign in to add your voice and help keep defenders ahead.
Want to contribute mitigation steps or threat intel context? Sign in or create an account to join the community discussion.
Actions
Updates to AI analysis require Pro Console access. Upgrade inside Console → Billing.
Need more coverage?
Upgrade to Pro Console for AI refresh and higher limits.
For incident response and remediation, OffSeq services can help resolve threats faster.
Latest Threats
Check if your credentials are on the dark web
Instant breach scanning across billions of leaked records. Free tier available.