CVE-2026-100651: Uncontrolled Resource Consumption in vllm-project vllm
vLLM before 0.29.0 fails to enforce decoder prompt-length validation on the disaggregated serving endpoint /inference/v1/generate. When the request contains a 'features' (multimodal) payload, vllm/entrypoints/serve/disagg/serving.py builds a multimodal EngineInput directly from the caller-supplied token_ids, and GenerateRequest.token_ids (vllm/entrypoints/serve/disagg/protocol.py) is not checked against model_config.max_model_len. For multimodal processors that report skip_prompt_length_check=True (for example Nemotron Parse, Whisper, and FireRedLID), InputProcessor._validate_prompt_len() returns immediately for both encoder and decoder prompts, so an overlong prompt becomes an EngineCoreRequest and reaches the worker input-batch copy into a fixed max_model_len-wide NumPy row. A client able to reach the endpoint on an affected model configuration can therefore submit an overlong token_ids list to trigger a worker failure and denial of service. Fixed in 0.29.0.
AI Analysis
Technical Summary
vLLM prior to version 0.29.0 fails to validate the length of decoder prompts on the disaggregated serving endpoint /inference/v1/generate when handling multimodal 'features' payloads. Specifically, for multimodal processors that skip prompt length checks, the system builds EngineInput directly from caller-supplied token_ids without verifying against model_config.max_model_len. This leads to an overlong prompt being processed and copied into a fixed-size NumPy array, causing worker failure and denial of service. The vulnerability is addressed in vLLM 0.29.0.
Potential Impact
An attacker with access to the affected endpoint can submit an overlong token_ids list, triggering a worker failure and causing denial of service. This impacts availability but does not indicate confidentiality or integrity compromise. The CVSS 4.0 score is 7.1 (high severity), reflecting network attack vector, low attack complexity, no privileges required, no user interaction, and high impact on availability.
Mitigation Recommendations
Upgrade to vLLM version 0.29.0 or later, where this vulnerability is fixed. No other mitigations are indicated or necessary as the fix is official and available.
CVE-2026-100651: Uncontrolled Resource Consumption in vllm-project vllm
Description
vLLM before 0.29.0 fails to enforce decoder prompt-length validation on the disaggregated serving endpoint /inference/v1/generate. When the request contains a 'features' (multimodal) payload, vllm/entrypoints/serve/disagg/serving.py builds a multimodal EngineInput directly from the caller-supplied token_ids, and GenerateRequest.token_ids (vllm/entrypoints/serve/disagg/protocol.py) is not checked against model_config.max_model_len. For multimodal processors that report skip_prompt_length_check=True (for example Nemotron Parse, Whisper, and FireRedLID), InputProcessor._validate_prompt_len() returns immediately for both encoder and decoder prompts, so an overlong prompt becomes an EngineCoreRequest and reaches the worker input-batch copy into a fixed max_model_len-wide NumPy row. A client able to reach the endpoint on an affected model configuration can therefore submit an overlong token_ids list to trigger a worker failure and denial of service. Fixed in 0.29.0.
CVSS v4.0
Score 7.1high
Affected software
vllm-project
vllm
pkg:github/vllm-project/vllmRun on your own infrastructure? Check whether these packages are installed with threat-finder — our free open-source scanner.
AI-Powered Analysis
Machine-generated threat intelligence
Technical Analysis
vLLM prior to version 0.29.0 fails to validate the length of decoder prompts on the disaggregated serving endpoint /inference/v1/generate when handling multimodal 'features' payloads. Specifically, for multimodal processors that skip prompt length checks, the system builds EngineInput directly from caller-supplied token_ids without verifying against model_config.max_model_len. This leads to an overlong prompt being processed and copied into a fixed-size NumPy array, causing worker failure and denial of service. The vulnerability is addressed in vLLM 0.29.0.
Potential Impact
An attacker with access to the affected endpoint can submit an overlong token_ids list, triggering a worker failure and causing denial of service. This impacts availability but does not indicate confidentiality or integrity compromise. The CVSS 4.0 score is 7.1 (high severity), reflecting network attack vector, low attack complexity, no privileges required, no user interaction, and high impact on availability.
Mitigation Recommendations
Upgrade to vLLM version 0.29.0 or later, where this vulnerability is fixed. No other mitigations are indicated or necessary as the fix is official and available.
Technical Details
- Data Version
- 5.2
- Assigner Short Name
- VulnCheck
- Date Reserved
- 2026-09-26T02:33:07.899Z
- Cvss Version
- 4.0
- State
- PUBLISHED
Threat ID: 6ab7c9a9f7a7c5410652fd3a
Added to database: 09/26/2026, 13:33:29 UTC
Last enriched: 09/26/2026, 14:02:54 UTC
Last updated: 09/27/2026, 01:57:11 UTC
Views: 11
Community Reviews
0 reviewsCrowdsource mitigation strategies, share intel context, and vote on the most helpful responses. Sign in to add your voice and help keep defenders ahead.
Want to contribute mitigation steps or threat intel context? Sign in or create an account to join the community discussion.
Actions
Updates to AI analysis require Pro Console access. Upgrade inside Console → Billing.
Need more coverage?
Upgrade to Pro Console for AI refresh and higher limits.
For incident response and remediation, OffSeq services can help resolve threats faster.
Latest Threats
Check if your credentials are on the dark web
Instant breach scanning across billions of leaked records. Free tier available.