CVE-2026-100652: Improper Input Validation in vllm-project vllm
vLLM versions 0.22.0 through 0.23.0 fail to validate stop_token_ids against vocabulary bounds in Rust HTTP and gRPC frontends, allowing out-of-vocabulary token IDs to reach MinTokensLogitsProcessor. Attackers can submit requests with min_tokens greater than zero and out-of-vocabulary stop_token_ids to trigger CUDA tensor indexing failures that leave EngineCore in a fatal state requiring service restart.
AI Analysis
Technical Summary
vLLM versions 0.22.0 through 0.23.0 fail to validate stop_token_ids against vocabulary bounds in Rust HTTP and gRPC frontends. Attackers can submit requests with min_tokens greater than zero and out-of-vocabulary stop_token_ids, triggering CUDA tensor indexing failures that cause the EngineCore to enter a fatal state, necessitating a service restart.
Potential Impact
Successful exploitation causes CUDA tensor indexing failures that leave the EngineCore in a fatal state, resulting in service disruption that requires a restart. This impacts availability of the vLLM service.
Mitigation Recommendations
Patch status is not yet confirmed — check the vendor advisory for current remediation guidance.
CVE-2026-100652: Improper Input Validation in vllm-project vllm
Description
vLLM versions 0.22.0 through 0.23.0 fail to validate stop_token_ids against vocabulary bounds in Rust HTTP and gRPC frontends, allowing out-of-vocabulary token IDs to reach MinTokensLogitsProcessor. Attackers can submit requests with min_tokens greater than zero and out-of-vocabulary stop_token_ids to trigger CUDA tensor indexing failures that leave EngineCore in a fatal state requiring service restart.
CVSS v4.0
Score 8.2high
Affected software
vllm-project
vllm
pkg:cargo/github/vllm-project/vllmRun on your own infrastructure? Check whether these packages are installed with threat-finder — our free open-source scanner.
AI-Powered Analysis
Machine-generated threat intelligence
Technical Analysis
vLLM versions 0.22.0 through 0.23.0 fail to validate stop_token_ids against vocabulary bounds in Rust HTTP and gRPC frontends. Attackers can submit requests with min_tokens greater than zero and out-of-vocabulary stop_token_ids, triggering CUDA tensor indexing failures that cause the EngineCore to enter a fatal state, necessitating a service restart.
Potential Impact
Successful exploitation causes CUDA tensor indexing failures that leave the EngineCore in a fatal state, resulting in service disruption that requires a restart. This impacts availability of the vLLM service.
Mitigation Recommendations
Patch status is not yet confirmed — check the vendor advisory for current remediation guidance.
Technical Details
- Data Version
- 5.2
- Assigner Short Name
- VulnCheck
- Date Reserved
- 2026-09-26T02:33:07.899Z
- Cvss Version
- 4.0
- State
- PUBLISHED
Threat ID: 6ab7c9a9f7a7c5410652fd3b
Added to database: 09/26/2026, 13:33:29 UTC
Last enriched: 09/26/2026, 14:02:50 UTC
Last updated: 09/27/2026, 01:57:11 UTC
Views: 11
Community Reviews
0 reviewsCrowdsource mitigation strategies, share intel context, and vote on the most helpful responses. Sign in to add your voice and help keep defenders ahead.
Want to contribute mitigation steps or threat intel context? Sign in or create an account to join the community discussion.
Actions
Updates to AI analysis require Pro Console access. Upgrade inside Console → Billing.
Need more coverage?
Upgrade to Pro Console for AI refresh and higher limits.
For incident response and remediation, OffSeq services can help resolve threats faster.
Latest Threats
Check if your credentials are on the dark web
Instant breach scanning across billions of leaked records. Free tier available.