CVE-2026-100654: Improper Validation of Array Index in vllm-project vllm
vLLM before 0.29.0 accepts user-controlled stop_token_ids on the OpenAI-compatible POST /v1/completions and POST /v1/chat/completions endpoints but validates only that the values are integers, not that each token id is within the model vocabulary/logits range. When min_tokens > 0, the stop token ids are used as logits indices to suppress stop tokens, so an out-of-range id reaches a CUDA indexing operation (index_put_) and triggers a device-side assertion. An authenticated API user can send a single malformed completion request that returns 500 Internal Server Error and puts EngineCore into a fatal state, causing subsequent requests to fail until the service is restarted (denial of service).
AI Analysis
Technical Summary
vLLM before version 0.29.0 accepts user-controlled stop_token_ids on the OpenAI-compatible POST /v1/completions and POST /v1/chat/completions endpoints. The validation only ensures the values are integers but does not verify that each token ID is within the model vocabulary or logits range. When the min_tokens parameter is greater than zero, these stop token IDs are used as indices into logits to suppress stop tokens. An out-of-range token ID causes a CUDA device-side assertion failure during an index_put_ operation. An authenticated API user can exploit this by sending a single malformed completion request that returns a 500 Internal Server Error and puts the EngineCore into a fatal state, causing subsequent requests to fail until the service is restarted, effectively resulting in a denial of service.
Potential Impact
An authenticated user can cause a denial of service by sending a single malformed request with out-of-range stop_token_ids. This triggers a device-side assertion failure in CUDA, crashing the EngineCore and making the service unavailable until manually restarted. There is no indication of data compromise or code execution beyond service disruption.
Mitigation Recommendations
No explicit patch or remediation is provided in the available data. Patch status is not yet confirmed — check the vendor advisory for current remediation guidance. Until a fix is available, avoid using vulnerable versions or restrict access to trusted users to prevent exploitation.
CVE-2026-100654: Improper Validation of Array Index in vllm-project vllm
Description
vLLM before 0.29.0 accepts user-controlled stop_token_ids on the OpenAI-compatible POST /v1/completions and POST /v1/chat/completions endpoints but validates only that the values are integers, not that each token id is within the model vocabulary/logits range. When min_tokens > 0, the stop token ids are used as logits indices to suppress stop tokens, so an out-of-range id reaches a CUDA indexing operation (index_put_) and triggers a device-side assertion. An authenticated API user can send a single malformed completion request that returns 500 Internal Server Error and puts EngineCore into a fatal state, causing subsequent requests to fail until the service is restarted (denial of service).
CVSS v4.0
Score 7.1high
Affected software
vllm-project
vllm
Run on your own infrastructure? Check whether these packages are installed with threat-finder — our free open-source scanner.
AI-Powered Analysis
Machine-generated threat intelligence
Technical Analysis
vLLM before version 0.29.0 accepts user-controlled stop_token_ids on the OpenAI-compatible POST /v1/completions and POST /v1/chat/completions endpoints. The validation only ensures the values are integers but does not verify that each token ID is within the model vocabulary or logits range. When the min_tokens parameter is greater than zero, these stop token IDs are used as indices into logits to suppress stop tokens. An out-of-range token ID causes a CUDA device-side assertion failure during an index_put_ operation. An authenticated API user can exploit this by sending a single malformed completion request that returns a 500 Internal Server Error and puts the EngineCore into a fatal state, causing subsequent requests to fail until the service is restarted, effectively resulting in a denial of service.
Potential Impact
An authenticated user can cause a denial of service by sending a single malformed request with out-of-range stop_token_ids. This triggers a device-side assertion failure in CUDA, crashing the EngineCore and making the service unavailable until manually restarted. There is no indication of data compromise or code execution beyond service disruption.
Mitigation Recommendations
No explicit patch or remediation is provided in the available data. Patch status is not yet confirmed — check the vendor advisory for current remediation guidance. Until a fix is available, avoid using vulnerable versions or restrict access to trusted users to prevent exploitation.
Technical Details
- Data Version
- 5.2
- Assigner Short Name
- VulnCheck
- Date Reserved
- 2026-09-26T02:33:07.899Z
- Cvss Version
- 4.0
- State
- PUBLISHED
Threat ID: 6ab7c9a9f7a7c5410652fd3d
Added to database: 09/26/2026, 13:33:29 UTC
Last enriched: 09/26/2026, 14:02:42 UTC
Last updated: 09/27/2026, 04:31:28 UTC
Views: 11
Community Reviews
0 reviewsCrowdsource mitigation strategies, share intel context, and vote on the most helpful responses. Sign in to add your voice and help keep defenders ahead.
Want to contribute mitigation steps or threat intel context? Sign in or create an account to join the community discussion.
Actions
Updates to AI analysis require Pro Console access. Upgrade inside Console → Billing.
Need more coverage?
Upgrade to Pro Console for AI refresh and higher limits.
For incident response and remediation, OffSeq services can help resolve threats faster.
Latest Threats
Check if your credentials are on the dark web
Instant breach scanning across billions of leaked records. Free tier available.