CVE-2026-93436: Missing Release of Memory after Effective Lifetime in vllm-project vllm
vLLM versions up to 0.29.0 have a memory management vulnerability where decode-side metadata for rejected inference requests is not properly released. This allows remote attackers to exhaust decode-worker memory by submitting requests with max_tokens=0, causing the worker to restart. The vulnerability has a high severity with a CVSS score of 8.7.
AI Analysis
Technical Summary
CVE-2026-93436 affects vLLM through version 0.29.0. The software fails to properly clean up decode-side metadata associated with rejected inference requests in prefill/decode disaggregated deployments. An attacker can exploit this by sending requests with max_tokens set to 0, which causes unbounded memory consumption in the decode worker process until it restarts, resulting in a denial of service condition.
Potential Impact
Remote attackers can cause a denial of service by exhausting the memory of decode-worker processes through crafted inference requests with max_tokens=0. This leads to worker restarts and potential service disruption.
Mitigation Recommendations
Patch status is not yet confirmed — check the vendor advisory for current remediation guidance.
CVE-2026-93436: Missing Release of Memory after Effective Lifetime in vllm-project vllm
Description
vLLM versions up to 0.29.0 have a memory management vulnerability where decode-side metadata for rejected inference requests is not properly released. This allows remote attackers to exhaust decode-worker memory by submitting requests with max_tokens=0, causing the worker to restart. The vulnerability has a high severity with a CVSS score of 8.7.
CVSS v4.0
Score 8.7high
Affected software
vllm-project
vllm
pkg:github/vllm-project/vllmRun on your own infrastructure? Check whether these packages are installed with threat-finder — our free open-source scanner.
AI-Powered Analysis
Machine-generated threat intelligence
Technical Analysis
CVE-2026-93436 affects vLLM through version 0.29.0. The software fails to properly clean up decode-side metadata associated with rejected inference requests in prefill/decode disaggregated deployments. An attacker can exploit this by sending requests with max_tokens set to 0, which causes unbounded memory consumption in the decode worker process until it restarts, resulting in a denial of service condition.
Potential Impact
Remote attackers can cause a denial of service by exhausting the memory of decode-worker processes through crafted inference requests with max_tokens=0. This leads to worker restarts and potential service disruption.
Mitigation Recommendations
Patch status is not yet confirmed — check the vendor advisory for current remediation guidance.
Technical Details
- Data Version
- 5.2
- Assigner Short Name
- VulnCheck
- Date Reserved
- 2026-09-17T21:50:02.601Z
- Cvss Version
- 4.0
- State
- PUBLISHED
Threat ID: 6aac6e1955bf5e2cf506c99a
Added to database: 09/17/2026, 22:47:53 UTC
Last enriched: 09/17/2026, 23:01:27 UTC
Last updated: 09/17/2026, 23:01:27 UTC
Views: 4
Community Reviews
0 reviewsCrowdsource mitigation strategies, share intel context, and vote on the most helpful responses. Sign in to add your voice and help keep defenders ahead.
Want to contribute mitigation steps or threat intel context? Sign in or create an account to join the community discussion.
Actions
Updates to AI analysis require Pro Console access. Upgrade inside Console → Billing.
Need more coverage?
Upgrade to Pro Console for AI refresh and higher limits.
For incident response and remediation, OffSeq services can help resolve threats faster.
Latest Threats
Check if your credentials are on the dark web
Instant breach scanning across billions of leaked records. Free tier available.