CVE-2026-94627: Missing Release of Memory after Effective Lifetime in vllm-project vllm
CVE-2026-94627 is a high-severity vulnerability in vllm up to version 0.29.0 where the Mooncake connector improperly manages GPU KV cache block ownership during concurrent child requests sharing a transfer ID. This leads to GPU memory exhaustion as orphaned KV cache blocks accumulate until the process restarts, preventing legitimate requests from executing.
AI Analysis
Technical Summary
The vLLM Mooncake connector through version 0.29.0 fails to properly release GPU KV cache blocks when concurrent child requests share a single transfer ID in prefill/decode disaggregated deployments. Attackers can exploit this by submitting completion requests with multiple prompts, causing orphaned KV cache blocks to accumulate and exhaust GPU memory. This condition persists until the process is restarted, effectively denying service to legitimate requests.
Potential Impact
Successful exploitation results in GPU memory exhaustion, causing denial of service by preventing legitimate requests from executing until the affected process is restarted. There is no indication of privilege escalation, data leakage, or code execution from the provided data.
Mitigation Recommendations
Patch status is not yet confirmed — check the vendor advisory for current remediation guidance. Until a fix is available, consider limiting concurrent child requests sharing a single transfer ID or restarting the process periodically to clear orphaned KV cache blocks.
CVE-2026-94627: Missing Release of Memory after Effective Lifetime in vllm-project vllm
Description
CVE-2026-94627 is a high-severity vulnerability in vllm up to version 0.29.0 where the Mooncake connector improperly manages GPU KV cache block ownership during concurrent child requests sharing a transfer ID. This leads to GPU memory exhaustion as orphaned KV cache blocks accumulate until the process restarts, preventing legitimate requests from executing.
CVSS v4.0
Score 8.7high
Affected software
vllm-project
vllm
pkg:github/vllm-project/vllmRun on your own infrastructure? Check whether these packages are installed with threat-finder — our free open-source scanner.
AI-Powered Analysis
Machine-generated threat intelligence
Technical Analysis
The vLLM Mooncake connector through version 0.29.0 fails to properly release GPU KV cache blocks when concurrent child requests share a single transfer ID in prefill/decode disaggregated deployments. Attackers can exploit this by submitting completion requests with multiple prompts, causing orphaned KV cache blocks to accumulate and exhaust GPU memory. This condition persists until the process is restarted, effectively denying service to legitimate requests.
Potential Impact
Successful exploitation results in GPU memory exhaustion, causing denial of service by preventing legitimate requests from executing until the affected process is restarted. There is no indication of privilege escalation, data leakage, or code execution from the provided data.
Mitigation Recommendations
Patch status is not yet confirmed — check the vendor advisory for current remediation guidance. Until a fix is available, consider limiting concurrent child requests sharing a single transfer ID or restarting the process periodically to clear orphaned KV cache blocks.
Technical Details
- Data Version
- 5.2
- Assigner Short Name
- VulnCheck
- Date Reserved
- 2026-09-21T21:42:30.080Z
- Cvss Version
- 4.0
- State
- PUBLISHED
Threat ID: 6ab1acdf55bf5e2cf58d35c5
Added to database: 09/21/2026, 22:17:03 UTC
Last enriched: 09/21/2026, 22:31:29 UTC
Last updated: 09/21/2026, 22:36:27 UTC
Views: 5
Community Reviews
0 reviewsCrowdsource mitigation strategies, share intel context, and vote on the most helpful responses. Sign in to add your voice and help keep defenders ahead.
Want to contribute mitigation steps or threat intel context? Sign in or create an account to join the community discussion.
Actions
Updates to AI analysis require Pro Console access. Upgrade inside Console → Billing.
Need more coverage?
Upgrade to Pro Console for AI refresh and higher limits.
For incident response and remediation, OffSeq services can help resolve threats faster.
Latest Threats
Check if your credentials are on the dark web
Instant breach scanning across billions of leaked records. Free tier available.