CVE-2026-34756: CWE-770: Allocation of Resources Without Limits or Throttling in vllm-project vllm
vLLM is an inference and serving engine for large language models (LLMs). From 0.1.0 to before 0.19.0, a Denial of Service vulnerability exists in the vLLM OpenAI-compatible API server. Due to the lack of an upper bound validation on the n parameter in the ChatCompletionRequest and CompletionRequest Pydantic models, an unauthenticated attacker can send a single HTTP request with an astronomically large n value. This completely blocks the Python asyncio event loop and causes immediate Out-Of-Memory crashes by allocating millions of request object copies in the heap before the request even reaches the scheduling queue. This vulnerability is fixed in 0.19.0.
AI Analysis
Technical Summary
The vLLM inference and serving engine for large language models has a resource allocation vulnerability (CWE-770) in versions >=0.1.0 and <0.19.0. Specifically, the OpenAI-compatible API server does not enforce an upper limit on the 'n' parameter in ChatCompletionRequest and CompletionRequest Pydantic models. An unauthenticated attacker can exploit this by sending a single HTTP request with an astronomically large 'n' value, causing the Python asyncio event loop to block and triggering immediate out-of-memory crashes due to massive heap allocations before the request reaches the scheduler. This denial of service vulnerability is resolved in vLLM 0.19.0. The CVSS 3.1 base score is 6.5 (medium severity) with network attack vector, low attack complexity, no privileges required, no user interaction, unchanged scope, no confidentiality or integrity impact, and high availability impact. Red Hat has published an advisory referencing this CVE but does not explicitly state a patch or fix version for their AI Inference Server product. Patch status is not yet confirmed — check the vendor advisory for current remediation guidance.
Potential Impact
An unauthenticated attacker can cause a denial of service by sending a single specially crafted HTTP request with a very large 'n' parameter value. This leads to blocking of the Python asyncio event loop and immediate out-of-memory crashes due to excessive memory allocation. The impact is limited to availability disruption; confidentiality and integrity are not affected.
Mitigation Recommendations
The vulnerability is fixed in vLLM version 0.19.0. Users should upgrade to vLLM 0.19.0 or later to remediate this issue. The Red Hat advisory referencing this CVE does not explicitly confirm a patch or fix availability for their AI Inference Server product, so users should consult the vendor advisory for the latest remediation guidance. Until patched, consider restricting access to the vulnerable API endpoints or implementing request size limits at network or application layers to mitigate exploitation risk.
CVE-2026-34756: CWE-770: Allocation of Resources Without Limits or Throttling in vllm-project vllm
Description
vLLM is an inference and serving engine for large language models (LLMs). From 0.1.0 to before 0.19.0, a Denial of Service vulnerability exists in the vLLM OpenAI-compatible API server. Due to the lack of an upper bound validation on the n parameter in the ChatCompletionRequest and CompletionRequest Pydantic models, an unauthenticated attacker can send a single HTTP request with an astronomically large n value. This completely blocks the Python asyncio event loop and causes immediate Out-Of-Memory crashes by allocating millions of request object copies in the heap before the request even reaches the scheduling queue. This vulnerability is fixed in 0.19.0.
CVSS v3.1
Score 6.5medium
Affected software
Run on your own infrastructure? Check whether these packages are installed with threat-finder — our free open-source scanner.
AI-Powered Analysis
Machine-generated threat intelligence
Technical Analysis
The vLLM inference and serving engine for large language models has a resource allocation vulnerability (CWE-770) in versions >=0.1.0 and <0.19.0. Specifically, the OpenAI-compatible API server does not enforce an upper limit on the 'n' parameter in ChatCompletionRequest and CompletionRequest Pydantic models. An unauthenticated attacker can exploit this by sending a single HTTP request with an astronomically large 'n' value, causing the Python asyncio event loop to block and triggering immediate out-of-memory crashes due to massive heap allocations before the request reaches the scheduler. This denial of service vulnerability is resolved in vLLM 0.19.0. The CVSS 3.1 base score is 6.5 (medium severity) with network attack vector, low attack complexity, no privileges required, no user interaction, unchanged scope, no confidentiality or integrity impact, and high availability impact. Red Hat has published an advisory referencing this CVE but does not explicitly state a patch or fix version for their AI Inference Server product. Patch status is not yet confirmed — check the vendor advisory for current remediation guidance.
Potential Impact
An unauthenticated attacker can cause a denial of service by sending a single specially crafted HTTP request with a very large 'n' parameter value. This leads to blocking of the Python asyncio event loop and immediate out-of-memory crashes due to excessive memory allocation. The impact is limited to availability disruption; confidentiality and integrity are not affected.
Mitigation Recommendations
The vulnerability is fixed in vLLM version 0.19.0. Users should upgrade to vLLM 0.19.0 or later to remediate this issue. The Red Hat advisory referencing this CVE does not explicitly confirm a patch or fix availability for their AI Inference Server product, so users should consult the vendor advisory for the latest remediation guidance. Until patched, consider restricting access to the vulnerable API endpoints or implementing request size limits at network or application layers to mitigate exploitation risk.
Technical Details
- Data Version
- 5.2
- Assigner Short Name
- GitHub_M
- Date Reserved
- 2026-03-30T19:17:10.225Z
- Cvss Version
- 3.1
- State
- PUBLISHED
- Remediation Level
- null
- Vendor Advisory Urls
- [{"url":"https://access.redhat.com/security/cve/CVE-2026-34756","vendor":"Red Hat"}]
Threat ID: 69d49831aaed68159aca0f94
Added to database: 04/07/2026, 05:37:53 UTC
Last enriched: 07/15/2026, 13:19:48 UTC
Last updated: 07/31/2026, 19:22:58 UTC
Views: 182
Community Reviews
0 reviewsCrowdsource mitigation strategies, share intel context, and vote on the most helpful responses. Sign in to add your voice and help keep defenders ahead.
Want to contribute mitigation steps or threat intel context? Sign in or create an account to join the community discussion.
Actions
Updates to AI analysis require Pro Console access. Upgrade inside Console → Billing.
Need more coverage?
Upgrade to Pro Console for AI refresh and higher limits.
For incident response and remediation, OffSeq services can help resolve threats faster.
Latest Threats
Check if your credentials are on the dark web
Instant breach scanning across billions of leaked records. Free tier available.