CVE-2026-69147: CWE-400: Uncontrolled Resource Consumption in vllm-project vllm
vLLM is an inference and serving engine for large language models. Prior to 0.28.0, request bodies for Chat Completions and Responses can set media_io_kwargs.video.video_backend to pynvvideocodec, and MediaConnector.fetch_video forwards that choice to VideoMediaIO even when startup configuration selected a software decoder. The engine's _reserve_mm_ipc_gpu_memory logic budgets decoder memory only from static configuration, so the request-selected VIDEO_LOADER_REGISTRY backend can create a CUDA context, decoder surfaces, and decoded-frame allocations that were not removed from the engine's KV-cache budget. An attacker able to submit video requests to a video-capable GPU deployment with PyNvVideoCodec installed can exhaust shared GPU memory, causing request failures, worker crashes, or denial of service. The first release containing the fix is version 0.28.0.
AI Analysis
Technical Summary
The vLLM inference engine for large language models has a resource consumption vulnerability (CWE-400) in versions before 0.28.0. Specifically, request bodies for Chat Completions and Responses can specify media_io_kwargs.video.video_backend as pynvvideocodec, which causes MediaConnector.fetch_video to use the VideoMediaIO backend regardless of the startup configuration. The engine's memory budgeting logic does not account for decoder memory allocated dynamically by this backend, allowing creation of CUDA contexts and decoder surfaces that exhaust shared GPU memory. This leads to request failures, worker crashes, or denial of service. The vulnerability is addressed in vLLM version 0.28.0.
Potential Impact
An attacker with the ability to submit video requests to a GPU deployment running vLLM with PyNvVideoCodec installed can exhaust shared GPU memory. This results in denial of service conditions such as request failures and worker crashes. There is no impact on confidentiality or integrity reported, only availability is affected.
Mitigation Recommendations
Upgrade vLLM to version 0.28.0 or later, which contains the fix for this uncontrolled resource consumption vulnerability. No other mitigation is indicated or required.
CVE-2026-69147: CWE-400: Uncontrolled Resource Consumption in vllm-project vllm
Description
vLLM is an inference and serving engine for large language models. Prior to 0.28.0, request bodies for Chat Completions and Responses can set media_io_kwargs.video.video_backend to pynvvideocodec, and MediaConnector.fetch_video forwards that choice to VideoMediaIO even when startup configuration selected a software decoder. The engine's _reserve_mm_ipc_gpu_memory logic budgets decoder memory only from static configuration, so the request-selected VIDEO_LOADER_REGISTRY backend can create a CUDA context, decoder surfaces, and decoded-frame allocations that were not removed from the engine's KV-cache budget. An attacker able to submit video requests to a video-capable GPU deployment with PyNvVideoCodec installed can exhaust shared GPU memory, causing request failures, worker crashes, or denial of service. The first release containing the fix is version 0.28.0.
CVSS v3.1
Score 6.5medium
Affected software
vllm-project
vllm
Run on your own infrastructure? Check whether these packages are installed with threat-finder — our free open-source scanner.
AI-Powered Analysis
Machine-generated threat intelligence
Technical Analysis
The vLLM inference engine for large language models has a resource consumption vulnerability (CWE-400) in versions before 0.28.0. Specifically, request bodies for Chat Completions and Responses can specify media_io_kwargs.video.video_backend as pynvvideocodec, which causes MediaConnector.fetch_video to use the VideoMediaIO backend regardless of the startup configuration. The engine's memory budgeting logic does not account for decoder memory allocated dynamically by this backend, allowing creation of CUDA contexts and decoder surfaces that exhaust shared GPU memory. This leads to request failures, worker crashes, or denial of service. The vulnerability is addressed in vLLM version 0.28.0.
Potential Impact
An attacker with the ability to submit video requests to a GPU deployment running vLLM with PyNvVideoCodec installed can exhaust shared GPU memory. This results in denial of service conditions such as request failures and worker crashes. There is no impact on confidentiality or integrity reported, only availability is affected.
Mitigation Recommendations
Upgrade vLLM to version 0.28.0 or later, which contains the fix for this uncontrolled resource consumption vulnerability. No other mitigation is indicated or required.
Technical Details
- Data Version
- 5.2
- Assigner Short Name
- GitHub_M
- Date Reserved
- 2026-08-03T15:20:30.218Z
- Cvss Version
- 3.1
- State
- PUBLISHED
Threat ID: 6aaad9a055bf5e2cf5f96e2e
Added to database: 09/16/2026, 18:02:08 UTC
Last enriched: 09/16/2026, 18:16:38 UTC
Last updated: 09/16/2026, 19:02:13 UTC
Views: 4
Community Reviews
0 reviewsCrowdsource mitigation strategies, share intel context, and vote on the most helpful responses. Sign in to add your voice and help keep defenders ahead.
Want to contribute mitigation steps or threat intel context? Sign in or create an account to join the community discussion.
Actions
Updates to AI analysis require Pro Console access. Upgrade inside Console → Billing.
Need more coverage?
Upgrade to Pro Console for AI refresh and higher limits.
For incident response and remediation, OffSeq services can help resolve threats faster.
Latest Threats
Check if your credentials are on the dark web
Instant breach scanning across billions of leaked records. Free tier available.