CVE-2026-53923: CWE-681: Incorrect Conversion between Numeric Types in vllm-project vllm
vLLM is an inference and serving engine for large language models (LLMs). From 0.5.5 until 0.23.1rc0, integer truncation of tensor dimensions in vLLM's GGUF dequantize kernels (csrc/quantization/gguf/gguf_kernel.cu) causes partial tensor processing. The output tensor is allocated at full size via torch::empty (uninitialized memory), but the dequantize CUDA kernel processes only a truncated number of elements. The unfilled portion of the output tensor retains whatever was previously in GPU memory. In multi-tenant inference deployments, this residual GPU memory may contain tensor data from other users' inference requests, constituting information disclosure. This vulnerability is fixed in 0.23.1rc0.
AI Analysis
Technical Summary
The vulnerability in vLLM arises from incorrect conversion between numeric types (integer truncation) in the GGUF dequantize CUDA kernel, leading to partial processing of tensor elements. While the output tensor is allocated at full size using torch::empty (which leaves memory uninitialized), the kernel processes fewer elements than allocated, leaving the remainder of the tensor with residual GPU memory content. In multi-tenant inference deployments, this residual memory may contain sensitive data from other users' inference requests, causing information disclosure. This vulnerability is addressed and fixed in vLLM version 0.23.1rc0.
Potential Impact
This vulnerability can lead to information disclosure in multi-tenant inference deployments of vLLM, where residual GPU memory from other users' tensor data may be exposed due to partial tensor processing and uninitialized memory usage. The CVSS 4.0 score is 5.3 (medium severity), reflecting the network attack vector, low attack complexity, no privileges required, user interaction needed, and limited confidentiality and integrity impact.
Mitigation Recommendations
A fix is available in vLLM version 0.23.1rc0. Users should upgrade to this version or later to remediate the vulnerability. No vendor advisory was provided, so patch status is based on the description stating the issue is fixed in 0.23.1rc0. Until upgraded, users should consider the risk of information disclosure in multi-tenant environments.
CVE-2026-53923: CWE-681: Incorrect Conversion between Numeric Types in vllm-project vllm
Description
vLLM is an inference and serving engine for large language models (LLMs). From 0.5.5 until 0.23.1rc0, integer truncation of tensor dimensions in vLLM's GGUF dequantize kernels (csrc/quantization/gguf/gguf_kernel.cu) causes partial tensor processing. The output tensor is allocated at full size via torch::empty (uninitialized memory), but the dequantize CUDA kernel processes only a truncated number of elements. The unfilled portion of the output tensor retains whatever was previously in GPU memory. In multi-tenant inference deployments, this residual GPU memory may contain tensor data from other users' inference requests, constituting information disclosure. This vulnerability is fixed in 0.23.1rc0.
CVSS v4.0
Score 5.3medium
Affected software
Run on your own infrastructure? Check whether these packages are installed with threat-finder — our free open-source scanner.
AI-Powered Analysis
Machine-generated threat intelligence
Technical Analysis
The vulnerability in vLLM arises from incorrect conversion between numeric types (integer truncation) in the GGUF dequantize CUDA kernel, leading to partial processing of tensor elements. While the output tensor is allocated at full size using torch::empty (which leaves memory uninitialized), the kernel processes fewer elements than allocated, leaving the remainder of the tensor with residual GPU memory content. In multi-tenant inference deployments, this residual memory may contain sensitive data from other users' inference requests, causing information disclosure. This vulnerability is addressed and fixed in vLLM version 0.23.1rc0.
Potential Impact
This vulnerability can lead to information disclosure in multi-tenant inference deployments of vLLM, where residual GPU memory from other users' tensor data may be exposed due to partial tensor processing and uninitialized memory usage. The CVSS 4.0 score is 5.3 (medium severity), reflecting the network attack vector, low attack complexity, no privileges required, user interaction needed, and limited confidentiality and integrity impact.
Mitigation Recommendations
A fix is available in vLLM version 0.23.1rc0. Users should upgrade to this version or later to remediate the vulnerability. No vendor advisory was provided, so patch status is based on the description stating the issue is fixed in 0.23.1rc0. Until upgraded, users should consider the risk of information disclosure in multi-tenant environments.
Technical Details
- Data Version
- 5.2
- Assigner Short Name
- GitHub_M
- Date Reserved
- 2026-06-11T15:46:12.316Z
- Cvss Version
- 4.0
- State
- PUBLISHED
- Remediation Level
- null
Threat ID: 6a39b9b1eed863c81e85ff9f
Added to database: 06/22/2026, 22:39:45 UTC
Last enriched: 06/22/2026, 22:54:29 UTC
Last updated: 08/06/2026, 12:41:11 UTC
Views: 92
Community Reviews
0 reviewsCrowdsource mitigation strategies, share intel context, and vote on the most helpful responses. Sign in to add your voice and help keep defenders ahead.
Want to contribute mitigation steps or threat intel context? Sign in or create an account to join the community discussion.
Actions
Updates to AI analysis require Pro Console access. Upgrade inside Console → Billing.
Need more coverage?
Upgrade to Pro Console for AI refresh and higher limits.
For incident response and remediation, OffSeq services can help resolve threats faster.
Latest Threats
Check if your credentials are on the dark web
Instant breach scanning across billions of leaked records. Free tier available.