CVE-2026-73557: CWE-362: Concurrent Execution using Shared Resource with Improper Synchronization ('Race Condition') in vllm-project vllm
CVE-2026-73557 is a race condition vulnerability in the vllm-project's vllm software affecting versions from 0.20.2rc0 up to but not including 0.26.0. It arises from improper synchronization of a process-global sparse tensor invariant guard in PyTorch 2.11.0 during concurrent prompt-embedding reconstruction within a single chat completion request. This can lead to invalid sparse tensor deserialization and subsequent unsafe dense tensor conversion. The vulnerability requires the --enable-prompt-embeds feature (default off) but does not require multiple workers or multimodal embeddings. The flaw allows a malicious payload to bypass the follow-up guard, potentially causing memory corruption or denial of service as documented in the related CVE-2025-62164. The CVSS 4.0 base score is 6.3 (medium severity).
AI Analysis
Technical Summary
This vulnerability is a concurrency issue (CWE-362) in vllm versions >=0.20.2rc0 <0.26.0 where the process-global sparse tensor invariant guard used in torch.sparse.check_sparse_tensor_invariants() is improperly synchronized. The guard uses save/enable/restore operations on a global flag in PyTorch 2.11.0. When two prompt-embedding parts are processed concurrently in one /v1/chat/completions request, one context may restore the global flag to false while the other is still inside its guarded section. This allows the second context to load an invalid sparse tensor without the invariant check, leading to unsafe tensor.to_dense() calls. The vulnerability builds on the previously reported CVE-2025-62164 but introduces a distinct concurrency root cause. Exploitation requires the --enable-prompt-embeds flag but not other optional features. The vulnerability was introduced after partial fixes for CVE-2025-62164 and is confirmed by deterministic testing with hash-verified source. The affected code is in vllm/renderers/embed_utils.py around the guarded loader. The vulnerability does not require API authentication and can be triggered in the default server configuration if --enable-prompt-embeds is enabled.
Potential Impact
An attacker can bypass the sparse tensor invariant guard due to a race condition, causing deserialization of malformed sparse tensors. This can lead to memory corruption, denial of service, or potentially code execution as previously documented for CVE-2025-62164. The concurrency flaw allows a malicious payload to be processed unsafely within a single chat completion request, increasing the risk of exploitation. However, the --enable-prompt-embeds feature is default off, reducing exposure unless explicitly enabled. API authentication is optional, so unauthenticated requests may trigger the vulnerability if the feature is enabled.
Mitigation Recommendations
No official patch or fix is explicitly stated in the provided data. Patch status is not yet confirmed — check the vendor advisory for current remediation guidance. Until a fix is available, avoid enabling the --enable-prompt-embeds feature in production environments. If this feature is required, consider applying additional synchronization controls or isolating usage to trusted inputs. Monitor vendor communications for updates and patches addressing this concurrency issue.
CVE-2026-73557: CWE-362: Concurrent Execution using Shared Resource with Improper Synchronization ('Race Condition') in vllm-project vllm
Description
CVE-2026-73557 is a race condition vulnerability in the vllm-project's vllm software affecting versions from 0.20.2rc0 up to but not including 0.26.0. It arises from improper synchronization of a process-global sparse tensor invariant guard in PyTorch 2.11.0 during concurrent prompt-embedding reconstruction within a single chat completion request. This can lead to invalid sparse tensor deserialization and subsequent unsafe dense tensor conversion. The vulnerability requires the --enable-prompt-embeds feature (default off) but does not require multiple workers or multimodal embeddings. The flaw allows a malicious payload to bypass the follow-up guard, potentially causing memory corruption or denial of service as documented in the related CVE-2025-62164. The CVSS 4.0 base score is 6.3 (medium severity).
CVSS v4.0
Score 6.3medium
Affected software
vllm-project
vllm
Run on your own infrastructure? Check whether these packages are installed with threat-finder — our free open-source scanner.
Weaknesses
AI-Powered Analysis
Machine-generated threat intelligence
Technical Analysis
This vulnerability is a concurrency issue (CWE-362) in vllm versions >=0.20.2rc0 <0.26.0 where the process-global sparse tensor invariant guard used in torch.sparse.check_sparse_tensor_invariants() is improperly synchronized. The guard uses save/enable/restore operations on a global flag in PyTorch 2.11.0. When two prompt-embedding parts are processed concurrently in one /v1/chat/completions request, one context may restore the global flag to false while the other is still inside its guarded section. This allows the second context to load an invalid sparse tensor without the invariant check, leading to unsafe tensor.to_dense() calls. The vulnerability builds on the previously reported CVE-2025-62164 but introduces a distinct concurrency root cause. Exploitation requires the --enable-prompt-embeds flag but not other optional features. The vulnerability was introduced after partial fixes for CVE-2025-62164 and is confirmed by deterministic testing with hash-verified source. The affected code is in vllm/renderers/embed_utils.py around the guarded loader. The vulnerability does not require API authentication and can be triggered in the default server configuration if --enable-prompt-embeds is enabled.
Potential Impact
An attacker can bypass the sparse tensor invariant guard due to a race condition, causing deserialization of malformed sparse tensors. This can lead to memory corruption, denial of service, or potentially code execution as previously documented for CVE-2025-62164. The concurrency flaw allows a malicious payload to be processed unsafely within a single chat completion request, increasing the risk of exploitation. However, the --enable-prompt-embeds feature is default off, reducing exposure unless explicitly enabled. API authentication is optional, so unauthenticated requests may trigger the vulnerability if the feature is enabled.
Mitigation Recommendations
No official patch or fix is explicitly stated in the provided data. Patch status is not yet confirmed — check the vendor advisory for current remediation guidance. Until a fix is available, avoid enabling the --enable-prompt-embeds feature in production environments. If this feature is required, consider applying additional synchronization controls or isolating usage to trusted inputs. Monitor vendor communications for updates and patches addressing this concurrency issue.
Technical Details
- Data Version
- 5.2
- Assigner Short Name
- GitHub_M
- Date Reserved
- 2026-08-12T20:53:46.380Z
- Cvss Version
- 4.0
- State
- PUBLISHED
Threat ID: 6a7ddecdbf8831d5395d0186
Added to database: 08/13/2026, 15:12:13 UTC
Last enriched: 09/13/2026, 13:03:56 UTC
Last updated: 09/26/2026, 01:47:44 UTC
Views: 60
Community Reviews
0 reviewsCrowdsource mitigation strategies, share intel context, and vote on the most helpful responses. Sign in to add your voice and help keep defenders ahead.
Want to contribute mitigation steps or threat intel context? Sign in or create an account to join the community discussion.
Actions
Updates to AI analysis require Pro Console access. Upgrade inside Console → Billing.
Need more coverage?
Upgrade to Pro Console for AI refresh and higher limits.
For incident response and remediation, OffSeq services can help resolve threats faster.
Latest Threats
Check if your credentials are on the dark web
Instant breach scanning across billions of leaked records. Free tier available.