Skip to main content

Threat Intelligence Database

Comprehensive database of the latest cyber threats affecting organizations worldwide. Filter and search to find specific threat intelligence relevant to your organization.

Pro Console Lifetime

Stop chasing alerts. Route them.

Start free, then upgrade once to turn Radar into an automated delivery engine for your security stack.

Custom feeds / Automations: email, Slack, webhooks, SIEM/MISP / API access (baseline limits)

View Plans & Pricing

API access activates after upgrading in Console -> Billing.

Breach by OffSeqOFFSEQFRIENDS — 25% OFF

Check if your credentials are on the dark web

Instant breach scanning across billions of leaked records. Free tier available.

Scan now

Filter Threats

Narrow down the results by type, severity, or affected countries

Search threats by title, CVE ID, or description. Maximum 100 characters.
Active filters (1):Package: pkg:github/vllm-project/vllm

Threat Intelligence

Click on any threat for detailed analysis and mitigation recommendations

CVE-2026-94627 is a high-severity vulnerability in vllm up to version 0.29.0 where the Mooncake connector improperly manages GPU KV cache block ownership during concurrent child requests sharing a transfer ID. This leads to GPU memory exhaustion as orphaned KV cache blocks accumulate until the process restarts, preventing legitimate requests from executing.

Join the discussion

vLLM versions up to 0.29.0 contain a vulnerability where the tp_size parameter in kv_transfer_params on OpenAI-compatible completion endpoints is not validated. This allows attackers to request memory allocations of arbitrary size, potentially exhausting system memory and causing the decode worker process to be terminated by the kernel's out-of-memory (OOM) killer.

Join the discussion

vLLM versions up to 0.29.0 have a resource exhaustion vulnerability in the MooncakeConnector component. Rejected prefill requests create ownerless transfer placeholders that are never reclaimed, allowing attackers to exhaust sender task pools. This causes delays of up to 480 seconds for valid requests, although health checks continue to report success.

Join the discussion

vLLM versions up to 0.29.0 have a denial of service vulnerability in the P2P KV offloading feature when configured with TieringOffloadingSpec and a peer-to-peer secondary tier. Attackers can cause the system to crash by supplying arbitrary remote host and port values that create unreachable peer sessions, exhausting ZeroMQ socket resources and triggering an uncaught error that stops all inference.

Join the discussion

vLLM versions up to 0.29.0 contain a denial of service vulnerability in the NIXL connector's prefix caching implementation. This flaw allows attackers to cause an assertion failure by submitting multi-prompt completion requests with varying prompt lengths, leading to termination of the decode worker. The affected worker remains unavailable until manually restarted, disrupting service availability.

Join the discussion
0

vLLM versions up to 0.29.0 have a denial of service vulnerability caused by an uncaught KeyError in the NIXL connector's metadata handling. This occurs when requests contain incomplete kv_transfer_params dictionary entries, leading the decode engine to terminate and causing all routed requests to fail until manually restarted.

Join the discussion

vLLM through 0.29.0 fails to properly validate bad_words token indices against the model's generation output width in SamplingParams.update_from_tokenizer(). Attackers can supply out-of-bounds token indices that corrupt logits memory of concurrent requests, causing different in-flight HTTP requests to return incorrect tokens.

Join the discussion

vLLM through 0.29.0 contains a memory corruption vulnerability in the Triton _bincount_kernel where prompt token IDs index the penalty prompt-presence bitset without bounds checking against vocabulary size. Attackers can submit multimodal audio requests with tokens equal to vocabulary size, causing out-of-bounds writes that corrupt concurrent requests' sampler state and alter repetition penalty behavior.

Join the discussion

vLLM before 0.29.0 validates allowed_token_ids against tokenizer length instead of model output logits width in SamplingParams._validate_allowed_token_ids(). Attackers can supply token IDs above the output vocabulary that pass validation, causing LogitBiasState to corrupt GPU logits state and allow concurrent requests to sample tokens outside their allowlists.

Join the discussion

vLLM versions before 0.28.0 fail to validate the lower bound of token IDs in the /v1/embeddings and /pooling endpoints, allowing unauthenticated attackers to crash the engine by submitting negative token IDs. A single request with a negative token ID triggers a CUDA device-side assertion that poisons the GPU context, causing all subsequent requests to fail until the process restarts.

Join the discussion

Showing 1 to 10 of 37 results

Filters:Package: pkg:github/vllm-project/vllm
Page 1 of 4
OffSeq TrainingCredly Certified

Lead Pen Test Professional

Technical5-day eLearningPECB Accredited
View courses