CVE-2026-43632: Use After Free in ggml-org llama.cpp
llama.cpp builds b7492 through the latest b9060 contains a use-after-free vulnerability in llama-server affecting six tokenization endpoints (/tokenize, /detokenize, /infill, /apply-template, /rerank, and /anthropic/count_tokens) that bypass the task queue and access ctx_server.vocab directly on HTTP worker threads. Attackers can exploit a time-of-check-time-of-use race condition where the main thread destroys and frees vocab after the synchronization lock is released but before the handler finishes using it, causing a crash or potential code execution when --sleep-idle-seconds is configured.
AI Analysis
Technical Summary
CVE-2026-43632 is a use-after-free vulnerability in the llama-server component of llama.cpp (ggml-org) affecting builds from b7492 through the latest b9060. The issue occurs in six tokenization endpoints (/tokenize, /detokenize, /infill, /apply-template, /rerank, and /anthropic/count_tokens) that bypass the task queue and directly access ctx_server.vocab on HTTP worker threads. A time-of-check-time-of-use race condition allows the main thread to destroy and free vocab after releasing a synchronization lock but before the handler completes its use, leading to potential crashes or code execution when the --sleep-idle-seconds option is enabled.
Potential Impact
Exploitation of this vulnerability can cause application crashes or potentially allow remote code execution due to the use-after-free condition in the tokenization endpoints of llama-server. The vulnerability requires network access and has a high complexity due to the race condition. No known exploits are reported in the wild at this time.
Mitigation Recommendations
Patch status is not yet confirmed — check the vendor advisory for current remediation guidance. No official fix or patch links are currently available. Users should monitor vendor communications for updates. Until a fix is released, consider disabling or restricting access to the affected tokenization endpoints or avoid using the --sleep-idle-seconds configuration option if feasible.
CVE-2026-43632: Use After Free in ggml-org llama.cpp
Description
llama.cpp builds b7492 through the latest b9060 contains a use-after-free vulnerability in llama-server affecting six tokenization endpoints (/tokenize, /detokenize, /infill, /apply-template, /rerank, and /anthropic/count_tokens) that bypass the task queue and access ctx_server.vocab directly on HTTP worker threads. Attackers can exploit a time-of-check-time-of-use race condition where the main thread destroys and frees vocab after the synchronization lock is released but before the handler finishes using it, causing a crash or potential code execution when --sleep-idle-seconds is configured.
CVSS v4.0
Score 9.2critical
Affected software
ggml-org
llama.cpp
Weaknesses
AI-Powered Analysis
Machine-generated threat intelligence
Technical Analysis
CVE-2026-43632 is a use-after-free vulnerability in the llama-server component of llama.cpp (ggml-org) affecting builds from b7492 through the latest b9060. The issue occurs in six tokenization endpoints (/tokenize, /detokenize, /infill, /apply-template, /rerank, and /anthropic/count_tokens) that bypass the task queue and directly access ctx_server.vocab on HTTP worker threads. A time-of-check-time-of-use race condition allows the main thread to destroy and free vocab after releasing a synchronization lock but before the handler completes its use, leading to potential crashes or code execution when the --sleep-idle-seconds option is enabled.
Potential Impact
Exploitation of this vulnerability can cause application crashes or potentially allow remote code execution due to the use-after-free condition in the tokenization endpoints of llama-server. The vulnerability requires network access and has a high complexity due to the race condition. No known exploits are reported in the wild at this time.
Mitigation Recommendations
Patch status is not yet confirmed — check the vendor advisory for current remediation guidance. No official fix or patch links are currently available. Users should monitor vendor communications for updates. Until a fix is released, consider disabling or restricting access to the affected tokenization endpoints or avoid using the --sleep-idle-seconds configuration option if feasible.
Technical Details
- Data Version
- 5.2
- Assigner Short Name
- VulnCheck
- Date Reserved
- 2026-05-01T18:22:45.641Z
- Cvss Version
- 4.0
- State
- PUBLISHED
Threat ID: 6a750702bf8831d5395f59cd
Added to database: 08/06/2026, 22:13:22 UTC
Last enriched: 08/14/2026, 15:52:48 UTC
Last updated: 09/21/2026, 22:01:34 UTC
Views: 28
Community Reviews
0 reviewsCrowdsource mitigation strategies, share intel context, and vote on the most helpful responses. Sign in to add your voice and help keep defenders ahead.
Want to contribute mitigation steps or threat intel context? Sign in or create an account to join the community discussion.
Actions
Updates to AI analysis require Pro Console access. Upgrade inside Console → Billing.
Need more coverage?
Upgrade to Pro Console for AI refresh and higher limits.
For incident response and remediation, OffSeq services can help resolve threats faster.
Latest Threats
Check if your credentials are on the dark web
Instant breach scanning across billions of leaked records. Free tier available.