CVE-2026-43632: CWE-416 Use After Free in ggml-org llama.cpp
A use-after-free vulnerability exists in llama.cpp builds starting from b7492 in the llama-server component affecting six tokenization endpoints. The flaw arises from a race condition where the main thread frees a vocabulary object after releasing a synchronization lock but before HTTP worker threads finish using it, potentially causing crashes or code execution when the --sleep-idle-seconds option is enabled.
AI Analysis
Technical Summary
CVE-2026-43632 is a critical use-after-free vulnerability (CWE-416) in the llama.cpp project by ggml-org, specifically in the llama-server's handling of six tokenization endpoints (/tokenize, /detokenize, /infill, /apply-template, /rerank, and /anthropic/count_tokens). The vulnerability stems from a time-of-check-time-of-use (TOCTOU) race condition where the main thread destroys and frees the ctx_server.vocab object after releasing a synchronization lock but before the HTTP worker thread completes its access. This unsafe concurrent access can lead to application crashes or potentially arbitrary code execution, particularly when the --sleep-idle-seconds configuration is used. The affected version explicitly identified is build b7492. No official patch or remediation guidance is currently provided by the vendor.
Potential Impact
Exploitation of this vulnerability can cause the llama-server to crash or potentially allow an attacker to execute arbitrary code remotely without authentication. The vulnerability affects multiple tokenization endpoints that bypass the task queue and directly access shared memory structures unsafely. The CVSS 4.0 score is 9.2 (critical), reflecting the network attack vector, high complexity, no privileges required, no user interaction, and high impact on confidentiality, integrity, and availability.
Mitigation Recommendations
Patch status is not yet confirmed — check the vendor advisory for current remediation guidance. Until an official fix is released, users should consider disabling or avoiding use of the --sleep-idle-seconds option and restrict access to the affected tokenization endpoints to trusted users only to reduce exposure. Monitoring for updates from the ggml-org project is recommended.
CVE-2026-43632: CWE-416 Use After Free in ggml-org llama.cpp
Description
A use-after-free vulnerability exists in llama.cpp builds starting from b7492 in the llama-server component affecting six tokenization endpoints. The flaw arises from a race condition where the main thread frees a vocabulary object after releasing a synchronization lock but before HTTP worker threads finish using it, potentially causing crashes or code execution when the --sleep-idle-seconds option is enabled.
CVSS v4.0
Score 9.2critical
Weaknesses
AI-Powered Analysis
Machine-generated threat intelligence
Technical Analysis
CVE-2026-43632 is a critical use-after-free vulnerability (CWE-416) in the llama.cpp project by ggml-org, specifically in the llama-server's handling of six tokenization endpoints (/tokenize, /detokenize, /infill, /apply-template, /rerank, and /anthropic/count_tokens). The vulnerability stems from a time-of-check-time-of-use (TOCTOU) race condition where the main thread destroys and frees the ctx_server.vocab object after releasing a synchronization lock but before the HTTP worker thread completes its access. This unsafe concurrent access can lead to application crashes or potentially arbitrary code execution, particularly when the --sleep-idle-seconds configuration is used. The affected version explicitly identified is build b7492. No official patch or remediation guidance is currently provided by the vendor.
Potential Impact
Exploitation of this vulnerability can cause the llama-server to crash or potentially allow an attacker to execute arbitrary code remotely without authentication. The vulnerability affects multiple tokenization endpoints that bypass the task queue and directly access shared memory structures unsafely. The CVSS 4.0 score is 9.2 (critical), reflecting the network attack vector, high complexity, no privileges required, no user interaction, and high impact on confidentiality, integrity, and availability.
Mitigation Recommendations
Patch status is not yet confirmed — check the vendor advisory for current remediation guidance. Until an official fix is released, users should consider disabling or avoiding use of the --sleep-idle-seconds option and restrict access to the affected tokenization endpoints to trusted users only to reduce exposure. Monitoring for updates from the ggml-org project is recommended.
Technical Details
- Data Version
- 5.2
- Assigner Short Name
- VulnCheck
- Date Reserved
- 2026-05-01T18:22:45.641Z
- Cvss Version
- 4.0
- State
- PUBLISHED
- Remediation Level
- null
Threat ID: 6a750702bf8831d5395f59cd
Added to database: 08/06/2026, 22:13:22 UTC
Last enriched: 08/06/2026, 23:11:24 UTC
Last updated: 08/07/2026, 02:40:17 UTC
Views: 5
Community Reviews
0 reviewsCrowdsource mitigation strategies, share intel context, and vote on the most helpful responses. Sign in to add your voice and help keep defenders ahead.
Want to contribute mitigation steps or threat intel context? Sign in or create an account to join the community discussion.
Actions
Updates to AI analysis require Pro Console access. Upgrade inside Console → Billing.
External Links
Need more coverage?
Upgrade to Pro Console for AI refresh and higher limits.
For incident response and remediation, OffSeq services can help resolve threats faster.
Latest Threats
Check if your credentials are on the dark web
Instant breach scanning across billions of leaked records. Free tier available.