CVE-2026-88049: CWE-787: Out-of-bounds Write in tesseract-ocr tesseract
Tesseract is an open source OCR engine. In version 5.5.3 and earlier, prior .traineddata hardening added bounds checks to NetworkIO::CopyTimeStepGeneral and NetworkIO::Randomize in src/lstm/networkio.cpp but left NetworkIO::WriteTimeStepPart and NetworkIO::AddTimeStepPart unchecked. In LSTM::Forward in src/lstm/lstm.cpp, source_ is sized from the independently deserialized na_ field while the WriteTimeStepPart count is ns_, which comes from the CI gate WeightMatrix dim1() value. A crafted NT_LSTM layer can make ns_ much larger than na_, causing a heap out-of-bounds write during the first recognition step on the default LSTM engine and resulting in heap corruption, a crash, or potentially controlled corruption. No fixed release is available as of this review.
AI Analysis
Technical Summary
Tesseract OCR versions 5.5.3 and earlier contain a heap out-of-bounds write vulnerability due to incomplete bounds checking in the LSTM network input/output functions. Specifically, NetworkIO::WriteTimeStepPart and NetworkIO::AddTimeStepPart lack proper bounds validation, while other related functions have such checks. The vulnerability occurs because the count ns_, derived from the CI gate WeightMatrix dimension, can be larger than the independently deserialized na_ field size, leading to a heap write beyond allocated memory during the first recognition step. This can cause heap corruption, application crashes, or potentially controlled memory corruption. No fixed release is currently available.
Potential Impact
Exploitation of this vulnerability can cause heap corruption resulting in application crashes or potentially allow an attacker to control memory corruption. This could impact the stability and security of applications using the vulnerable Tesseract OCR engine. However, no known exploits are reported in the wild as of the publication date.
Mitigation Recommendations
No official fix or patch is currently available for this vulnerability. Users should monitor the vendor's advisories for updates. Until a patch is released, avoid processing untrusted or specially crafted NT_LSTM layers that could trigger the vulnerability.
CVE-2026-88049: CWE-787: Out-of-bounds Write in tesseract-ocr tesseract
Description
Tesseract is an open source OCR engine. In version 5.5.3 and earlier, prior .traineddata hardening added bounds checks to NetworkIO::CopyTimeStepGeneral and NetworkIO::Randomize in src/lstm/networkio.cpp but left NetworkIO::WriteTimeStepPart and NetworkIO::AddTimeStepPart unchecked. In LSTM::Forward in src/lstm/lstm.cpp, source_ is sized from the independently deserialized na_ field while the WriteTimeStepPart count is ns_, which comes from the CI gate WeightMatrix dim1() value. A crafted NT_LSTM layer can make ns_ much larger than na_, causing a heap out-of-bounds write during the first recognition step on the default LSTM engine and resulting in heap corruption, a crash, or potentially controlled corruption. No fixed release is available as of this review.
CVSS v4.0
Score 8.6high
Affected software
tesseract-ocr
tesseract
pkg:github/tesseract-ocr/tesseractRun on your own infrastructure? Check whether these packages are installed with threat-finder — our free open-source scanner.
Weaknesses
AI-Powered Analysis
Machine-generated threat intelligence
Technical Analysis
Tesseract OCR versions 5.5.3 and earlier contain a heap out-of-bounds write vulnerability due to incomplete bounds checking in the LSTM network input/output functions. Specifically, NetworkIO::WriteTimeStepPart and NetworkIO::AddTimeStepPart lack proper bounds validation, while other related functions have such checks. The vulnerability occurs because the count ns_, derived from the CI gate WeightMatrix dimension, can be larger than the independently deserialized na_ field size, leading to a heap write beyond allocated memory during the first recognition step. This can cause heap corruption, application crashes, or potentially controlled memory corruption. No fixed release is currently available.
Potential Impact
Exploitation of this vulnerability can cause heap corruption resulting in application crashes or potentially allow an attacker to control memory corruption. This could impact the stability and security of applications using the vulnerable Tesseract OCR engine. However, no known exploits are reported in the wild as of the publication date.
Mitigation Recommendations
No official fix or patch is currently available for this vulnerability. Users should monitor the vendor's advisories for updates. Until a patch is released, avoid processing untrusted or specially crafted NT_LSTM layers that could trigger the vulnerability.
Technical Details
- Data Version
- 5.2
- Assigner Short Name
- GitHub_M
- Date Reserved
- 2026-09-09T21:22:45.433Z
- Cvss Version
- 4.0
- State
- PUBLISHED
Threat ID: 6aa2dccdf76d26f025573a87
Added to database: 09/10/2026, 16:37:33 UTC
Last enriched: 09/10/2026, 16:51:19 UTC
Last updated: 09/10/2026, 17:37:58 UTC
Views: 5
Community Reviews
0 reviewsCrowdsource mitigation strategies, share intel context, and vote on the most helpful responses. Sign in to add your voice and help keep defenders ahead.
Want to contribute mitigation steps or threat intel context? Sign in or create an account to join the community discussion.
Actions
Updates to AI analysis require Pro Console access. Upgrade inside Console → Billing.
Need more coverage?
Upgrade to Pro Console for AI refresh and higher limits.
For incident response and remediation, OffSeq services can help resolve threats faster.
Latest Threats
Check if your credentials are on the dark web
Instant breach scanning across billions of leaked records. Free tier available.