CVE-2026-88052: CWE-129: Improper Validation of Array Index in tesseract-ocr tesseract
Tesseract is an open source OCR engine. In version 5.5.3 and earlier, UNICHARSET::load_via_fgets in src/ccutil/unicharset.cpp trusts the declared unichar count as a loop bound and uses id as an unchecked index into the unichars vector. unichar_insert_backwards_compatible can leave the vector unchanged for an empty, duplicate, or already-encodable representation, causing id to become larger than unichars.size(). Subsequent set_* calls and the write to unichars[id].properties.enabled then write UNICHAR_PROPERTIES beyond the vector during initialization in both the default LSTM and legacy engines, causing heap corruption, a crash, or potentially controlled corruption. No fixed release is available as of this review.
AI Analysis
Technical Summary
The vulnerability occurs in tesseract-ocr's UNICHARSET::load_via_fgets function in src/ccutil/unicharset.cpp, where the declared unichar count is trusted as a loop bound and the 'id' variable is used as an unchecked index into the unichars vector. The function unichar_insert_backwards_compatible may leave the vector unchanged for empty, duplicate, or already-encodable representations, causing 'id' to exceed the vector size. Subsequent writes to unichars[id].properties.enabled result in heap corruption, crashes, or potentially controlled memory corruption during initialization in both the default LSTM and legacy OCR engines. This issue affects versions 5.5.3 and earlier. As of the latest review, no patch or fixed release is available.
Potential Impact
Exploitation of this vulnerability can lead to heap corruption, causing application crashes or potentially controlled memory corruption. This can result in denial of service or arbitrary code execution within the context of the OCR engine. The CVSS 3.1 score is 7.8 (high), reflecting the potential for high confidentiality, integrity, and availability impact. No known exploits are reported in the wild at this time.
Mitigation Recommendations
No official fix or patch is currently available for this vulnerability. Users should monitor the vendor's advisories for updates and consider applying any future patches promptly. Until a fix is released, avoid processing untrusted input with affected versions of tesseract-ocr to reduce risk.
CVE-2026-88052: CWE-129: Improper Validation of Array Index in tesseract-ocr tesseract
Description
Tesseract is an open source OCR engine. In version 5.5.3 and earlier, UNICHARSET::load_via_fgets in src/ccutil/unicharset.cpp trusts the declared unichar count as a loop bound and uses id as an unchecked index into the unichars vector. unichar_insert_backwards_compatible can leave the vector unchanged for an empty, duplicate, or already-encodable representation, causing id to become larger than unichars.size(). Subsequent set_* calls and the write to unichars[id].properties.enabled then write UNICHAR_PROPERTIES beyond the vector during initialization in both the default LSTM and legacy engines, causing heap corruption, a crash, or potentially controlled corruption. No fixed release is available as of this review.
CVSS v3.1
Score 7.8high
Affected software
tesseract-ocr
tesseract
pkg:github/tesseract-ocr/tesseractRun on your own infrastructure? Check whether these packages are installed with threat-finder — our free open-source scanner.
AI-Powered Analysis
Machine-generated threat intelligence
Technical Analysis
The vulnerability occurs in tesseract-ocr's UNICHARSET::load_via_fgets function in src/ccutil/unicharset.cpp, where the declared unichar count is trusted as a loop bound and the 'id' variable is used as an unchecked index into the unichars vector. The function unichar_insert_backwards_compatible may leave the vector unchanged for empty, duplicate, or already-encodable representations, causing 'id' to exceed the vector size. Subsequent writes to unichars[id].properties.enabled result in heap corruption, crashes, or potentially controlled memory corruption during initialization in both the default LSTM and legacy OCR engines. This issue affects versions 5.5.3 and earlier. As of the latest review, no patch or fixed release is available.
Potential Impact
Exploitation of this vulnerability can lead to heap corruption, causing application crashes or potentially controlled memory corruption. This can result in denial of service or arbitrary code execution within the context of the OCR engine. The CVSS 3.1 score is 7.8 (high), reflecting the potential for high confidentiality, integrity, and availability impact. No known exploits are reported in the wild at this time.
Mitigation Recommendations
No official fix or patch is currently available for this vulnerability. Users should monitor the vendor's advisories for updates and consider applying any future patches promptly. Until a fix is released, avoid processing untrusted input with affected versions of tesseract-ocr to reduce risk.
Technical Details
- Data Version
- 5.2
- Assigner Short Name
- GitHub_M
- Date Reserved
- 2026-09-09T21:22:45.433Z
- Cvss Version
- 3.1
- State
- PUBLISHED
Threat ID: 6aa2ee7b555a9c516207cb20
Added to database: 09/10/2026, 17:52:59 UTC
Last enriched: 09/10/2026, 18:06:24 UTC
Last updated: 09/10/2026, 18:44:06 UTC
Views: 7
Community Reviews
0 reviewsCrowdsource mitigation strategies, share intel context, and vote on the most helpful responses. Sign in to add your voice and help keep defenders ahead.
Want to contribute mitigation steps or threat intel context? Sign in or create an account to join the community discussion.
Actions
Updates to AI analysis require Pro Console access. Upgrade inside Console → Billing.
Need more coverage?
Upgrade to Pro Console for AI refresh and higher limits.
For incident response and remediation, OffSeq services can help resolve threats faster.
Latest Threats
Check if your credentials are on the dark web
Instant breach scanning across billions of leaked records. Free tier available.