Gptline: NLTK: Pl196xCorpusReader has quadratic ReDoS on malformed TEI blocks (CVE-2026-81725)
The Pl196xCorpusReader component in the gptline product has a Regular Expression Denial of Service (ReDoS) vulnerability (CVE-2026-81725) caused by quadratic CPU growth when parsing malformed TEI blocks with many unmatched opening tags. This affects versions from 1.0.8 up to but not including 1.0.8_23. The vulnerability arises from lazy regex patterns that repeatedly rescan attacker-controlled text, leading to excessive CPU consumption and potential parser thread stalling. No official patch is currently available, but a fix involving replacing the regex parser with a linear parser or bounded tokenizer is recommended.
AI Analysis
Technical Summary
CVE-2026-81725 is a ReDoS vulnerability in the nltk.corpus.reader.pl196x.Pl196xCorpusReader, specifically in the TEICorpusView.read_block method. The issue is due to lazy '.*?' regexes that rescan malformed TEI blocks containing many unmatched opening tags, causing quadratic CPU time growth during parsing operations such as words() and tagged_words(). Both published version 3.9.4 and source version 3.10.0-rc2 demonstrate the issue. The vulnerability allows an attacker supplying malformed corpus files to cause heavy CPU usage and stall parser threads. A patch is not yet available, but remediation involves replacing the regex-based parser with a linear or bounded tokenizer approach.
Potential Impact
Attackers can supply malformed TEI corpus files with many unmatched opening tags to cause the Pl196xCorpusReader to consume excessive CPU resources, leading to denial of service by stalling parser threads. This impacts applications that parse attacker-influenced corpus files using public reader APIs such as words() and tagged_words(). There is no indication of data breach or code execution, only resource exhaustion.
Mitigation Recommendations
A patch is available but not yet released. The vendor recommends replacing the lazy regex parser with a linear parser or bounded tokenizer and adding regression tests to ensure near-linear parsing time on malformed inputs. Until a patch is released, users should avoid processing untrusted or attacker-controlled corpus files with Pl196xCorpusReader to mitigate risk.
Gptline: NLTK: Pl196xCorpusReader has quadratic ReDoS on malformed TEI blocks (CVE-2026-81725)
Description
The Pl196xCorpusReader component in the gptline product has a Regular Expression Denial of Service (ReDoS) vulnerability (CVE-2026-81725) caused by quadratic CPU growth when parsing malformed TEI blocks with many unmatched opening tags. This affects versions from 1.0.8 up to but not including 1.0.8_23. The vulnerability arises from lazy regex patterns that repeatedly rescan attacker-controlled text, leading to excessive CPU consumption and potential parser thread stalling. No official patch is currently available, but a fix involving replacing the regex parser with a linear parser or bounded tokenizer is recommended.
CVSS v4.0
Affected software
Run on your own infrastructure? Check whether these packages are installed with threat-finder — our free open-source scanner.
AI-Powered Analysis
Machine-generated threat intelligence
Technical Analysis
CVE-2026-81725 is a ReDoS vulnerability in the nltk.corpus.reader.pl196x.Pl196xCorpusReader, specifically in the TEICorpusView.read_block method. The issue is due to lazy '.*?' regexes that rescan malformed TEI blocks containing many unmatched opening tags, causing quadratic CPU time growth during parsing operations such as words() and tagged_words(). Both published version 3.9.4 and source version 3.10.0-rc2 demonstrate the issue. The vulnerability allows an attacker supplying malformed corpus files to cause heavy CPU usage and stall parser threads. A patch is not yet available, but remediation involves replacing the regex-based parser with a linear or bounded tokenizer approach.
Potential Impact
Attackers can supply malformed TEI corpus files with many unmatched opening tags to cause the Pl196xCorpusReader to consume excessive CPU resources, leading to denial of service by stalling parser threads. This impacts applications that parse attacker-influenced corpus files using public reader APIs such as words() and tagged_words(). There is no indication of data breach or code execution, only resource exhaustion.
Mitigation Recommendations
A patch is available but not yet released. The vendor recommends replacing the lazy regex parser with a linear parser or bounded tokenizer and adding regression tests to ensure near-linear parsing time on malformed inputs. Until a patch is released, users should avoid processing untrusted or attacker-controlled corpus files with Pl196xCorpusReader to mitigate risk.
Technical Details
- Gcve Source
- db.gcve.eu
- Osv Id
- BREW-gptline-CVE-2026-81725
- Osv Schema Version
- 1.7.3
- Ecosystems
- ["Homebrew"]
- Cvss Version
- 4.0
Threat ID: 6aac8e5955bf5e2cf5491aeb
Added to database: 09/18/2026, 01:05:29 UTC
Last enriched: 09/18/2026, 01:50:51 UTC
Last updated: 09/18/2026, 02:07:21 UTC
Views: 4
Community Reviews
0 reviewsCrowdsource mitigation strategies, share intel context, and vote on the most helpful responses. Sign in to add your voice and help keep defenders ahead.
Want to contribute mitigation steps or threat intel context? Sign in or create an account to join the community discussion.
Actions
Updates to AI analysis require Pro Console access. Upgrade inside Console → Billing.
Need more coverage?
Upgrade to Pro Console for AI refresh and higher limits.
For incident response and remediation, OffSeq services can help resolve threats faster.
Latest Threats
Check if your credentials are on the dark web
Instant breach scanning across billions of leaked records. Free tier available.