Script 'mail_helper' called by obssrc Hello community, here is the log from the commit of package tesseract-ocr for openSUSE:Factory checked in at 2026-09-24 22:55:32 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ Comparing /work/SRC/openSUSE:Factory/tesseract-ocr (Old) and /work/SRC/openSUSE:Factory/.tesseract-ocr.new.383539 (New) ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Package is "tesseract-ocr" Thu Sep 24 22:55:32 2026 rev:26 rq:1379978 version:5.5.3 Changes: -------- --- /work/SRC/openSUSE:Factory/tesseract-ocr/tesseract-ocr.changes 2026-07-28 17:49:24.766645934 +0200 +++ /work/SRC/openSUSE:Factory/.tesseract-ocr.new.383539/tesseract-ocr.changes 2026-09-24 22:55:49.589114671 +0200 @@ -1,0 +2,34 @@ +Wed Sep 23 12:35:46 UTC 2026 - Martin Pluskal <[email protected]> + +- CVE-2026-88047: stack buffer overflow in + Classify::ReadNormProtos on crafted traineddata (boo#1280925) + * tesseract-CVE-2026-88047.patch +- CVE-2026-88048: heap out-of-bounds write/read in + FullyConnected::Forward via dimension mismatch (boo#1280929) + * tesseract-CVE-2026-88048.patch +- CVE-2026-88049: heap out-of-bounds write in LSTM::Forward + via na_/gate-matrix dimension mismatch (boo#1280930) + * tesseract-CVE-2026-88049.patch +- CVE-2026-88050: out-of-bounds write in UnicharCompress + via unvalidated recoder code values (boo#1280931) + * tesseract-CVE-2026-88050.patch +- CVE-2026-88051: heap out-of-bounds write in + GenericVector<T>::read via reserved/size_used_ mismatch + (boo#1280932) + * tesseract-CVE-2026-88051.patch +- CVE-2026-88052: heap out-of-bounds write in + UNICHARSET::load_via_fgets via count/insert + desynchronization (boo#1280933) + * tesseract-CVE-2026-88052.patch +- CVE-2026-88053: heap out-of-bounds write in + Classify::ReadIntTemplates via unvalidated counts in + crafted traineddata (boo#1280934) + * tesseract-CVE-2026-88053.patch +- CVE-2026-88054: denial of service via empty-stack + dereference at model load (boo#1280935) + * tesseract-CVE-2026-88054.patch +- CVE-2026-73067 (boo#1275623): heap out-of-bounds read in + SquishedDawg on crafted model, already fixed in the + shipped 5.5.3 (DAWG edge-structure validation). + +------------------------------------------------------------------- New: ---- tesseract-CVE-2026-88047.patch tesseract-CVE-2026-88048.patch tesseract-CVE-2026-88049.patch tesseract-CVE-2026-88050.patch tesseract-CVE-2026-88051.patch tesseract-CVE-2026-88052.patch tesseract-CVE-2026-88053.patch tesseract-CVE-2026-88054.patch ----------(New B)---------- New: Classify::ReadNormProtos on crafted traineddata (boo#1280925) * tesseract-CVE-2026-88047.patch - CVE-2026-88048: heap out-of-bounds write/read in New: FullyConnected::Forward via dimension mismatch (boo#1280929) * tesseract-CVE-2026-88048.patch - CVE-2026-88049: heap out-of-bounds write in LSTM::Forward New: via na_/gate-matrix dimension mismatch (boo#1280930) * tesseract-CVE-2026-88049.patch - CVE-2026-88050: out-of-bounds write in UnicharCompress New: via unvalidated recoder code values (boo#1280931) * tesseract-CVE-2026-88050.patch - CVE-2026-88051: heap out-of-bounds write in New: (boo#1280932) * tesseract-CVE-2026-88051.patch - CVE-2026-88052: heap out-of-bounds write in New: desynchronization (boo#1280933) * tesseract-CVE-2026-88052.patch - CVE-2026-88053: heap out-of-bounds write in New: crafted traineddata (boo#1280934) * tesseract-CVE-2026-88053.patch - CVE-2026-88054: denial of service via empty-stack New: dereference at model load (boo#1280935) * tesseract-CVE-2026-88054.patch - CVE-2026-73067 (boo#1275623): heap out-of-bounds read in ----------(New E)---------- ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ Other differences: ------------------ ++++++ tesseract-ocr.spec ++++++ --- /var/tmp/diff_new_pack.YgKCla/_old 2026-09-24 22:55:50.297144260 +0200 +++ /var/tmp/diff_new_pack.YgKCla/_new 2026-09-24 22:55:50.299144343 +0200 @@ -25,6 +25,22 @@ URL: https://github.com/tesseract-ocr/tesseract Source0: https://github.com/tesseract-ocr/tesseract/archive/refs/tags/%{version}.tar.gz#/tesseract-%{version}.tar.gz Source99: baselibs.conf +# PATCH-FIX-UPSTREAM tesseract-CVE-2026-88047.patch GHSA-5j2p-r5vc-q7f3 (upstream commit 1bda5079) -- bound stack buffer in Classify::ReadNormProtos (CVE-2026-88047, boo#1280925) +Patch0: tesseract-CVE-2026-88047.patch +# PATCH-FIX-UPSTREAM tesseract-CVE-2026-88048.patch GHSA-q44c-23p6-5mw6 (upstream commit 103dc134) -- validate FullyConnected layer dims vs weight matrix (CVE-2026-88048, boo#1280929) +Patch1: tesseract-CVE-2026-88048.patch +# PATCH-FIX-UPSTREAM tesseract-CVE-2026-88049.patch GHSA-jgq8-pprg-vc68 (upstream commit b494ac18) -- bound LSTM::Forward WriteTimeStepPart count (CVE-2026-88049, boo#1280930) +Patch2: tesseract-CVE-2026-88049.patch +# PATCH-FIX-UPSTREAM tesseract-CVE-2026-88050.patch GHSA-7v9h-3q3m-w68g (upstream commit c94a5532) -- reject negative recoder code values (CVE-2026-88050, boo#1280931) +Patch3: tesseract-CVE-2026-88050.patch +# PATCH-FIX-UPSTREAM tesseract-CVE-2026-88051.patch GHSA-88qp-4g94-3rf3 (upstream commit 56e09ca1) -- cap GenericVector::read reserved/size_used_ (CVE-2026-88051, boo#1280932) +Patch4: tesseract-CVE-2026-88051.patch +# PATCH-FIX-UPSTREAM tesseract-CVE-2026-88052.patch GHSA-2hm8-q5c7-c373 (upstream commit 2d04d640) -- validate UNICHARSET load loop bound and index (CVE-2026-88052, boo#1280933) +Patch5: tesseract-CVE-2026-88052.patch +# PATCH-FIX-UPSTREAM tesseract-CVE-2026-88053.patch GHSA-rphx-x795-5qjv (upstream commit 8b057468) -- validate inttemp counts against MAX_* bounds (CVE-2026-88053, boo#1280934) +Patch6: tesseract-CVE-2026-88053.patch +# PATCH-FIX-UPSTREAM tesseract-CVE-2026-88054.patch GHSA-f6h7-cqr4-6fx4 (upstream commit 55277123) -- reject zero-length network stack in Plumbing::DeSerialize (CVE-2026-88054, boo#1280935) +Patch7: tesseract-CVE-2026-88054.patch BuildRequires: autoconf BuildRequires: automake BuildRequires: curl-devel ++++++ tesseract-CVE-2026-88047.patch ++++++ >From 1bda5079b1c8a7e25f523486837426903d29ce84 Mon Sep 17 00:00:00 2001 From: Stefan Weil <[email protected]> Date: Mon, 24 Aug 2026 13:39:22 +0200 Subject: [PATCH] Limit unichar extraction in ReadNormProtos to the buffer size Classify::ReadNormProtos parsed each normproto line with `stream >> unichar >> NumProtos` into a char unichar[2 * UNICHAR_LEN + 1] stack buffer, but char* extraction from an istream has no length limit (the stream width was never set). A crafted TESSDATA_NORMPROTO component in a .traineddata file whose first proto-line token exceeds 60 characters (the 100-byte line buffer allows up to 99) overflows the stack buffer during legacy engine initialization (CWE-121). Toolchain note: Apple's libc++ provides a C++20 array overload of operator>>(basic_istream&, char(&)[N]) that implicitly bounds the extraction to the array size, so builds against that standard library are incidentally protected. Standard libraries without that overload (e.g. libstdc++) still take the unbounded char* overload, so the explicit width limit below makes the behavior defined on all toolchains. Key changes: - normmatch.cpp: read the unichar token with std::setw(2 * UNICHAR_LEN + 1); char* extraction takes at most width - 1 characters, which fits the buffer exactly. Overlong tokens are truncated and the line is rejected like any other unparseable line. - unittest: add normproto_test covering a 99-character token (the maximum a 100-byte line can hold), a token of exactly 2 * UNICHAR_LEN characters, and a well-formed component. A minimal reproduction of the unbounded extraction crashes an ASan build with a stack-buffer-overflow. Reported-by: Tristan Madani <[email protected]> Assisted-by: OpenCode / qwen3.8-27b-thinking (Alibaba Cloud) Signed-off-by: Stefan Weil <[email protected]> --- # Backport note: test-only hunks (Makefile.am, unittest/) dropped - # the OBS build runs no %%check. Source fix hunks are verbatim upstream. # Makefile.am | 5 ++ src/classify/normmatch.cpp | 6 +- unittest/CMakeLists.txt | 1 + unittest/normproto_test.cc | 111 +++++++++++++++++++++++++++++++++++++ 4 files changed, 122 insertions(+), 1 deletion(-) create mode 100644 unittest/normproto_test.cc diff --git a/src/classify/normmatch.cpp b/src/classify/normmatch.cpp index d5bd7e6a28..1ce7c9283b 100644 --- a/src/classify/normmatch.cpp +++ b/src/classify/normmatch.cpp @@ -28,6 +28,7 @@ #include <cmath> #include <cstdio> +#include <iomanip> // for std::setw #include <sstream> // for std::istringstream namespace tesseract { @@ -190,7 +191,10 @@ NORM_PROTOS *Classify::ReadNormProtos(TFile *fp) { while (fp->FGets(line, kMaxLineSize) != nullptr) { std::istringstream stream(line); stream.imbue(std::locale::classic()); - stream >> unichar >> NumProtos; + // unichar holds at most 2 * UNICHAR_LEN characters; the width limit + // (width - 1 characters for char* extraction) keeps the extraction + // from overflowing the buffer on overlong lines. + stream >> std::setw(2 * UNICHAR_LEN + 1) >> unichar >> NumProtos; if (stream.fail()) { continue; } ++++++ tesseract-CVE-2026-88048.patch ++++++ >From 103dc134eb36411ddc6833ec20aa2c76795bd0ff Mon Sep 17 00:00:00 2001 From: Stefan Weil <[email protected]> Date: Fri, 21 Aug 2026 15:23:59 +0200 Subject: [PATCH] Validate FullyConnected weight matrix dimensions at load FullyConnected::DeSerialize read the WeightMatrix without checking that its dimensions match the layer's declared ni/no. MatrixDotVector drives the dot product from the matrix dimensions (writes w.dim1() results, reads w.dim2()-1 inputs) while the scratch buffers are sized from no_ and ni_, so a crafted .traineddata with a mismatched matrix performed a heap out-of-bounds write (up to 65535 rows) and out-of- bounds read on the first recognition step (crash, or heap corruption with attacker-influenced size and content). Key changes: - weightmatrix.h: add Dim1()/Dim2() accessors for the active weight matrix (int or float), alongside the existing NumOutputs(). - fullyconnected.cpp: reject the layer in DeSerialize unless Dim1() == no_ and Dim2() == ni_ + 1 (the second dimension includes the bias column), the layout that InitWeightsFloat always produces. - unittest: new fullyconnected_test that builds a Softmax network with mismatched matrix dimensions and expects CreateFromFile to return nullptr; on unpatched code the test reaches Forward and ASan catches the out-of-bounds write in MatrixDotVector. A second test verifies that matching dimensions are still accepted. Reported-by: Zhixi "Jace" Sun <[email protected]> Assisted-by: OpenCode / qwen3.8-27b-thinking (Alibaba Cloud) Signed-off-by: Stefan Weil <[email protected]> --- # Backport note: test-only hunks (Makefile.am, unittest/) dropped - # the OBS build runs no %%check. Source fix hunks are verbatim upstream. # Makefile.am | 5 ++ src/lstm/fullyconnected.cpp | 11 ++- src/lstm/weightmatrix.h | 7 ++ unittest/fullyconnected_test.cc | 115 ++++++++++++++++++++++++++++++++ 4 files changed, 137 insertions(+), 1 deletion(-) create mode 100644 unittest/fullyconnected_test.cc diff --git a/src/lstm/fullyconnected.cpp b/src/lstm/fullyconnected.cpp index a0ee1b02dd..d999f57529 100644 --- a/src/lstm/fullyconnected.cpp +++ b/src/lstm/fullyconnected.cpp @@ -121,7 +121,16 @@ bool FullyConnected::Serialize(TFile *fp) const { // Reads from the given file. Returns false in case of error. bool FullyConnected::DeSerialize(TFile *fp) { - return weights_.DeSerialize(IsTraining(), fp); + if (!weights_.DeSerialize(IsTraining(), fp)) { + return false; + } + // The weight matrix must match the declared sizes (the second dimension + // includes the bias column); otherwise Forward would read or write + // outside the scratch buffers sized from ni_ and no_. + if (weights_.Dim1() != no_ || weights_.Dim2() != ni_ + 1) { + return false; + } + return true; } // Runs forward propagation of activations on the input line. diff --git a/src/lstm/weightmatrix.h b/src/lstm/weightmatrix.h index eaca3ffb09..6c5ab23659 100644 --- a/src/lstm/weightmatrix.h +++ b/src/lstm/weightmatrix.h @@ -107,6 +107,13 @@ class WeightMatrix { int NumOutputs() const { return int_mode_ ? wi_.dim1() : wf_.dim1(); } + // The dimensions of the active weight matrix (wi_ in int mode, else wf_). + int Dim1() const { + return int_mode_ ? wi_.dim1() : wf_.dim1(); + } + int Dim2() const { + return int_mode_ ? wi_.dim2() : wf_.dim2(); + } // Provides one set of weights. Only used by peep weight maxpool. const TFloat *GetWeights(int index) const { return wf_[index]; ++++++ tesseract-CVE-2026-88049.patch ++++++ >From b494ac18925f9d9aff9ef5815475de9943ab19bf Mon Sep 17 00:00:00 2001 From: Stefan Weil <[email protected]> Date: Fri, 21 Aug 2026 16:41:51 +0200 Subject: [PATCH] Validate LSTM gate matrix dimensions against na_/no_ at load LSTM::DeSerialize read na_ from the untrusted TESSDATA_LSTM component and derived ns_ from the CI gate matrix's dim1, but never checked that the deserialized dimensions were mutually consistent. The forward pass sizes its buffers from na_, no_ and ns_ while the gate matrices drive their own dimensions, so a crafted .traineddata performed heap out-of-bounds writes and reads during the first recognition step, e.g. WriteTimeStepPart writing ns_ floats at offset ni_+nf_ into a source_ buffer sized from na_ (up to ~256 KB with an attacker-chosen gate matrix dim1). The bounds assertions added by 2f4d2f4 (CVE-2026-73066) covered only NetworkIO::CopyTimeStepGeneral and Randomize, leaving WriteTimeStepPart and AddTimeStepPart unguarded. Key changes: - weightmatrix.h: add Dim1()/Dim2() accessors for the active weight matrix (int or float), alongside the existing NumOutputs(). - lstm.cpp: reject the layer in DeSerialize unless na_ == ni_ + nf_ + (is_2d_ ? 2 : 1) * ns_, every deserialized gate has Dim1() == ns_ and Dim2() == na_ + 1 (the layout InitWeightsFloat always produces), ns_ == no_ for plain NT_LSTM/NT_LSTM_SUMMARY, and the softmax layer's sizes match ns_/no_ for the softmax variants. - networkio.cpp: add the same defense-in-depth bounds assertions to WriteTimeStepPart and AddTimeStepPart as 2f4d2f4 added to CopyTimeStepGeneral and Randomize. - unittest: add lstm_layer_test with crafted NT_LSTM layers for the na_ mismatch, gate dim1 mismatch, and gate dim2 mismatch cases (each rejected at load; on unpatched code the tests reach Forward and ASan catches the out-of-bounds write in WriteTimeStepPart), plus a positive control that a consistent layer loads and runs. Reported-by: Zhixi "Jace" Sun <[email protected]> Assisted-by: OpenCode / qwen3.8-27b-thinking (Alibaba Cloud) Signed-off-by: Stefan Weil <[email protected]> --- # Backport note: test-only hunks (Makefile.am, unittest/) dropped - # the OBS build runs no %%check. Source fix hunks are verbatim upstream. # Makefile.am | 5 + src/lstm/lstm.cpp | 20 ++++ src/lstm/networkio.cpp | 2 + unittest/lstm_layer_test.cc | 177 ++++++++++++++++++++++++++++++++++++ 4 files changed, 204 insertions(+) create mode 100644 unittest/lstm_layer_test.cc diff --git a/src/lstm/lstm.cpp b/src/lstm/lstm.cpp index 7bfd19b524..bfffa5deab 100644 --- a/src/lstm/lstm.cpp +++ b/src/lstm/lstm.cpp @@ -274,12 +274,32 @@ bool LSTM::DeSerialize(TFile *fp) { is_2d_ = na_ - nf_ == ni_ + 2 * ns_; } } + // The deserialized dimensions must be mutually consistent: the forward + // pass sizes its buffers from na_, no_ and ns_ while the gate matrices + // drive their own dimensions. + if (na_ != ni_ + nf_ + (is_2d_ ? 2 : 1) * ns_) { + return false; + } + for (int w = 0; w < WT_COUNT; ++w) { + if (w == GFS && !Is2D()) { + continue; + } + if (gate_weights_[w].Dim1() != ns_ || gate_weights_[w].Dim2() != na_ + 1) { + return false; + } + } + if ((type_ == NT_LSTM || type_ == NT_LSTM_SUMMARY) && ns_ != no_) { + return false; + } delete softmax_; if (type_ == NT_LSTM_SOFTMAX || type_ == NT_LSTM_SOFTMAX_ENCODED) { softmax_ = static_cast<FullyConnected *>(Network::CreateFromFile(fp)); if (softmax_ == nullptr) { return false; } + if (softmax_->NumInputs() != ns_ || softmax_->NumOutputs() != no_) { + return false; + } } else { softmax_ = nullptr; } diff --git a/src/lstm/networkio.cpp b/src/lstm/networkio.cpp index 8e1c6679c7..7a31034bf9 100644 --- a/src/lstm/networkio.cpp +++ b/src/lstm/networkio.cpp @@ -641,6 +641,7 @@ void NetworkIO::AddTimeStep(int t, TFloat *inout) const { // Adds part of a single timestep to floats. void NetworkIO::AddTimeStepPart(int t, int offset, int num_features, float *inout) const { + ASSERT_HOST(offset + num_features <= NumFeatures()); if (int_mode_) { const int8_t *line = i_[t] + offset; for (int i = 0; i < num_features; ++i) { @@ -662,6 +663,7 @@ void NetworkIO::WriteTimeStep(int t, const TFloat *input) { // Writes a single timestep from floats in the range [-1, 1] writing only // num_features elements of input to (*this)[t], starting at offset. void NetworkIO::WriteTimeStepPart(int t, int offset, int num_features, const TFloat *input) { + ASSERT_HOST(offset + num_features <= NumFeatures()); if (int_mode_) { int8_t *line = i_[t] + offset; for (int i = 0; i < num_features; ++i) { ++++++ tesseract-CVE-2026-88050.patch ++++++ >From c94a5532ee04db5a4919542832fd94caee5ea58f Mon Sep 17 00:00:00 2001 From: Stefan Weil <[email protected]> Date: Fri, 21 Aug 2026 16:56:51 +0200 Subject: [PATCH] Reject recoder code values outside the sane range at load RecodedCharID::DeSerialize (hardened by 82727cc to check length_ only) still read the individual code values as raw signed int32. UnicharCompress::ComputeCodeRange computes code_range_ as 1 plus the maximum code using a signed > comparison, so a code value of -1 never raises the maximum and yields code_range_ = 0. SetupDecoder then resizes is_valid_start_ to 0 and writes is_valid_start_[code(0)] on the size-0 vector<bool>, an out-of-bounds write at a wild wrapped index (deterministic crash) on LSTMRecognizer load. A code value of INT32_MAX instead wraps code_range_ negative and makes resize() throw. Key changes: - unicharcompress.h: validate each deserialized code value to be within [0, UINT16_MAX), the same arbitrary cap used elsewhere for .traineddata counts; reject the recoder otherwise. - unicharcompress.h/.cpp: add defense-in-depth bounds assertions to IsValidFirstCode and SetupDecoder, matching the style of the NetworkIO assertions from 2f4d2f4. - unittest: add recoder_test, which feeds a crafted UnicharCompress with a -1 code and an INT32_MAX code and expects DeSerialize to fail (on unpatched code the -1 case dies on the SEGV in SetupDecoder, the huge case on the uncaught length_error), plus a positive control that valid codes load and answer code_range()/IsValidFirstCode() correctly. Reported-by: Zhixi "Jace" Sun <[email protected]> Assisted-by: OpenCode / qwen3.8-27b-thinking (Alibaba Cloud) Signed-off-by: Stefan Weil <[email protected]> --- # Backport note: test-only hunks (Makefile.am, unittest/) dropped - # the OBS build runs no %%check. Source fix hunks are verbatim upstream. # Makefile.am | 5 ++ src/ccutil/unicharcompress.cpp | 1 + src/ccutil/unicharcompress.h | 13 ++++- unittest/recoder_test.cc | 91 ++++++++++++++++++++++++++++++++++ 4 files changed, 109 insertions(+), 1 deletion(-) create mode 100644 unittest/recoder_test.cc diff --git a/src/ccutil/unicharcompress.cpp b/src/ccutil/unicharcompress.cpp index f909f76eeb..b3dd440f26 100644 --- a/src/ccutil/unicharcompress.cpp +++ b/src/ccutil/unicharcompress.cpp @@ -400,6 +400,7 @@ void UnicharCompress::SetupDecoder() { for (unsigned c = 0; c < encoder_.size(); ++c) { const RecodedCharID &code = encoder_[c]; decoder_[code] = c; + ASSERT_HOST(code(0) >= 0 && code(0) < code_range_); is_valid_start_[code(0)] = true; RecodedCharID prefix = code; uint32_t len = code.length(); diff --git a/src/ccutil/unicharcompress.h b/src/ccutil/unicharcompress.h index 9b696d033c..6919b7d259 100644 --- a/src/ccutil/unicharcompress.h +++ b/src/ccutil/unicharcompress.h @@ -84,7 +84,17 @@ class RecodedCharID { if (length_ > kMaxCodeLen) { return false; } - return fp->DeSerialize(&code_[0], length_); + if (!fp->DeSerialize(&code_[0], length_)) { + return false; + } + // Code values index arrays sized from the maximum code; reject values + // that are out of the sane range for a recoded alphabet. + for (uint32_t i = 0; i < length_; ++i) { + if (code_[i] < 0 || code_[i] >= static_cast<int32_t>(UINT16_MAX)) { + return false; + } + } + return true; } bool operator==(const RecodedCharID &other) const { if (length_ != other.length_) { @@ -190,6 +200,7 @@ class TESS_API UnicharCompress { int DecodeUnichar(const RecodedCharID &code) const; // Returns true if the given code is a valid start or single code. bool IsValidFirstCode(int code) const { + ASSERT_HOST(code >= 0 && code < code_range_); return is_valid_start_[code]; } // Returns a list of valid non-final next codes for a given prefix code, ++++++ tesseract-CVE-2026-88051.patch ++++++ >From 56e09ca12e751623fe796ce1554ce704bffd2ef0 Mon Sep 17 00:00:00 2001 From: Stefan Weil <[email protected]> Date: Fri, 21 Aug 2026 20:29:37 +0200 Subject: [PATCH] Validate vector counts in GenericVector::read The callback form of GenericVector::read read two independent int32 fields from the file: reserved sized the allocation via reserve(), while size_used_ drove the element loop. Neither was capped and no size_used_ <= reserved invariant was checked, so a crafted .traineddata (e.g. the fontinfo table of a version >= 4 inttemp component) performed a heap out-of-bounds write during legacy engine initialization. Key changes: - genericvector.h: reject negative or over-limit reserved (matching the 50000000 cap of the DeSerialize overloads) and reject size_used_ < 0 or size_used_ > reserved before entering the read loop. Legit files always satisfy size_used_ <= reserved, as write() persists size_reserved_ first. - unittest: add genericvector_test covering size_used_ beyond reserved (on unpatched code the ASan build dies on a heap-buffer-overflow in the read loop), negative counts, and a consistent vector that must still load. Reported-by: Zhixi "Jace" Sun <[email protected]> Assisted-by: OpenCode / qwen3.8-27b-thinking (Alibaba Cloud) Signed-off-by: Stefan Weil <[email protected]> --- # Backport note: test-only hunks (Makefile.am, unittest/) dropped - # the OBS build runs no %%check. Source fix hunks are verbatim upstream. # Makefile.am | 5 ++ src/ccutil/genericvector.h | 10 ++++ unittest/genericvector_test.cc | 91 ++++++++++++++++++++++++++++++++++ 3 files changed, 106 insertions(+) create mode 100644 unittest/genericvector_test.cc diff --git a/src/ccutil/genericvector.h b/src/ccutil/genericvector.h index 4a5bbe12d6..cbe1e203d2 100644 --- a/src/ccutil/genericvector.h +++ b/src/ccutil/genericvector.h @@ -654,10 +654,20 @@ bool GenericVector<T>::read(TFile *f, const std::function<bool(TFile *, T *)> &c if (f->FReadEndian(&reserved, sizeof(reserved), 1) != 1) { return false; } + // Arbitrarily limit the number of elements to protect against bad data. + const uint32_t limit = 50000000; + if (reserved < 0 || static_cast<uint32_t>(reserved) > limit) { + return false; + } reserve(reserved); if (f->FReadEndian(&size_used_, sizeof(size_used_), 1) != 1) { return false; } + // size_used_ is an independent file field; without this check the reads + // below land past the end of the buffer sized from reserved. + if (size_used_ < 0 || size_used_ > reserved) { + return false; + } if (cb != nullptr) { for (int i = 0; i < size_used_; ++i) { if (!cb(f, data_ + i)) { ++++++ tesseract-CVE-2026-88052.patch ++++++ >From 2d04d640db2e8c7e3bab2369d599343b5a8b8443 Mon Sep 17 00:00:00 2001 From: Stefan Weil <[email protected]> Date: Fri, 21 Aug 2026 18:34:33 +0200 Subject: [PATCH] Reject unicharset files whose inserts desync id from unichars UNICHARSET::load_via_fgets reads the unichar count via sscanf and trusts it as the loop bound, indexing the unichars vector with the loop index id via the unchecked set_* accessors. unichar_insert is a no-op for duplicate (or empty) representations, so once any insert no-ops, unichars.size() falls behind id and the subsequent set_*(id, ...) and unichars[id].properties writes land past the end of the vector - a deterministic heap out-of-bounds write (including a std::string assignment via set_normed) for every remaining line, on both the LSTM and legacy init paths. A malformed unicharset with a duplicate line (e.g. two identical entries) triggers it; a non-positive header count likewise loads an empty unicharset "successfully". Key changes: - unicharset.cpp: reject unicharset_size <= 0, and after each insert verify the vector actually grew to id + 1; on mismatch report the offending line and reject the file instead of writing out of bounds. - unittest: add unicharset_load_test with a duplicate-representation unicharset (on unpatched code the test dies on the container-overflow in load_via_fgets), a zero and a negative count, and a positive control that a valid unicharset still loads. Reported-by: Zhixi "Jace" Sun <[email protected]> Assisted-by: OpenCode / qwen3.8-27b-thinking (Alibaba Cloud) Signed-off-by: Stefan Weil <[email protected]> --- # Backport note: test-only hunks (Makefile.am, unittest/) dropped - # the OBS build runs no %%check. Source fix hunks are verbatim upstream. # Makefile.am | 5 +++ src/ccutil/unicharset.cpp | 12 ++++++ unittest/unicharset_load_test.cc | 64 ++++++++++++++++++++++++++++++++ 3 files changed, 81 insertions(+) create mode 100644 unittest/unicharset_load_test.cc diff --git a/src/ccutil/unicharset.cpp b/src/ccutil/unicharset.cpp index b29ec3b7fe..0e72ae48d0 100644 --- a/src/ccutil/unicharset.cpp +++ b/src/ccutil/unicharset.cpp @@ -791,6 +791,9 @@ bool UNICHARSET::load_via_fgets( sscanf(buffer, "%d", &unicharset_size) != 1) { return false; } + if (unicharset_size <= 0) { + return false; + } for (UNICHAR_ID id = 0; id < unicharset_size; ++id) { char unichar[256]; unsigned int properties; @@ -884,6 +887,15 @@ bool UNICHARSET::load_via_fgets( } else { this->unichar_insert_backwards_compatible(unichar); } + // A duplicate or empty representation makes the insert a no-op, + // desynchronizing id from the unichars vector; the set_* calls and + // unichars[id] below would then write out of bounds. The file is + // malformed, so reject it. + if (size() != static_cast<size_t>(id) + 1) { + fprintf(stderr, "%s:%d unichar %d has a duplicate or empty representation\n", + __FILE__, __LINE__, id); + return false; + } this->set_isalpha(id, properties & ISALPHA_MASK); this->set_islower(id, properties & ISLOWER_MASK); ++++++ tesseract-CVE-2026-88053.patch ++++++ >From 8b0574680f3b22f246ade6a4c8e3029104255c63 Mon Sep 17 00:00:00 2001 From: Stefan Weil <[email protected]> Date: Fri, 21 Aug 2026 13:47:58 +0200 Subject: [PATCH] Fix out-of-bounds writes in .traineddata inttemp deserialization Classify::ReadIntTemplates reads NumClassPruners, NumClasses and NumProtoSets from the untrusted inttemp component of a .traineddata file and used them as loop bounds writing pointers into fixed-size arrays (ClassPruners[], Class[], ProtoSets[]) without validation. A crafted file with counts larger than the array capacity caused heap out-of-bounds pointer writes during legacy engine initialization, before any OCR is performed. Key changes: - intproto.cpp: validate unicharset_size, NumClassPruners, NumClasses, per-class NumProtos/NumProtoSets/NumConfigs and the old-format (version < 2) class ids against the array capacities and fail the load instead of writing out of bounds. Never trust file-sourced counts to bound the destructor loops. - adaptive.cpp/h: same treatment in ReadAdaptedTemplates and ReadAdaptedClass (NumTempProtos, NumConfigs); the ADAPT_TEMPLATES_STRUCT default constructor now initializes all members and the destructor is null-safe. - adaptmatch.cpp, tface.cpp, tessedit.cpp: propagate the read failure through InitAdaptiveClassifier and program_editup so the language load fails gracefully instead of continuing with a broken state. - intproto.h, adaptive.h: replace the fixed-size member arrays by std::array (identical on-disk layout, value-initialized). - unittest: add intproto_test, which builds a minimal traineddata with corrupt inttemp counts and expects a graceful init failure. On unpatched code the test trips ASan on the original heap-buffer-overflow in ReadIntTemplates. Reported-by: Zhixi "Jace" Sun <[email protected]> Assisted-by: OpenCode / qwen3.8-27b-thinking (Alibaba Cloud) Signed-off-by: Stefan Weil <[email protected]> --- # Backport note: test-only hunks (Makefile.am, unittest/) dropped - # the OBS build runs no %%check. Upstream's std::array modernization # (intproto.h, adaptive.h) omitted: 5.5.3 still passes these members # as bare pointers elsewhere (eg adaptmatch.cpp MasterMatcher call, # intproto.cpp AddIntProto loop). intproto.cpp ctor/dtor hunk replaced # with a minimal destructor loop-bound fix for the 5.5.3 code. # Makefile.am | 7 ++ src/ccmain/tessedit.cpp | 4 +- src/classify/adaptive.cpp | 91 +++++++++++++++++--------- src/classify/adaptive.h | 7 +- src/classify/adaptmatch.cpp | 27 +++++--- src/classify/classify.h | 2 +- src/classify/intproto.cpp | 77 +++++++++++++++------- src/classify/intproto.h | 12 ++-- src/wordrec/tface.cpp | 7 +- src/wordrec/wordrec.h | 4 +- unittest/CMakeLists.txt | 1 + unittest/intproto_test.cc | 124 ++++++++++++++++++++++++++++++++++++ 12 files changed, 288 insertions(+), 75 deletions(-) create mode 100644 unittest/intproto_test.cc diff --git a/src/ccmain/tessedit.cpp b/src/ccmain/tessedit.cpp index 412b8c15be..cfbeca6211 100644 --- a/src/ccmain/tessedit.cpp +++ b/src/ccmain/tessedit.cpp @@ -418,7 +418,9 @@ int Tesseract::init_tesseract_internal(const std::string &textbase, // If only LSTM will be used, skip loading Tesseract classifier's // pre-trained templates and dictionary. bool init_tesseract = tessedit_ocr_engine_mode != OEM_LSTM_ONLY; - program_editup(textbase, init_tesseract ? mgr : nullptr, init_tesseract ? mgr : nullptr); + if (!program_editup(textbase, init_tesseract ? mgr : nullptr, init_tesseract ? mgr : nullptr)) { + return -1; + } return 0; // Normal exit } diff --git a/src/classify/adaptive.cpp b/src/classify/adaptive.cpp index e3297c0e0c..96850ab75b 100644 --- a/src/classify/adaptive.cpp +++ b/src/classify/adaptive.cpp @@ -67,10 +67,6 @@ ADAPT_CLASS_STRUCT::ADAPT_CLASS_STRUCT() : TempProtos(NIL_LIST) { zero_all_bits(PermProtos, WordsInVectorOfSize(MAX_NUM_PROTOS)); zero_all_bits(PermConfigs, WordsInVectorOfSize(MAX_NUM_CONFIGS)); - - for (int i = 0; i < MAX_NUM_CONFIGS; i++) { - TempConfigFor(this, i) = nullptr; - } } ADAPT_CLASS_STRUCT::~ADAPT_CLASS_STRUCT() { @@ -92,25 +88,24 @@ ADAPT_CLASS_STRUCT::~ADAPT_CLASS_STRUCT() { /// Constructor for adapted templates. /// Add an empty class for each char in unicharset to the newly created templates. -ADAPT_TEMPLATES_STRUCT::ADAPT_TEMPLATES_STRUCT(UNICHARSET &unicharset) { - Templates = new INT_TEMPLATES_STRUCT; - NumPermClasses = 0; - NumNonEmptyClasses = 0; - - /* Insert an empty class for each unichar id in unicharset */ - for (unsigned i = 0; i < MAX_NUM_CLASSES; i++) { - Class[i] = nullptr; - if (i < unicharset.size()) { - AddAdaptedClass(this, new ADAPT_CLASS_STRUCT, i); - } +ADAPT_TEMPLATES_STRUCT::ADAPT_TEMPLATES_STRUCT(UNICHARSET &unicharset) : + Templates(new INT_TEMPLATES_STRUCT), NumNonEmptyClasses(0), NumPermClasses(0) { + // Insert an empty class for each unichar id in unicharset. + // Class is value-initialized to nullptr in-class. + for (unsigned i = 0; i < unicharset.size(); i++) { + AddAdaptedClass(this, new ADAPT_CLASS_STRUCT, i); } } ADAPT_TEMPLATES_STRUCT::~ADAPT_TEMPLATES_STRUCT() { - for (unsigned i = 0; i < (Templates)->NumClasses; i++) { - delete Class[i]; + if (Templates != nullptr) { + // NumClasses comes from an untrusted file, so never trust it to bound + // the loop over the fixed-size Class[] array. + for (unsigned i = 0; i < (Templates)->NumClasses && i < MAX_NUM_CLASSES; i++) { + delete Class[i]; + } + delete Templates; } - delete Templates; } // Returns FontinfoId of the given config of the given adapted class. @@ -180,23 +175,32 @@ void Classify::PrintAdaptedTemplates(FILE *File, ADAPT_TEMPLATES_STRUCT *Templat * @note Globals: none */ ADAPT_CLASS_STRUCT *ReadAdaptedClass(TFile *fp) { - int NumTempProtos; - int NumConfigs; + int NumTempProtos = 0; + int NumConfigs = 0; int i; ADAPT_CLASS_STRUCT *Class; - /* first read high level adapted class structure */ + // first read high level adapted class structure Class = new ADAPT_CLASS_STRUCT; fp->FRead(Class, sizeof(ADAPT_CLASS_STRUCT), 1); - /* then read in the definitions of the permanent protos and configs */ + // then read in the definitions of the permanent protos and configs Class->PermProtos = NewBitVector(MAX_NUM_PROTOS); Class->PermConfigs = NewBitVector(MAX_NUM_CONFIGS); fp->FRead(Class->PermProtos, sizeof(uint32_t), WordsInVectorOfSize(MAX_NUM_PROTOS)); fp->FRead(Class->PermConfigs, sizeof(uint32_t), WordsInVectorOfSize(MAX_NUM_CONFIGS)); - /* then read in the list of temporary protos */ + // then read in the list of temporary protos fp->FRead(&NumTempProtos, sizeof(int), 1); + if (NumTempProtos < 0 || NumTempProtos > MAX_NUM_PROTOS) { + tprintf("Bad read of adapted class!\n"); + // Reset file-sourced pointers so the destructor does not delete them. + for (i = 0; i < MAX_NUM_CONFIGS; i++) { + Class->Config[i].Temp = nullptr; + } + delete Class; + return nullptr; + } Class->TempProtos = NIL_LIST; for (i = 0; i < NumTempProtos; i++) { auto TempProto = new TEMP_PROTO_STRUCT; @@ -204,8 +208,20 @@ ADAPT_CLASS_STRUCT *ReadAdaptedClass(TFile *fp) { Class->TempProtos = push_last(Class->TempProtos, TempProto); } - /* then read in the adapted configs */ + // then read in the adapted configs fp->FRead(&NumConfigs, sizeof(int), 1); + // NumConfigs is used as a loop bound that writes into the fixed-size + // Config[] array, so reject a corrupt or malicious file instead of + // writing out of bounds. + if (NumConfigs < 0 || NumConfigs > MAX_NUM_CONFIGS) { + tprintf("Bad read of adapted class!\n"); + // Reset file-sourced pointers so the destructor does not delete them. + for (i = 0; i < MAX_NUM_CONFIGS; i++) { + Class->Config[i].Temp = nullptr; + } + delete Class; + return nullptr; + } for (i = 0; i < NumConfigs; i++) { if (test_bit(Class->PermConfigs, i)) { Class->Config[i].Perm = ReadPermConfig(fp); @@ -231,18 +247,35 @@ ADAPT_CLASS_STRUCT *ReadAdaptedClass(TFile *fp) { ADAPT_TEMPLATES_STRUCT *Classify::ReadAdaptedTemplates(TFile *fp) { auto Templates = new ADAPT_TEMPLATES_STRUCT; - /* first read the high level adaptive template struct */ - fp->FRead(Templates, sizeof(ADAPT_TEMPLATES_STRUCT), 1); + // first read in the high level adaptive template struct + if (fp->FRead(Templates, sizeof(ADAPT_TEMPLATES_STRUCT), 1) != 1) { + tprintf("Bad read of adapted templates!\n"); + delete Templates; + return nullptr; + } + // The Class[] array was just filled with pointers read from the file; + // those are not valid allocations, so reset it before storing real ones. + for (unsigned i = 0; i < MAX_NUM_CLASSES; i++) { + Templates->Class[i] = nullptr; + } - /* then read in the basic integer templates */ + // then read in the basic integer templates Templates->Templates = ReadIntTemplates(fp); + if (Templates->Templates == nullptr) { + delete Templates; + return nullptr; + } - /* then read in the adaptive info for each class */ + // then read in the adaptive info for each class for (unsigned i = 0; i < (Templates->Templates)->NumClasses; i++) { Templates->Class[i] = ReadAdaptedClass(fp); + if (Templates->Class[i] == nullptr) { + tprintf("Bad read of adapted templates (class %u)!\n", i); + delete Templates; + return nullptr; + } } return (Templates); - } /* ReadAdaptedTemplates */ /*---------------------------------------------------------------------------*/ diff --git a/src/classify/adaptive.h b/src/classify/adaptive.h index 652e1fbdcb..fde3ef2270 100644 --- a/src/classify/adaptive.h +++ b/src/classify/adaptive.h @@ -67,9 +67,9 @@ class ADAPT_TEMPLATES_STRUCT { public: - ADAPT_TEMPLATES_STRUCT() = default; + ADAPT_TEMPLATES_STRUCT() : Templates(nullptr), NumNonEmptyClasses(0), NumPermClasses(0) {} ADAPT_TEMPLATES_STRUCT(UNICHARSET &unicharset); ~ADAPT_TEMPLATES_STRUCT(); INT_TEMPLATES_STRUCT *Templates; int NumNonEmptyClasses; uint8_t NumPermClasses; ADAPT_CLASS_STRUCT *Class[MAX_NUM_CLASSES]; }; diff --git a/src/classify/adaptmatch.cpp b/src/classify/adaptmatch.cpp index f9f38fbe67..30d30a810e 100644 --- a/src/classify/adaptmatch.cpp +++ b/src/classify/adaptmatch.cpp @@ -524,9 +524,9 @@ void Classify::EndAdaptiveClassifier() { * classify_use_pre_adapted_templates * enables use of pre-adapted templates */ -void Classify::InitAdaptiveClassifier(TessdataManager *mgr) { +bool Classify::InitAdaptiveClassifier(TessdataManager *mgr) { if (!CLASSIFY_ENABLE_ADAPTIVE_MATCHER_OVERRIDE) { - return; + return true; } if (AllProtosOn != nullptr) { EndAdaptiveClassifier(); // Don't leak with multiple inits. @@ -538,6 +538,11 @@ void Classify::InitAdaptiveClassifier(TessdataManager *mgr) { TFile fp; ASSERT_HOST(mgr->GetComponent(TESSDATA_INTTEMP, &fp)); PreTrainedTemplates = ReadIntTemplates(&fp); + if (PreTrainedTemplates == nullptr) { + tprintf("Error: invalid inttemp component in traineddata, " + "cannot initialize the legacy engine.\n"); + return false; + } if (mgr->GetComponent(TESSDATA_SHAPE_TABLE, &fp)) { shape_table_ = new ShapeTable(unicharset); @@ -580,17 +585,23 @@ void Classify::InitAdaptiveClassifier(TessdataManager *mgr) { tprintf("\nReading pre-adapted templates from %s ...\n", Filename.c_str()); fflush(stdout); AdaptedTemplates = ReadAdaptedTemplates(&fp); - tprintf("\n"); - PrintAdaptedTemplates(stdout, AdaptedTemplates); + if (AdaptedTemplates == nullptr) { + tprintf("Error: invalid pre-adapted templates in %s, ignoring.\n", Filename.c_str()); + AdaptedTemplates = new ADAPT_TEMPLATES_STRUCT(unicharset); + } else { + tprintf("\n"); + PrintAdaptedTemplates(stdout, AdaptedTemplates); - for (unsigned i = 0; i < AdaptedTemplates->Templates->NumClasses; i++) { - BaselineCutoffs[i] = CharNormCutoffs[i]; + for (unsigned i = 0; i < AdaptedTemplates->Templates->NumClasses; i++) { + BaselineCutoffs[i] = CharNormCutoffs[i]; + } } } } else { delete AdaptedTemplates; AdaptedTemplates = new ADAPT_TEMPLATES_STRUCT(unicharset); } + return true; } /* InitAdaptiveClassifier */ void Classify::ResetAdaptiveClassifierInternal() { @@ -1243,8 +1254,8 @@ UNICHAR_ID *Classify::BaselineClassifier(TBLOB *Blob, } MasterMatcher(Templates->Templates, int_features.size(), &int_features[0], CharNormArray, - Templates->Class, matcher_debug_flags, 0, Blob->bounding_box(), Results->CPResults, - Results); + Templates->Class, matcher_debug_flags, 0, Blob->bounding_box(), + Results->CPResults, Results); delete[] CharNormArray; CLASS_ID ClassId = Results->best_unichar_id; diff --git a/src/classify/classify.h b/src/classify/classify.h index 1a511c2826..8ebf451f31 100644 --- a/src/classify/classify.h +++ b/src/classify/classify.h @@ -164,7 +164,7 @@ class TESS_API Classify : public CCStruct { // provided to explicitly clarify the character segmentation. void LearnPieces(const char *fontname, int start, int length, float threshold, CharSegmentationType segmentation, const char *correct_text, WERD_RES *word); - void InitAdaptiveClassifier(TessdataManager *mgr); + bool InitAdaptiveClassifier(TessdataManager *mgr); void InitAdaptedClass(TBLOB *Blob, CLASS_ID ClassId, int FontinfoId, ADAPT_CLASS_STRUCT *Class, ADAPT_TEMPLATES_STRUCT *Templates); void AmbigClassifier(const std::vector<INT_FEATURE_STRUCT> &int_features, diff --git a/src/classify/intproto.cpp b/src/classify/intproto.cpp index b199ae2893..15cd260d9c 100644 --- a/src/classify/intproto.cpp +++ b/src/classify/intproto.cpp @@ -596,5 +596,7 @@ INT_CLASS_STRUCT::~INT_CLASS_STRUCT() { INT_CLASS_STRUCT::~INT_CLASS_STRUCT() { - for (int i = 0; i < NumProtoSets; i++) { + // NumProtoSets comes from an untrusted file, so never trust it to bound + // the loop over the fixed-size ProtoSets[] array. + for (int i = 0; i < NumProtoSets && i < MAX_NUM_PROTO_SETS; i++) { delete ProtoSets[i]; } } @@ -613,8 +613,10 @@ INT_TEMPLATES_STRUCT::~INT_TEMPLATES_STRUCT() { INT_TEMPLATES_STRUCT::~INT_TEMPLATES_STRUCT() { - for (unsigned i = 0; i < NumClasses; i++) { + // The counts come from an untrusted file, so never trust them to bound + // the loops over the fixed-size arrays. + for (unsigned i = 0; i < NumClasses && i < MAX_NUM_CLASSES; i++) { delete Class[i]; } - for (unsigned i = 0; i < NumClassPruners; i++) { + for (unsigned i = 0; i < NumClassPruners && i < MAX_NUM_CLASS_PRUNERS; i++) { delete ClassPruners[i]; } } @@ -669,6 +663,20 @@ INT_TEMPLATES_STRUCT *Classify::ReadIntTemplates(TFile *fp) { Templates->NumClasses = version_id; } + // The counts read from the file are used as loop bounds that write into + // fixed-size arrays (Class[], ClassPruners[], TempClassPruner[] and + // IndexFor[]), so reject a corrupt or malicious file instead of writing + // out of bounds. + if (unicharset_size > MAX_NUM_CLASSES || + Templates->NumClassPruners > MAX_NUM_CLASS_PRUNERS || + Templates->NumClasses > MAX_NUM_CLASSES) { + tprintf("Error: invalid counts in inttemp: unicharset_size=%u, NumClassPruners=%u, " + "NumClasses=%u\n", + unicharset_size, Templates->NumClassPruners, Templates->NumClasses); + delete Templates; + return nullptr; + } + if (version_id < 3) { MaxNumConfigs = OLD_MAX_NUM_CONFIGS; WerdsPerConfigVec = OLD_WERDS_PER_CONFIG_VEC; @@ -708,6 +716,16 @@ INT_TEMPLATES_STRUCT *Classify::ReadIntTemplates(TFile *fp) { max_class_id = ClassIdFor[i]; } } + // Class ids index Class[] and (divided by CLASSES_PER_CP) ClassPruners[], + // so reject a corrupt or malicious file instead of writing out of bounds. + if (max_class_id >= MAX_NUM_CLASSES) { + tprintf("Error: class id %u in inttemp exceeds MAX_NUM_CLASSES\n", max_class_id); + for (unsigned i = 0; i < Templates->NumClassPruners; i++) { + delete TempClassPruner[i]; + } + delete Templates; + return nullptr; + } for (int i = 0; i <= CPrunerIdFor(max_class_id); i++) { Templates->ClassPruners[i] = new CLASS_PRUNER_STRUCT; memset(Templates->ClassPruners[i], 0, sizeof(CLASS_PRUNER_STRUCT)); @@ -777,8 +795,20 @@ INT_TEMPLATES_STRUCT *Classify::ReadIntTemplates(TFile *fp) { } } unsigned num_configs = version_id < 4 ? MaxNumConfigs : Class->NumConfigs; - ASSERT_HOST(num_configs <= MaxNumConfigs); - if (fp->FReadEndian(Class->ConfigLengths, sizeof(uint16_t), num_configs) != num_configs) { + // Class->NumProtoSets is used as a loop bound that writes into the + // fixed-size ProtoSets[] array, so reject a corrupt or malicious file + // instead of writing out of bounds. + if (Class->NumProtos > MAX_NUM_PROTOS || Class->NumProtoSets > MAX_NUM_PROTO_SETS || + num_configs > MaxNumConfigs) { + tprintf("Error: invalid counts for class %u in inttemp: NumProtos=%u, NumProtoSets=%u, " + "NumConfigs=%u\n", + i, Class->NumProtos, Class->NumProtoSets, Class->NumConfigs); + Class->NumProtoSets = 0; // no proto sets allocated yet; keep destructor safe + delete Class; + delete Templates; + return nullptr; + } + if (fp->FReadEndian(Class->ConfigLengths, sizeof(uint16_t), num_configs) != num_configs) { tprintf("Bad read of inttemp!\n"); } if (version_id < 2) { @@ -812,8 +842,9 @@ INT_TEMPLATES_STRUCT *Classify::ReadIntTemplates(TFile *fp) { fp->FRead(&ProtoSet->Protos[x].Angle, sizeof(ProtoSet->Protos[x].Angle), 1) != 1) { tprintf("Bad read of inttemp!\n"); } - if (fp->FReadEndian(&ProtoSet->Protos[x].Configs, sizeof(ProtoSet->Protos[x].Configs[0]), - WerdsPerConfigVec) != WerdsPerConfigVec) { + if (fp->FReadEndian(ProtoSet->Protos[x].Configs, + sizeof(ProtoSet->Protos[x].Configs[0]), WerdsPerConfigVec) != + WerdsPerConfigVec) { tprintf("Bad read of inttemp!\n"); } } diff --git a/src/wordrec/tface.cpp b/src/wordrec/tface.cpp index c058d4abed..146f0b01ff 100644 --- a/src/wordrec/tface.cpp +++ b/src/wordrec/tface.cpp @@ -36,14 +36,16 @@ namespace tesseract { * init_permute determines whether to initialize the permute functions * and Dawg models. */ -void Wordrec::program_editup(const std::string &textbase, TessdataManager *init_classifier, +bool Wordrec::program_editup(const std::string &textbase, TessdataManager *init_classifier, TessdataManager *init_dict) { if (!textbase.empty()) { imagefile = textbase; } #ifndef DISABLED_LEGACY_ENGINE InitFeatureDefs(&feature_defs_); - InitAdaptiveClassifier(init_classifier); + if (!InitAdaptiveClassifier(init_classifier)) { + return false; + } if (init_dict) { getDict().SetupForLoad(Dict::GlobalDawgCache()); getDict().Load(lang, init_dict); @@ -51,6 +53,7 @@ void Wordrec::program_editup(const std::string &textbase, TessdataManager *init_ } pass2_ok_split = chop_ok_split; #endif // ndef DISABLED_LEGACY_ENGINE + return true; } /** diff --git a/src/wordrec/wordrec.h b/src/wordrec/wordrec.h index a50ad871ae..ecf67f02d4 100644 --- a/src/wordrec/wordrec.h +++ b/src/wordrec/wordrec.h @@ -50,7 +50,7 @@ class TESS_API Wordrec : public Classify { virtual ~Wordrec() = default; // tface.cpp - void program_editup(const std::string &textbase, TessdataManager *init_classifier, + bool program_editup(const std::string &textbase, TessdataManager *init_classifier, TessdataManager *init_dict); void program_editdown(); int end_recog(); @@ -243,7 +243,7 @@ class TESS_API Wordrec : public Classify { } // tface.cpp - void program_editup(const std::string &textbase, TessdataManager *init_classifier, + bool program_editup(const std::string &textbase, TessdataManager *init_classifier, TessdataManager *init_dict); void cc_recog(WERD_RES *word); void program_editdown(); ++++++ tesseract-CVE-2026-88054.patch ++++++ >From 552771236b0d80cbdb0c7dd856120fa21a4672e5 Mon Sep 17 00:00:00 2001 From: Stefan Weil <[email protected]> Date: Fri, 21 Aug 2026 14:38:55 +0200 Subject: [PATCH] Reject empty network stacks in LSTM .traineddata deserialization Plumbing::DeSerialize read the network stack size from the untrusted TESSDATA_LSTM component and only rejected size > 10000. A crafted .traineddata with a top-level Series/Parallel/Reversed layer with an empty stack survived load; LSTMRecognizer initialization then called network_->CacheXScaleFactor(network_->XScaleFactor()), which dereferences stack_[0] on the empty vector: - Series: Series::CacheXScaleFactor (series.cpp) - Parallel/Reversed: inherited Plumbing::XScaleFactor The virtual call through the wild pointer crashed the process at initialization (deterministic denial of service). Key changes: - plumbing.cpp: reject size == 0 for all plumbing types, and size < 2 for NT_SERIES (Series::Forward requires two or more networks and always aborts on one). - tessedit.cpp: fail the language load gracefully when the LSTM model cannot be loaded, instead of aborting via ASSERT_HOST, so TessBaseAPI::Init returns -1 on corrupt traineddata. - unittest: add plumbing_test, which builds a minimal traineddata with empty/undersized LSTM plumbing stacks and expects a graceful init failure. On unpatched code the tests die on the original SEGV in Series::CacheXScaleFactor / Plumbing::XScaleFactor. Reported-by: Zhixi "Jace" Sun <[email protected]> Assisted-by: OpenCode / qwen3.8-27b-thinking (Alibaba Cloud) Signed-off-by: Stefan Weil <[email protected]> --- # Backport note: test-only hunks (Makefile.am, unittest/) dropped - # the OBS build runs no %%check. Source fix hunks are verbatim upstream. # Makefile.am | 5 ++ src/ccmain/tessedit.cpp | 7 +- src/lstm/plumbing.cpp | 6 ++ unittest/plumbing_test.cc | 165 ++++++++++++++++++++++++++++++++++++++ 4 files changed, 182 insertions(+), 1 deletion(-) create mode 100644 unittest/plumbing_test.cc diff --git a/src/ccmain/tessedit.cpp b/src/ccmain/tessedit.cpp index cfbeca6211..aaa1599bfc 100644 --- a/src/ccmain/tessedit.cpp +++ b/src/ccmain/tessedit.cpp @@ -169,7 +169,12 @@ bool Tesseract::init_tesseract_lang_data(const std::string &language, OcrEngineM #endif // ndef DISABLED_LEGACY_ENGINE if (mgr->IsComponentAvailable(TESSDATA_LSTM)) { lstm_recognizer_ = new LSTMRecognizer(language_data_path_prefix.c_str()); - ASSERT_HOST(lstm_recognizer_->Load(this->params(), lstm_use_matrix ? language : "", mgr)); + if (!lstm_recognizer_->Load(this->params(), lstm_use_matrix ? language : "", mgr)) { + delete lstm_recognizer_; + lstm_recognizer_ = nullptr; + tprintf("Error: Failed to load the LSTM model from %s\n", tessdata_path.c_str()); + return false; + } } else { #ifdef DISABLED_LEGACY_ENGINE // The legacy engine is compiled out, so we cannot fall back to it. diff --git a/src/lstm/plumbing.cpp b/src/lstm/plumbing.cpp index 7288bd02b5..908a751381 100644 --- a/src/lstm/plumbing.cpp +++ b/src/lstm/plumbing.cpp @@ -228,6 +228,12 @@ bool Plumbing::DeSerialize(TFile *fp) { if (size > 10000) { return false; } + // Reject empty stacks: XScaleFactor, CacheXScaleFactor and other methods + // unconditionally dereference stack_[0] during network initialization. + // A Series needs at least two networks (see Series::Forward). + if (size == 0 || (type() == NT_SERIES && size == 1)) { + return false; + } for (uint32_t i = 0; i < size; ++i) { Network *network = CreateFromFile(fp); if (network == nullptr) {
