Script 'mail_helper' called by obssrc
Hello community,

here is the log from the commit of package tesseract-ocr for openSUSE:Factory 
checked in at 2026-09-24 22:55:32
++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Comparing /work/SRC/openSUSE:Factory/tesseract-ocr (Old)
 and      /work/SRC/openSUSE:Factory/.tesseract-ocr.new.383539 (New)
++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++

Package is "tesseract-ocr"

Thu Sep 24 22:55:32 2026 rev:26 rq:1379978 version:5.5.3

Changes:
--------
--- /work/SRC/openSUSE:Factory/tesseract-ocr/tesseract-ocr.changes      
2026-07-28 17:49:24.766645934 +0200
+++ /work/SRC/openSUSE:Factory/.tesseract-ocr.new.383539/tesseract-ocr.changes  
2026-09-24 22:55:49.589114671 +0200
@@ -1,0 +2,34 @@
+Wed Sep 23 12:35:46 UTC 2026 - Martin Pluskal <[email protected]>
+
+- CVE-2026-88047: stack buffer overflow in
+  Classify::ReadNormProtos on crafted traineddata (boo#1280925)
+  * tesseract-CVE-2026-88047.patch
+- CVE-2026-88048: heap out-of-bounds write/read in
+  FullyConnected::Forward via dimension mismatch (boo#1280929)
+  * tesseract-CVE-2026-88048.patch
+- CVE-2026-88049: heap out-of-bounds write in LSTM::Forward
+  via na_/gate-matrix dimension mismatch (boo#1280930)
+  * tesseract-CVE-2026-88049.patch
+- CVE-2026-88050: out-of-bounds write in UnicharCompress
+  via unvalidated recoder code values (boo#1280931)
+  * tesseract-CVE-2026-88050.patch
+- CVE-2026-88051: heap out-of-bounds write in
+  GenericVector<T>::read via reserved/size_used_ mismatch
+  (boo#1280932)
+  * tesseract-CVE-2026-88051.patch
+- CVE-2026-88052: heap out-of-bounds write in
+  UNICHARSET::load_via_fgets via count/insert
+  desynchronization (boo#1280933)
+  * tesseract-CVE-2026-88052.patch
+- CVE-2026-88053: heap out-of-bounds write in
+  Classify::ReadIntTemplates via unvalidated counts in
+  crafted traineddata (boo#1280934)
+  * tesseract-CVE-2026-88053.patch
+- CVE-2026-88054: denial of service via empty-stack
+  dereference at model load (boo#1280935)
+  * tesseract-CVE-2026-88054.patch
+- CVE-2026-73067 (boo#1275623): heap out-of-bounds read in
+  SquishedDawg on crafted model, already fixed in the
+  shipped 5.5.3 (DAWG edge-structure validation).
+
+-------------------------------------------------------------------

New:
----
  tesseract-CVE-2026-88047.patch
  tesseract-CVE-2026-88048.patch
  tesseract-CVE-2026-88049.patch
  tesseract-CVE-2026-88050.patch
  tesseract-CVE-2026-88051.patch
  tesseract-CVE-2026-88052.patch
  tesseract-CVE-2026-88053.patch
  tesseract-CVE-2026-88054.patch

----------(New B)----------
  New:  Classify::ReadNormProtos on crafted traineddata (boo#1280925)
  * tesseract-CVE-2026-88047.patch
- CVE-2026-88048: heap out-of-bounds write/read in
  New:  FullyConnected::Forward via dimension mismatch (boo#1280929)
  * tesseract-CVE-2026-88048.patch
- CVE-2026-88049: heap out-of-bounds write in LSTM::Forward
  New:  via na_/gate-matrix dimension mismatch (boo#1280930)
  * tesseract-CVE-2026-88049.patch
- CVE-2026-88050: out-of-bounds write in UnicharCompress
  New:  via unvalidated recoder code values (boo#1280931)
  * tesseract-CVE-2026-88050.patch
- CVE-2026-88051: heap out-of-bounds write in
  New:  (boo#1280932)
  * tesseract-CVE-2026-88051.patch
- CVE-2026-88052: heap out-of-bounds write in
  New:  desynchronization (boo#1280933)
  * tesseract-CVE-2026-88052.patch
- CVE-2026-88053: heap out-of-bounds write in
  New:  crafted traineddata (boo#1280934)
  * tesseract-CVE-2026-88053.patch
- CVE-2026-88054: denial of service via empty-stack
  New:  dereference at model load (boo#1280935)
  * tesseract-CVE-2026-88054.patch
- CVE-2026-73067 (boo#1275623): heap out-of-bounds read in
----------(New E)----------

++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++

Other differences:
------------------
++++++ tesseract-ocr.spec ++++++
--- /var/tmp/diff_new_pack.YgKCla/_old  2026-09-24 22:55:50.297144260 +0200
+++ /var/tmp/diff_new_pack.YgKCla/_new  2026-09-24 22:55:50.299144343 +0200
@@ -25,6 +25,22 @@
 URL:            https://github.com/tesseract-ocr/tesseract
 Source0:        
https://github.com/tesseract-ocr/tesseract/archive/refs/tags/%{version}.tar.gz#/tesseract-%{version}.tar.gz
 Source99:       baselibs.conf
+# PATCH-FIX-UPSTREAM tesseract-CVE-2026-88047.patch GHSA-5j2p-r5vc-q7f3 
(upstream commit 1bda5079) -- bound stack buffer in Classify::ReadNormProtos 
(CVE-2026-88047, boo#1280925)
+Patch0:         tesseract-CVE-2026-88047.patch
+# PATCH-FIX-UPSTREAM tesseract-CVE-2026-88048.patch GHSA-q44c-23p6-5mw6 
(upstream commit 103dc134) -- validate FullyConnected layer dims vs weight 
matrix (CVE-2026-88048, boo#1280929)
+Patch1:         tesseract-CVE-2026-88048.patch
+# PATCH-FIX-UPSTREAM tesseract-CVE-2026-88049.patch GHSA-jgq8-pprg-vc68 
(upstream commit b494ac18) -- bound LSTM::Forward WriteTimeStepPart count 
(CVE-2026-88049, boo#1280930)
+Patch2:         tesseract-CVE-2026-88049.patch
+# PATCH-FIX-UPSTREAM tesseract-CVE-2026-88050.patch GHSA-7v9h-3q3m-w68g 
(upstream commit c94a5532) -- reject negative recoder code values 
(CVE-2026-88050, boo#1280931)
+Patch3:         tesseract-CVE-2026-88050.patch
+# PATCH-FIX-UPSTREAM tesseract-CVE-2026-88051.patch GHSA-88qp-4g94-3rf3 
(upstream commit 56e09ca1) -- cap GenericVector::read reserved/size_used_ 
(CVE-2026-88051, boo#1280932)
+Patch4:         tesseract-CVE-2026-88051.patch
+# PATCH-FIX-UPSTREAM tesseract-CVE-2026-88052.patch GHSA-2hm8-q5c7-c373 
(upstream commit 2d04d640) -- validate UNICHARSET load loop bound and index 
(CVE-2026-88052, boo#1280933)
+Patch5:         tesseract-CVE-2026-88052.patch
+# PATCH-FIX-UPSTREAM tesseract-CVE-2026-88053.patch GHSA-rphx-x795-5qjv 
(upstream commit 8b057468) -- validate inttemp counts against MAX_* bounds 
(CVE-2026-88053, boo#1280934)
+Patch6:         tesseract-CVE-2026-88053.patch
+# PATCH-FIX-UPSTREAM tesseract-CVE-2026-88054.patch GHSA-f6h7-cqr4-6fx4 
(upstream commit 55277123) -- reject zero-length network stack in 
Plumbing::DeSerialize (CVE-2026-88054, boo#1280935)
+Patch7:         tesseract-CVE-2026-88054.patch
 BuildRequires:  autoconf
 BuildRequires:  automake
 BuildRequires:  curl-devel

++++++ tesseract-CVE-2026-88047.patch ++++++
>From 1bda5079b1c8a7e25f523486837426903d29ce84 Mon Sep 17 00:00:00 2001
From: Stefan Weil <[email protected]>
Date: Mon, 24 Aug 2026 13:39:22 +0200
Subject: [PATCH] Limit unichar extraction in ReadNormProtos to the buffer size

Classify::ReadNormProtos parsed each normproto line with
`stream >> unichar >> NumProtos` into a char unichar[2 * UNICHAR_LEN + 1]
stack buffer, but char* extraction from an istream has no length limit
(the stream width was never set). A crafted TESSDATA_NORMPROTO component
in a .traineddata file whose first proto-line token exceeds 60
characters (the 100-byte line buffer allows up to 99) overflows the
stack buffer during legacy engine initialization (CWE-121).

Toolchain note: Apple's libc++ provides a C++20 array overload of
operator>>(basic_istream&, char(&)[N]) that implicitly bounds the
extraction to the array size, so builds against that standard library
are incidentally protected. Standard libraries without that overload
(e.g. libstdc++) still take the unbounded char* overload, so the
explicit width limit below makes the behavior defined on all
toolchains.

Key changes:
- normmatch.cpp: read the unichar token with
  std::setw(2 * UNICHAR_LEN + 1); char* extraction takes at most
  width - 1 characters, which fits the buffer exactly. Overlong
  tokens are truncated and the line is rejected like any other
  unparseable line.
- unittest: add normproto_test covering a 99-character token (the
  maximum a 100-byte line can hold), a token of exactly 2 *
  UNICHAR_LEN characters, and a well-formed component. A minimal
  reproduction of the unbounded extraction crashes an ASan build
  with a stack-buffer-overflow.

Reported-by: Tristan Madani <[email protected]>
Assisted-by: OpenCode / qwen3.8-27b-thinking (Alibaba Cloud)
Signed-off-by: Stefan Weil <[email protected]>
---
# Backport note: test-only hunks (Makefile.am, unittest/) dropped -
# the OBS build runs no %%check. Source fix hunks are verbatim upstream.
#
 Makefile.am                |   5 ++
 src/classify/normmatch.cpp |   6 +-
 unittest/CMakeLists.txt    |   1 +
 unittest/normproto_test.cc | 111 +++++++++++++++++++++++++++++++++++++
 4 files changed, 122 insertions(+), 1 deletion(-)
 create mode 100644 unittest/normproto_test.cc

diff --git a/src/classify/normmatch.cpp b/src/classify/normmatch.cpp
index d5bd7e6a28..1ce7c9283b 100644
--- a/src/classify/normmatch.cpp
+++ b/src/classify/normmatch.cpp
@@ -28,6 +28,7 @@
 
 #include <cmath>
 #include <cstdio>
+#include <iomanip> // for std::setw
 #include <sstream> // for std::istringstream
 
 namespace tesseract {
@@ -190,7 +191,10 @@ NORM_PROTOS *Classify::ReadNormProtos(TFile *fp) {
   while (fp->FGets(line, kMaxLineSize) != nullptr) {
     std::istringstream stream(line);
     stream.imbue(std::locale::classic());
-    stream >> unichar >> NumProtos;
+    // unichar holds at most 2 * UNICHAR_LEN characters; the width limit
+    // (width - 1 characters for char* extraction) keeps the extraction
+    // from overflowing the buffer on overlong lines.
+    stream >> std::setw(2 * UNICHAR_LEN + 1) >> unichar >> NumProtos;
     if (stream.fail()) {
       continue;
     }

++++++ tesseract-CVE-2026-88048.patch ++++++
>From 103dc134eb36411ddc6833ec20aa2c76795bd0ff Mon Sep 17 00:00:00 2001
From: Stefan Weil <[email protected]>
Date: Fri, 21 Aug 2026 15:23:59 +0200
Subject: [PATCH] Validate FullyConnected weight matrix dimensions at load

FullyConnected::DeSerialize read the WeightMatrix without checking
that its dimensions match the layer's declared ni/no. MatrixDotVector
drives the dot product from the matrix dimensions (writes w.dim1()
results, reads w.dim2()-1 inputs) while the scratch buffers are sized
from no_ and ni_, so a crafted .traineddata with a mismatched matrix
performed a heap out-of-bounds write (up to 65535 rows) and out-of-
bounds read on the first recognition step (crash, or heap corruption
with attacker-influenced size and content).

Key changes:
- weightmatrix.h: add Dim1()/Dim2() accessors for the active weight
  matrix (int or float), alongside the existing NumOutputs().
- fullyconnected.cpp: reject the layer in DeSerialize unless
  Dim1() == no_ and Dim2() == ni_ + 1 (the second dimension
  includes the bias column), the layout that InitWeightsFloat
  always produces.
- unittest: new fullyconnected_test that builds a Softmax network
  with mismatched matrix dimensions and expects CreateFromFile to
  return nullptr; on unpatched code the test reaches Forward and
  ASan catches the out-of-bounds write in MatrixDotVector. A
  second test verifies that matching dimensions are still accepted.

Reported-by: Zhixi "Jace" Sun <[email protected]>
Assisted-by: OpenCode / qwen3.8-27b-thinking (Alibaba Cloud)
Signed-off-by: Stefan Weil <[email protected]>
---
# Backport note: test-only hunks (Makefile.am, unittest/) dropped -
# the OBS build runs no %%check. Source fix hunks are verbatim upstream.
#
 Makefile.am                     |   5 ++
 src/lstm/fullyconnected.cpp     |  11 ++-
 src/lstm/weightmatrix.h         |   7 ++
 unittest/fullyconnected_test.cc | 115 ++++++++++++++++++++++++++++++++
 4 files changed, 137 insertions(+), 1 deletion(-)
 create mode 100644 unittest/fullyconnected_test.cc

diff --git a/src/lstm/fullyconnected.cpp b/src/lstm/fullyconnected.cpp
index a0ee1b02dd..d999f57529 100644
--- a/src/lstm/fullyconnected.cpp
+++ b/src/lstm/fullyconnected.cpp
@@ -121,7 +121,16 @@ bool FullyConnected::Serialize(TFile *fp) const {
 
 // Reads from the given file. Returns false in case of error.
 bool FullyConnected::DeSerialize(TFile *fp) {
-  return weights_.DeSerialize(IsTraining(), fp);
+  if (!weights_.DeSerialize(IsTraining(), fp)) {
+    return false;
+  }
+  // The weight matrix must match the declared sizes (the second dimension
+  // includes the bias column); otherwise Forward would read or write
+  // outside the scratch buffers sized from ni_ and no_.
+  if (weights_.Dim1() != no_ || weights_.Dim2() != ni_ + 1) {
+    return false;
+  }
+  return true;
 }
 
 // Runs forward propagation of activations on the input line.
diff --git a/src/lstm/weightmatrix.h b/src/lstm/weightmatrix.h
index eaca3ffb09..6c5ab23659 100644
--- a/src/lstm/weightmatrix.h
+++ b/src/lstm/weightmatrix.h
@@ -107,6 +107,13 @@ class WeightMatrix {
   int NumOutputs() const {
     return int_mode_ ? wi_.dim1() : wf_.dim1();
   }
+  // The dimensions of the active weight matrix (wi_ in int mode, else wf_).
+  int Dim1() const {
+    return int_mode_ ? wi_.dim1() : wf_.dim1();
+  }
+  int Dim2() const {
+    return int_mode_ ? wi_.dim2() : wf_.dim2();
+  }
   // Provides one set of weights. Only used by peep weight maxpool.
   const TFloat *GetWeights(int index) const {
     return wf_[index];

++++++ tesseract-CVE-2026-88049.patch ++++++
>From b494ac18925f9d9aff9ef5815475de9943ab19bf Mon Sep 17 00:00:00 2001
From: Stefan Weil <[email protected]>
Date: Fri, 21 Aug 2026 16:41:51 +0200
Subject: [PATCH] Validate LSTM gate matrix dimensions against na_/no_ at load

LSTM::DeSerialize read na_ from the untrusted TESSDATA_LSTM component
and derived ns_ from the CI gate matrix's dim1, but never checked that
the deserialized dimensions were mutually consistent. The forward pass
sizes its buffers from na_, no_ and ns_ while the gate matrices drive
their own dimensions, so a crafted .traineddata performed heap
out-of-bounds writes and reads during the first recognition step, e.g.
WriteTimeStepPart writing ns_ floats at offset ni_+nf_ into a source_
buffer sized from na_ (up to ~256 KB with an attacker-chosen gate
matrix dim1).

The bounds assertions added by 2f4d2f4 (CVE-2026-73066) covered only
NetworkIO::CopyTimeStepGeneral and Randomize, leaving
WriteTimeStepPart and AddTimeStepPart unguarded.

Key changes:
- weightmatrix.h: add Dim1()/Dim2() accessors for the active weight
  matrix (int or float), alongside the existing NumOutputs().
- lstm.cpp: reject the layer in DeSerialize unless na_ ==
  ni_ + nf_ + (is_2d_ ? 2 : 1) * ns_, every deserialized gate has
  Dim1() == ns_ and Dim2() == na_ + 1 (the layout InitWeightsFloat
  always produces), ns_ == no_ for plain NT_LSTM/NT_LSTM_SUMMARY, and
  the softmax layer's sizes match ns_/no_ for the softmax variants.
- networkio.cpp: add the same defense-in-depth bounds assertions to
  WriteTimeStepPart and AddTimeStepPart as 2f4d2f4 added to
  CopyTimeStepGeneral and Randomize.
- unittest: add lstm_layer_test with crafted NT_LSTM layers for the
  na_ mismatch, gate dim1 mismatch, and gate dim2 mismatch cases
  (each rejected at load; on unpatched code the tests reach Forward
  and ASan catches the out-of-bounds write in WriteTimeStepPart),
  plus a positive control that a consistent layer loads and runs.

Reported-by: Zhixi "Jace" Sun <[email protected]>
Assisted-by: OpenCode / qwen3.8-27b-thinking (Alibaba Cloud)
Signed-off-by: Stefan Weil <[email protected]>
---
# Backport note: test-only hunks (Makefile.am, unittest/) dropped -
# the OBS build runs no %%check. Source fix hunks are verbatim upstream.
#
 Makefile.am                 |   5 +
 src/lstm/lstm.cpp           |  20 ++++
 src/lstm/networkio.cpp      |   2 +
 unittest/lstm_layer_test.cc | 177 ++++++++++++++++++++++++++++++++++++
 4 files changed, 204 insertions(+)
 create mode 100644 unittest/lstm_layer_test.cc

diff --git a/src/lstm/lstm.cpp b/src/lstm/lstm.cpp
index 7bfd19b524..bfffa5deab 100644
--- a/src/lstm/lstm.cpp
+++ b/src/lstm/lstm.cpp
@@ -274,12 +274,32 @@ bool LSTM::DeSerialize(TFile *fp) {
       is_2d_ = na_ - nf_ == ni_ + 2 * ns_;
     }
   }
+  // The deserialized dimensions must be mutually consistent: the forward
+  // pass sizes its buffers from na_, no_ and ns_ while the gate matrices
+  // drive their own dimensions.
+  if (na_ != ni_ + nf_ + (is_2d_ ? 2 : 1) * ns_) {
+    return false;
+  }
+  for (int w = 0; w < WT_COUNT; ++w) {
+    if (w == GFS && !Is2D()) {
+      continue;
+    }
+    if (gate_weights_[w].Dim1() != ns_ || gate_weights_[w].Dim2() != na_ + 1) {
+      return false;
+    }
+  }
+  if ((type_ == NT_LSTM || type_ == NT_LSTM_SUMMARY) && ns_ != no_) {
+    return false;
+  }
   delete softmax_;
   if (type_ == NT_LSTM_SOFTMAX || type_ == NT_LSTM_SOFTMAX_ENCODED) {
     softmax_ = static_cast<FullyConnected *>(Network::CreateFromFile(fp));
     if (softmax_ == nullptr) {
       return false;
     }
+    if (softmax_->NumInputs() != ns_ || softmax_->NumOutputs() != no_) {
+      return false;
+    }
   } else {
     softmax_ = nullptr;
   }
diff --git a/src/lstm/networkio.cpp b/src/lstm/networkio.cpp
index 8e1c6679c7..7a31034bf9 100644
--- a/src/lstm/networkio.cpp
+++ b/src/lstm/networkio.cpp
@@ -641,6 +641,7 @@ void NetworkIO::AddTimeStep(int t, TFloat *inout) const {
 
 // Adds part of a single timestep to floats.
 void NetworkIO::AddTimeStepPart(int t, int offset, int num_features, float 
*inout) const {
+  ASSERT_HOST(offset + num_features <= NumFeatures());
   if (int_mode_) {
     const int8_t *line = i_[t] + offset;
     for (int i = 0; i < num_features; ++i) {
@@ -662,6 +663,7 @@ void NetworkIO::WriteTimeStep(int t, const TFloat *input) {
 // Writes a single timestep from floats in the range [-1, 1] writing only
 // num_features elements of input to (*this)[t], starting at offset.
 void NetworkIO::WriteTimeStepPart(int t, int offset, int num_features, const 
TFloat *input) {
+  ASSERT_HOST(offset + num_features <= NumFeatures());
   if (int_mode_) {
     int8_t *line = i_[t] + offset;
     for (int i = 0; i < num_features; ++i) {

++++++ tesseract-CVE-2026-88050.patch ++++++
>From c94a5532ee04db5a4919542832fd94caee5ea58f Mon Sep 17 00:00:00 2001
From: Stefan Weil <[email protected]>
Date: Fri, 21 Aug 2026 16:56:51 +0200
Subject: [PATCH] Reject recoder code values outside the sane range at load

RecodedCharID::DeSerialize (hardened by 82727cc to check length_
only) still read the individual code values as raw signed int32.
UnicharCompress::ComputeCodeRange computes code_range_ as 1 plus the
maximum code using a signed > comparison, so a code value of -1
never raises the maximum and yields code_range_ = 0. SetupDecoder
then resizes is_valid_start_ to 0 and writes is_valid_start_[code(0)]
on the size-0 vector<bool>, an out-of-bounds write at a wild wrapped
index (deterministic crash) on LSTMRecognizer load. A code value of
INT32_MAX instead wraps code_range_ negative and makes resize()
throw.

Key changes:
- unicharcompress.h: validate each deserialized code value to be
  within [0, UINT16_MAX), the same arbitrary cap used elsewhere for
  .traineddata counts; reject the recoder otherwise.
- unicharcompress.h/.cpp: add defense-in-depth bounds assertions to
  IsValidFirstCode and SetupDecoder, matching the style of the
  NetworkIO assertions from 2f4d2f4.
- unittest: add recoder_test, which feeds a crafted UnicharCompress
  with a -1 code and an INT32_MAX code and expects DeSerialize to
  fail (on unpatched code the -1 case dies on the SEGV in
  SetupDecoder, the huge case on the uncaught length_error), plus a
  positive control that valid codes load and answer
  code_range()/IsValidFirstCode() correctly.

Reported-by: Zhixi "Jace" Sun <[email protected]>
Assisted-by: OpenCode / qwen3.8-27b-thinking (Alibaba Cloud)
Signed-off-by: Stefan Weil <[email protected]>
---
# Backport note: test-only hunks (Makefile.am, unittest/) dropped -
# the OBS build runs no %%check. Source fix hunks are verbatim upstream.
#
 Makefile.am                    |  5 ++
 src/ccutil/unicharcompress.cpp |  1 +
 src/ccutil/unicharcompress.h   | 13 ++++-
 unittest/recoder_test.cc       | 91 ++++++++++++++++++++++++++++++++++
 4 files changed, 109 insertions(+), 1 deletion(-)
 create mode 100644 unittest/recoder_test.cc

diff --git a/src/ccutil/unicharcompress.cpp b/src/ccutil/unicharcompress.cpp
index f909f76eeb..b3dd440f26 100644
--- a/src/ccutil/unicharcompress.cpp
+++ b/src/ccutil/unicharcompress.cpp
@@ -400,6 +400,7 @@ void UnicharCompress::SetupDecoder() {
   for (unsigned c = 0; c < encoder_.size(); ++c) {
     const RecodedCharID &code = encoder_[c];
     decoder_[code] = c;
+    ASSERT_HOST(code(0) >= 0 && code(0) < code_range_);
     is_valid_start_[code(0)] = true;
     RecodedCharID prefix = code;
     uint32_t len = code.length();
diff --git a/src/ccutil/unicharcompress.h b/src/ccutil/unicharcompress.h
index 9b696d033c..6919b7d259 100644
--- a/src/ccutil/unicharcompress.h
+++ b/src/ccutil/unicharcompress.h
@@ -84,7 +84,17 @@ class RecodedCharID {
     if (length_ > kMaxCodeLen) {
       return false;
     }
-    return fp->DeSerialize(&code_[0], length_);
+    if (!fp->DeSerialize(&code_[0], length_)) {
+      return false;
+    }
+    // Code values index arrays sized from the maximum code; reject values
+    // that are out of the sane range for a recoded alphabet.
+    for (uint32_t i = 0; i < length_; ++i) {
+      if (code_[i] < 0 || code_[i] >= static_cast<int32_t>(UINT16_MAX)) {
+        return false;
+      }
+    }
+    return true;
   }
   bool operator==(const RecodedCharID &other) const {
     if (length_ != other.length_) {
@@ -190,6 +200,7 @@ class TESS_API UnicharCompress {
   int DecodeUnichar(const RecodedCharID &code) const;
   // Returns true if the given code is a valid start or single code.
   bool IsValidFirstCode(int code) const {
+    ASSERT_HOST(code >= 0 && code < code_range_);
     return is_valid_start_[code];
   }
   // Returns a list of valid non-final next codes for a given prefix code,

++++++ tesseract-CVE-2026-88051.patch ++++++
>From 56e09ca12e751623fe796ce1554ce704bffd2ef0 Mon Sep 17 00:00:00 2001
From: Stefan Weil <[email protected]>
Date: Fri, 21 Aug 2026 20:29:37 +0200
Subject: [PATCH] Validate vector counts in GenericVector::read

The callback form of GenericVector::read read two independent int32
fields from the file: reserved sized the allocation via reserve(),
while size_used_ drove the element loop. Neither was capped and no
size_used_ <= reserved invariant was checked, so a crafted
.traineddata (e.g. the fontinfo table of a version >= 4 inttemp
component) performed a heap out-of-bounds write during legacy
engine initialization.

Key changes:
- genericvector.h: reject negative or over-limit reserved
  (matching the 50000000 cap of the DeSerialize overloads) and
  reject size_used_ < 0 or size_used_ > reserved before entering
  the read loop. Legit files always satisfy size_used_ <= reserved,
  as write() persists size_reserved_ first.
- unittest: add genericvector_test covering size_used_ beyond
  reserved (on unpatched code the ASan build dies on a
  heap-buffer-overflow in the read loop), negative counts, and a
  consistent vector that must still load.

Reported-by: Zhixi "Jace" Sun <[email protected]>
Assisted-by: OpenCode / qwen3.8-27b-thinking (Alibaba Cloud)
Signed-off-by: Stefan Weil <[email protected]>
---
# Backport note: test-only hunks (Makefile.am, unittest/) dropped -
# the OBS build runs no %%check. Source fix hunks are verbatim upstream.
#
 Makefile.am                    |  5 ++
 src/ccutil/genericvector.h     | 10 ++++
 unittest/genericvector_test.cc | 91 ++++++++++++++++++++++++++++++++++
 3 files changed, 106 insertions(+)
 create mode 100644 unittest/genericvector_test.cc

diff --git a/src/ccutil/genericvector.h b/src/ccutil/genericvector.h
index 4a5bbe12d6..cbe1e203d2 100644
--- a/src/ccutil/genericvector.h
+++ b/src/ccutil/genericvector.h
@@ -654,10 +654,20 @@ bool GenericVector<T>::read(TFile *f, const 
std::function<bool(TFile *, T *)> &c
   if (f->FReadEndian(&reserved, sizeof(reserved), 1) != 1) {
     return false;
   }
+  // Arbitrarily limit the number of elements to protect against bad data.
+  const uint32_t limit = 50000000;
+  if (reserved < 0 || static_cast<uint32_t>(reserved) > limit) {
+    return false;
+  }
   reserve(reserved);
   if (f->FReadEndian(&size_used_, sizeof(size_used_), 1) != 1) {
     return false;
   }
+  // size_used_ is an independent file field; without this check the reads
+  // below land past the end of the buffer sized from reserved.
+  if (size_used_ < 0 || size_used_ > reserved) {
+    return false;
+  }
   if (cb != nullptr) {
     for (int i = 0; i < size_used_; ++i) {
       if (!cb(f, data_ + i)) {

++++++ tesseract-CVE-2026-88052.patch ++++++
>From 2d04d640db2e8c7e3bab2369d599343b5a8b8443 Mon Sep 17 00:00:00 2001
From: Stefan Weil <[email protected]>
Date: Fri, 21 Aug 2026 18:34:33 +0200
Subject: [PATCH] Reject unicharset files whose inserts desync id from unichars

UNICHARSET::load_via_fgets reads the unichar count via sscanf and
trusts it as the loop bound, indexing the unichars vector with the
loop index id via the unchecked set_* accessors. unichar_insert is a
no-op for duplicate (or empty) representations, so once any insert
no-ops, unichars.size() falls behind id and the subsequent
set_*(id, ...) and unichars[id].properties writes land past the end
of the vector - a deterministic heap out-of-bounds write (including a
std::string assignment via set_normed) for every remaining line, on
both the LSTM and legacy init paths. A malformed unicharset with a
duplicate line (e.g. two identical entries) triggers it; a
non-positive header count likewise loads an empty unicharset
"successfully".

Key changes:
- unicharset.cpp: reject unicharset_size <= 0, and after each insert
  verify the vector actually grew to id + 1; on mismatch report the
  offending line and reject the file instead of writing out of bounds.
- unittest: add unicharset_load_test with a duplicate-representation
  unicharset (on unpatched code the test dies on the
  container-overflow in load_via_fgets), a zero and a negative count,
  and a positive control that a valid unicharset still loads.

Reported-by: Zhixi "Jace" Sun <[email protected]>
Assisted-by: OpenCode / qwen3.8-27b-thinking (Alibaba Cloud)
Signed-off-by: Stefan Weil <[email protected]>
---
# Backport note: test-only hunks (Makefile.am, unittest/) dropped -
# the OBS build runs no %%check. Source fix hunks are verbatim upstream.
#
 Makefile.am                      |  5 +++
 src/ccutil/unicharset.cpp        | 12 ++++++
 unittest/unicharset_load_test.cc | 64 ++++++++++++++++++++++++++++++++
 3 files changed, 81 insertions(+)
 create mode 100644 unittest/unicharset_load_test.cc

diff --git a/src/ccutil/unicharset.cpp b/src/ccutil/unicharset.cpp
index b29ec3b7fe..0e72ae48d0 100644
--- a/src/ccutil/unicharset.cpp
+++ b/src/ccutil/unicharset.cpp
@@ -791,6 +791,9 @@ bool UNICHARSET::load_via_fgets(
       sscanf(buffer, "%d", &unicharset_size) != 1) {
     return false;
   }
+  if (unicharset_size <= 0) {
+    return false;
+  }
   for (UNICHAR_ID id = 0; id < unicharset_size; ++id) {
     char unichar[256];
     unsigned int properties;
@@ -884,6 +887,15 @@ bool UNICHARSET::load_via_fgets(
     } else {
       this->unichar_insert_backwards_compatible(unichar);
     }
+    // A duplicate or empty representation makes the insert a no-op,
+    // desynchronizing id from the unichars vector; the set_* calls and
+    // unichars[id] below would then write out of bounds. The file is
+    // malformed, so reject it.
+    if (size() != static_cast<size_t>(id) + 1) {
+      fprintf(stderr, "%s:%d unichar %d has a duplicate or empty 
representation\n",
+              __FILE__, __LINE__, id);
+      return false;
+    }
 
     this->set_isalpha(id, properties & ISALPHA_MASK);
     this->set_islower(id, properties & ISLOWER_MASK);

++++++ tesseract-CVE-2026-88053.patch ++++++
>From 8b0574680f3b22f246ade6a4c8e3029104255c63 Mon Sep 17 00:00:00 2001
From: Stefan Weil <[email protected]>
Date: Fri, 21 Aug 2026 13:47:58 +0200
Subject: [PATCH] Fix out-of-bounds writes in .traineddata inttemp
 deserialization

Classify::ReadIntTemplates reads NumClassPruners, NumClasses and
NumProtoSets from the untrusted inttemp component of a .traineddata
file and used them as loop bounds writing pointers into fixed-size
arrays (ClassPruners[], Class[], ProtoSets[]) without validation. A
crafted file with counts larger than the array capacity caused heap
out-of-bounds pointer writes during legacy engine initialization,
before any OCR is performed.

Key changes:
- intproto.cpp: validate unicharset_size, NumClassPruners, NumClasses,
  per-class NumProtos/NumProtoSets/NumConfigs and the old-format
  (version < 2) class ids against the array capacities and fail the
  load instead of writing out of bounds. Never trust file-sourced
  counts to bound the destructor loops.
- adaptive.cpp/h: same treatment in ReadAdaptedTemplates and
  ReadAdaptedClass (NumTempProtos, NumConfigs); the
  ADAPT_TEMPLATES_STRUCT default constructor now initializes all
  members and the destructor is null-safe.
- adaptmatch.cpp, tface.cpp, tessedit.cpp: propagate the read failure
  through InitAdaptiveClassifier and program_editup so the language
  load fails gracefully instead of continuing with a broken state.
- intproto.h, adaptive.h: replace the fixed-size member arrays by
  std::array (identical on-disk layout, value-initialized).
- unittest: add intproto_test, which builds a minimal traineddata
  with corrupt inttemp counts and expects a graceful init failure.
  On unpatched code the test trips ASan on the original
  heap-buffer-overflow in ReadIntTemplates.

Reported-by: Zhixi "Jace" Sun <[email protected]>
Assisted-by: OpenCode / qwen3.8-27b-thinking (Alibaba Cloud)
Signed-off-by: Stefan Weil <[email protected]>
---
# Backport note: test-only hunks (Makefile.am, unittest/) dropped -
# the OBS build runs no %%check. Upstream's std::array modernization
# (intproto.h, adaptive.h) omitted: 5.5.3 still passes these members
# as bare pointers elsewhere (eg adaptmatch.cpp MasterMatcher call,
# intproto.cpp AddIntProto loop). intproto.cpp ctor/dtor hunk replaced
# with a minimal destructor loop-bound fix for the 5.5.3 code.
#
 Makefile.am                 |   7 ++
 src/ccmain/tessedit.cpp     |   4 +-
 src/classify/adaptive.cpp   |  91 +++++++++++++++++---------
 src/classify/adaptive.h     |   7 +-
 src/classify/adaptmatch.cpp |  27 +++++---
 src/classify/classify.h     |   2 +-
 src/classify/intproto.cpp   |  77 +++++++++++++++-------
 src/classify/intproto.h     |  12 ++--
 src/wordrec/tface.cpp       |   7 +-
 src/wordrec/wordrec.h       |   4 +-
 unittest/CMakeLists.txt     |   1 +
 unittest/intproto_test.cc   | 124 ++++++++++++++++++++++++++++++++++++
 12 files changed, 288 insertions(+), 75 deletions(-)
 create mode 100644 unittest/intproto_test.cc

diff --git a/src/ccmain/tessedit.cpp b/src/ccmain/tessedit.cpp
index 412b8c15be..cfbeca6211 100644
--- a/src/ccmain/tessedit.cpp
+++ b/src/ccmain/tessedit.cpp
@@ -418,7 +418,9 @@ int Tesseract::init_tesseract_internal(const std::string 
&textbase,
   // If only LSTM will be used, skip loading Tesseract classifier's
   // pre-trained templates and dictionary.
   bool init_tesseract = tessedit_ocr_engine_mode != OEM_LSTM_ONLY;
-  program_editup(textbase, init_tesseract ? mgr : nullptr, init_tesseract ? 
mgr : nullptr);
+  if (!program_editup(textbase, init_tesseract ? mgr : nullptr, init_tesseract 
? mgr : nullptr)) {
+    return -1;
+  }
   return 0; // Normal exit
 }
 
diff --git a/src/classify/adaptive.cpp b/src/classify/adaptive.cpp
index e3297c0e0c..96850ab75b 100644
--- a/src/classify/adaptive.cpp
+++ b/src/classify/adaptive.cpp
@@ -67,10 +67,6 @@ ADAPT_CLASS_STRUCT::ADAPT_CLASS_STRUCT() :
   TempProtos(NIL_LIST) {
   zero_all_bits(PermProtos, WordsInVectorOfSize(MAX_NUM_PROTOS));
   zero_all_bits(PermConfigs, WordsInVectorOfSize(MAX_NUM_CONFIGS));
-
-  for (int i = 0; i < MAX_NUM_CONFIGS; i++) {
-    TempConfigFor(this, i) = nullptr;
-  }
 }
 
 ADAPT_CLASS_STRUCT::~ADAPT_CLASS_STRUCT() {
@@ -92,25 +88,24 @@ ADAPT_CLASS_STRUCT::~ADAPT_CLASS_STRUCT() {
 
 /// Constructor for adapted templates.
 /// Add an empty class for each char in unicharset to the newly created 
templates.
-ADAPT_TEMPLATES_STRUCT::ADAPT_TEMPLATES_STRUCT(UNICHARSET &unicharset) {
-  Templates = new INT_TEMPLATES_STRUCT;
-  NumPermClasses = 0;
-  NumNonEmptyClasses = 0;
-
-  /* Insert an empty class for each unichar id in unicharset */
-  for (unsigned i = 0; i < MAX_NUM_CLASSES; i++) {
-    Class[i] = nullptr;
-    if (i < unicharset.size()) {
-      AddAdaptedClass(this, new ADAPT_CLASS_STRUCT, i);
-    }
+ADAPT_TEMPLATES_STRUCT::ADAPT_TEMPLATES_STRUCT(UNICHARSET &unicharset) :
+  Templates(new INT_TEMPLATES_STRUCT), NumNonEmptyClasses(0), 
NumPermClasses(0) {
+  // Insert an empty class for each unichar id in unicharset.
+  // Class is value-initialized to nullptr in-class.
+  for (unsigned i = 0; i < unicharset.size(); i++) {
+    AddAdaptedClass(this, new ADAPT_CLASS_STRUCT, i);
   }
 }
 
 ADAPT_TEMPLATES_STRUCT::~ADAPT_TEMPLATES_STRUCT() {
-  for (unsigned i = 0; i < (Templates)->NumClasses; i++) {
-    delete Class[i];
+  if (Templates != nullptr) {
+    // NumClasses comes from an untrusted file, so never trust it to bound
+    // the loop over the fixed-size Class[] array.
+    for (unsigned i = 0; i < (Templates)->NumClasses && i < MAX_NUM_CLASSES; 
i++) {
+      delete Class[i];
+    }
+    delete Templates;
   }
-  delete Templates;
 }
 
 // Returns FontinfoId of the given config of the given adapted class.
@@ -180,23 +175,32 @@ void Classify::PrintAdaptedTemplates(FILE *File, 
ADAPT_TEMPLATES_STRUCT *Templat
  * @note Globals: none
  */
 ADAPT_CLASS_STRUCT *ReadAdaptedClass(TFile *fp) {
-  int NumTempProtos;
-  int NumConfigs;
+  int NumTempProtos = 0;
+  int NumConfigs = 0;
   int i;
   ADAPT_CLASS_STRUCT *Class;
 
-  /* first read high level adapted class structure */
+  // first read high level adapted class structure
   Class = new ADAPT_CLASS_STRUCT;
   fp->FRead(Class, sizeof(ADAPT_CLASS_STRUCT), 1);
 
-  /* then read in the definitions of the permanent protos and configs */
+  // then read in the definitions of the permanent protos and configs
   Class->PermProtos = NewBitVector(MAX_NUM_PROTOS);
   Class->PermConfigs = NewBitVector(MAX_NUM_CONFIGS);
   fp->FRead(Class->PermProtos, sizeof(uint32_t), 
WordsInVectorOfSize(MAX_NUM_PROTOS));
   fp->FRead(Class->PermConfigs, sizeof(uint32_t), 
WordsInVectorOfSize(MAX_NUM_CONFIGS));
 
-  /* then read in the list of temporary protos */
+  // then read in the list of temporary protos
   fp->FRead(&NumTempProtos, sizeof(int), 1);
+  if (NumTempProtos < 0 || NumTempProtos > MAX_NUM_PROTOS) {
+    tprintf("Bad read of adapted class!\n");
+    // Reset file-sourced pointers so the destructor does not delete them.
+    for (i = 0; i < MAX_NUM_CONFIGS; i++) {
+      Class->Config[i].Temp = nullptr;
+    }
+    delete Class;
+    return nullptr;
+  }
   Class->TempProtos = NIL_LIST;
   for (i = 0; i < NumTempProtos; i++) {
     auto TempProto = new TEMP_PROTO_STRUCT;
@@ -204,8 +208,20 @@ ADAPT_CLASS_STRUCT *ReadAdaptedClass(TFile *fp) {
     Class->TempProtos = push_last(Class->TempProtos, TempProto);
   }
 
-  /* then read in the adapted configs */
+  // then read in the adapted configs
   fp->FRead(&NumConfigs, sizeof(int), 1);
+  // NumConfigs is used as a loop bound that writes into the fixed-size
+  // Config[] array, so reject a corrupt or malicious file instead of
+  // writing out of bounds.
+  if (NumConfigs < 0 || NumConfigs > MAX_NUM_CONFIGS) {
+    tprintf("Bad read of adapted class!\n");
+    // Reset file-sourced pointers so the destructor does not delete them.
+    for (i = 0; i < MAX_NUM_CONFIGS; i++) {
+      Class->Config[i].Temp = nullptr;
+    }
+    delete Class;
+    return nullptr;
+  }
   for (i = 0; i < NumConfigs; i++) {
     if (test_bit(Class->PermConfigs, i)) {
       Class->Config[i].Perm = ReadPermConfig(fp);
@@ -231,18 +247,35 @@ ADAPT_CLASS_STRUCT *ReadAdaptedClass(TFile *fp) {
 ADAPT_TEMPLATES_STRUCT *Classify::ReadAdaptedTemplates(TFile *fp) {
   auto Templates = new ADAPT_TEMPLATES_STRUCT;
 
-  /* first read the high level adaptive template struct */
-  fp->FRead(Templates, sizeof(ADAPT_TEMPLATES_STRUCT), 1);
+  // first read in the high level adaptive template struct
+  if (fp->FRead(Templates, sizeof(ADAPT_TEMPLATES_STRUCT), 1) != 1) {
+    tprintf("Bad read of adapted templates!\n");
+    delete Templates;
+    return nullptr;
+  }
+  // The Class[] array was just filled with pointers read from the file;
+  // those are not valid allocations, so reset it before storing real ones.
+  for (unsigned i = 0; i < MAX_NUM_CLASSES; i++) {
+    Templates->Class[i] = nullptr;
+  }
 
-  /* then read in the basic integer templates */
+  // then read in the basic integer templates
   Templates->Templates = ReadIntTemplates(fp);
+  if (Templates->Templates == nullptr) {
+    delete Templates;
+    return nullptr;
+  }
 
-  /* then read in the adaptive info for each class */
+  // then read in the adaptive info for each class
   for (unsigned i = 0; i < (Templates->Templates)->NumClasses; i++) {
     Templates->Class[i] = ReadAdaptedClass(fp);
+    if (Templates->Class[i] == nullptr) {
+      tprintf("Bad read of adapted templates (class %u)!\n", i);
+      delete Templates;
+      return nullptr;
+    }
   }
   return (Templates);
-
 } /* ReadAdaptedTemplates */
 
 /*---------------------------------------------------------------------------*/
diff --git a/src/classify/adaptive.h b/src/classify/adaptive.h
index 652e1fbdcb..fde3ef2270 100644
--- a/src/classify/adaptive.h
+++ b/src/classify/adaptive.h
@@ -67,9 +67,9 @@ class ADAPT_TEMPLATES_STRUCT {
 public:
-  ADAPT_TEMPLATES_STRUCT() = default;
+  ADAPT_TEMPLATES_STRUCT() : Templates(nullptr), NumNonEmptyClasses(0), 
NumPermClasses(0) {}
   ADAPT_TEMPLATES_STRUCT(UNICHARSET &unicharset);
   ~ADAPT_TEMPLATES_STRUCT();
   INT_TEMPLATES_STRUCT *Templates;
   int NumNonEmptyClasses;
   uint8_t NumPermClasses;
   ADAPT_CLASS_STRUCT *Class[MAX_NUM_CLASSES];
 };
diff --git a/src/classify/adaptmatch.cpp b/src/classify/adaptmatch.cpp
index f9f38fbe67..30d30a810e 100644
--- a/src/classify/adaptmatch.cpp
+++ b/src/classify/adaptmatch.cpp
@@ -524,9 +524,9 @@ void Classify::EndAdaptiveClassifier() {
  *      classify_use_pre_adapted_templates
  *                            enables use of pre-adapted templates
  */
-void Classify::InitAdaptiveClassifier(TessdataManager *mgr) {
+bool Classify::InitAdaptiveClassifier(TessdataManager *mgr) {
   if (!CLASSIFY_ENABLE_ADAPTIVE_MATCHER_OVERRIDE) {
-    return;
+    return true;
   }
   if (AllProtosOn != nullptr) {
     EndAdaptiveClassifier(); // Don't leak with multiple inits.
@@ -538,6 +538,11 @@ void Classify::InitAdaptiveClassifier(TessdataManager 
*mgr) {
     TFile fp;
     ASSERT_HOST(mgr->GetComponent(TESSDATA_INTTEMP, &fp));
     PreTrainedTemplates = ReadIntTemplates(&fp);
+    if (PreTrainedTemplates == nullptr) {
+      tprintf("Error: invalid inttemp component in traineddata, "
+              "cannot initialize the legacy engine.\n");
+      return false;
+    }
 
     if (mgr->GetComponent(TESSDATA_SHAPE_TABLE, &fp)) {
       shape_table_ = new ShapeTable(unicharset);
@@ -580,17 +585,23 @@ void Classify::InitAdaptiveClassifier(TessdataManager 
*mgr) {
       tprintf("\nReading pre-adapted templates from %s ...\n", 
Filename.c_str());
       fflush(stdout);
       AdaptedTemplates = ReadAdaptedTemplates(&fp);
-      tprintf("\n");
-      PrintAdaptedTemplates(stdout, AdaptedTemplates);
+      if (AdaptedTemplates == nullptr) {
+        tprintf("Error: invalid pre-adapted templates in %s, ignoring.\n", 
Filename.c_str());
+        AdaptedTemplates = new ADAPT_TEMPLATES_STRUCT(unicharset);
+      } else {
+        tprintf("\n");
+        PrintAdaptedTemplates(stdout, AdaptedTemplates);
 
-      for (unsigned i = 0; i < AdaptedTemplates->Templates->NumClasses; i++) {
-        BaselineCutoffs[i] = CharNormCutoffs[i];
+        for (unsigned i = 0; i < AdaptedTemplates->Templates->NumClasses; i++) 
{
+          BaselineCutoffs[i] = CharNormCutoffs[i];
+        }
       }
     }
   } else {
     delete AdaptedTemplates;
     AdaptedTemplates = new ADAPT_TEMPLATES_STRUCT(unicharset);
   }
+  return true;
 } /* InitAdaptiveClassifier */
 
 void Classify::ResetAdaptiveClassifierInternal() {
@@ -1243,8 +1254,8 @@ UNICHAR_ID *Classify::BaselineClassifier(TBLOB *Blob,
   }
 
   MasterMatcher(Templates->Templates, int_features.size(), &int_features[0], 
CharNormArray,
-                Templates->Class, matcher_debug_flags, 0, 
Blob->bounding_box(), Results->CPResults,
-                Results);
+                Templates->Class, matcher_debug_flags, 0, Blob->bounding_box(),
+                Results->CPResults, Results);
 
   delete[] CharNormArray;
   CLASS_ID ClassId = Results->best_unichar_id;
diff --git a/src/classify/classify.h b/src/classify/classify.h
index 1a511c2826..8ebf451f31 100644
--- a/src/classify/classify.h
+++ b/src/classify/classify.h
@@ -164,7 +164,7 @@ class TESS_API Classify : public CCStruct {
   // provided to explicitly clarify the character segmentation.
   void LearnPieces(const char *fontname, int start, int length, float 
threshold,
                    CharSegmentationType segmentation, const char 
*correct_text, WERD_RES *word);
-  void InitAdaptiveClassifier(TessdataManager *mgr);
+  bool InitAdaptiveClassifier(TessdataManager *mgr);
   void InitAdaptedClass(TBLOB *Blob, CLASS_ID ClassId, int FontinfoId, 
ADAPT_CLASS_STRUCT *Class,
                         ADAPT_TEMPLATES_STRUCT *Templates);
   void AmbigClassifier(const std::vector<INT_FEATURE_STRUCT> &int_features,
diff --git a/src/classify/intproto.cpp b/src/classify/intproto.cpp
index b199ae2893..15cd260d9c 100644
--- a/src/classify/intproto.cpp
+++ b/src/classify/intproto.cpp
@@ -596,5 +596,7 @@ INT_CLASS_STRUCT::~INT_CLASS_STRUCT() {
 INT_CLASS_STRUCT::~INT_CLASS_STRUCT() {
-  for (int i = 0; i < NumProtoSets; i++) {
+  // NumProtoSets comes from an untrusted file, so never trust it to bound
+  // the loop over the fixed-size ProtoSets[] array.
+  for (int i = 0; i < NumProtoSets && i < MAX_NUM_PROTO_SETS; i++) {
     delete ProtoSets[i];
   }
 }
@@ -613,8 +613,10 @@ INT_TEMPLATES_STRUCT::~INT_TEMPLATES_STRUCT() {
 INT_TEMPLATES_STRUCT::~INT_TEMPLATES_STRUCT() {
-  for (unsigned i = 0; i < NumClasses; i++) {
+  // The counts come from an untrusted file, so never trust them to bound
+  // the loops over the fixed-size arrays.
+  for (unsigned i = 0; i < NumClasses && i < MAX_NUM_CLASSES; i++) {
     delete Class[i];
   }
-  for (unsigned i = 0; i < NumClassPruners; i++) {
+  for (unsigned i = 0; i < NumClassPruners && i < MAX_NUM_CLASS_PRUNERS; i++) {
     delete ClassPruners[i];
   }
 }
@@ -669,6 +663,20 @@ INT_TEMPLATES_STRUCT *Classify::ReadIntTemplates(TFile 
*fp) {
     Templates->NumClasses = version_id;
   }
 
+  // The counts read from the file are used as loop bounds that write into
+  // fixed-size arrays (Class[], ClassPruners[], TempClassPruner[] and
+  // IndexFor[]), so reject a corrupt or malicious file instead of writing
+  // out of bounds.
+  if (unicharset_size > MAX_NUM_CLASSES ||
+      Templates->NumClassPruners > MAX_NUM_CLASS_PRUNERS ||
+      Templates->NumClasses > MAX_NUM_CLASSES) {
+    tprintf("Error: invalid counts in inttemp: unicharset_size=%u, 
NumClassPruners=%u, "
+            "NumClasses=%u\n",
+            unicharset_size, Templates->NumClassPruners, 
Templates->NumClasses);
+    delete Templates;
+    return nullptr;
+  }
+
   if (version_id < 3) {
     MaxNumConfigs = OLD_MAX_NUM_CONFIGS;
     WerdsPerConfigVec = OLD_WERDS_PER_CONFIG_VEC;
@@ -708,6 +716,16 @@ INT_TEMPLATES_STRUCT *Classify::ReadIntTemplates(TFile 
*fp) {
         max_class_id = ClassIdFor[i];
       }
     }
+    // Class ids index Class[] and (divided by CLASSES_PER_CP) ClassPruners[],
+    // so reject a corrupt or malicious file instead of writing out of bounds.
+    if (max_class_id >= MAX_NUM_CLASSES) {
+      tprintf("Error: class id %u in inttemp exceeds MAX_NUM_CLASSES\n", 
max_class_id);
+      for (unsigned i = 0; i < Templates->NumClassPruners; i++) {
+        delete TempClassPruner[i];
+      }
+      delete Templates;
+      return nullptr;
+    }
     for (int i = 0; i <= CPrunerIdFor(max_class_id); i++) {
       Templates->ClassPruners[i] = new CLASS_PRUNER_STRUCT;
       memset(Templates->ClassPruners[i], 0, sizeof(CLASS_PRUNER_STRUCT));
@@ -777,8 +795,20 @@ INT_TEMPLATES_STRUCT *Classify::ReadIntTemplates(TFile 
*fp) {
       }
     }
     unsigned num_configs = version_id < 4 ? MaxNumConfigs : Class->NumConfigs;
-    ASSERT_HOST(num_configs <= MaxNumConfigs);
-    if (fp->FReadEndian(Class->ConfigLengths, sizeof(uint16_t), num_configs) 
!= num_configs) {
+    // Class->NumProtoSets is used as a loop bound that writes into the
+    // fixed-size ProtoSets[] array, so reject a corrupt or malicious file
+    // instead of writing out of bounds.
+    if (Class->NumProtos > MAX_NUM_PROTOS || Class->NumProtoSets > 
MAX_NUM_PROTO_SETS ||
+        num_configs > MaxNumConfigs) {
+      tprintf("Error: invalid counts for class %u in inttemp: NumProtos=%u, 
NumProtoSets=%u, "
+              "NumConfigs=%u\n",
+              i, Class->NumProtos, Class->NumProtoSets, Class->NumConfigs);
+      Class->NumProtoSets = 0; // no proto sets allocated yet; keep destructor 
safe
+      delete Class;
+      delete Templates;
+      return nullptr;
+    }
+    if (fp->FReadEndian(Class->ConfigLengths, sizeof(uint16_t), num_configs) 
!= num_configs) {
       tprintf("Bad read of inttemp!\n");
     }
     if (version_id < 2) {
@@ -812,8 +842,9 @@ INT_TEMPLATES_STRUCT *Classify::ReadIntTemplates(TFile *fp) 
{
             fp->FRead(&ProtoSet->Protos[x].Angle, 
sizeof(ProtoSet->Protos[x].Angle), 1) != 1) {
           tprintf("Bad read of inttemp!\n");
         }
-        if (fp->FReadEndian(&ProtoSet->Protos[x].Configs, 
sizeof(ProtoSet->Protos[x].Configs[0]),
-                            WerdsPerConfigVec) != WerdsPerConfigVec) {
+        if (fp->FReadEndian(ProtoSet->Protos[x].Configs,
+                            sizeof(ProtoSet->Protos[x].Configs[0]), 
WerdsPerConfigVec) !=
+            WerdsPerConfigVec) {
           tprintf("Bad read of inttemp!\n");
         }
       }
diff --git a/src/wordrec/tface.cpp b/src/wordrec/tface.cpp
index c058d4abed..146f0b01ff 100644
--- a/src/wordrec/tface.cpp
+++ b/src/wordrec/tface.cpp
@@ -36,14 +36,16 @@ namespace tesseract {
  * init_permute determines whether to initialize the permute functions
  * and Dawg models.
  */
-void Wordrec::program_editup(const std::string &textbase, TessdataManager 
*init_classifier,
+bool Wordrec::program_editup(const std::string &textbase, TessdataManager 
*init_classifier,
                              TessdataManager *init_dict) {
   if (!textbase.empty()) {
     imagefile = textbase;
   }
 #ifndef DISABLED_LEGACY_ENGINE
   InitFeatureDefs(&feature_defs_);
-  InitAdaptiveClassifier(init_classifier);
+  if (!InitAdaptiveClassifier(init_classifier)) {
+    return false;
+  }
   if (init_dict) {
     getDict().SetupForLoad(Dict::GlobalDawgCache());
     getDict().Load(lang, init_dict);
@@ -51,6 +53,7 @@ void Wordrec::program_editup(const std::string &textbase, 
TessdataManager *init_
   }
   pass2_ok_split = chop_ok_split;
 #endif // ndef DISABLED_LEGACY_ENGINE
+  return true;
 }
 
 /**
diff --git a/src/wordrec/wordrec.h b/src/wordrec/wordrec.h
index a50ad871ae..ecf67f02d4 100644
--- a/src/wordrec/wordrec.h
+++ b/src/wordrec/wordrec.h
@@ -50,7 +50,7 @@ class TESS_API Wordrec : public Classify {
   virtual ~Wordrec() = default;
 
   // tface.cpp
-  void program_editup(const std::string &textbase, TessdataManager 
*init_classifier,
+  bool program_editup(const std::string &textbase, TessdataManager 
*init_classifier,
                       TessdataManager *init_dict);
   void program_editdown();
   int end_recog();
@@ -243,7 +243,7 @@ class TESS_API Wordrec : public Classify {
   }
 
   // tface.cpp
-  void program_editup(const std::string &textbase, TessdataManager 
*init_classifier,
+  bool program_editup(const std::string &textbase, TessdataManager 
*init_classifier,
                       TessdataManager *init_dict);
   void cc_recog(WERD_RES *word);
   void program_editdown();

++++++ tesseract-CVE-2026-88054.patch ++++++
>From 552771236b0d80cbdb0c7dd856120fa21a4672e5 Mon Sep 17 00:00:00 2001
From: Stefan Weil <[email protected]>
Date: Fri, 21 Aug 2026 14:38:55 +0200
Subject: [PATCH] Reject empty network stacks in LSTM .traineddata
 deserialization

Plumbing::DeSerialize read the network stack size from the untrusted
TESSDATA_LSTM component and only rejected size > 10000. A crafted
.traineddata with a top-level Series/Parallel/Reversed layer with an
empty stack survived load; LSTMRecognizer initialization then called
network_->CacheXScaleFactor(network_->XScaleFactor()), which
dereferences stack_[0] on the empty vector:
- Series: Series::CacheXScaleFactor (series.cpp)
- Parallel/Reversed: inherited Plumbing::XScaleFactor
The virtual call through the wild pointer crashed the process at
initialization (deterministic denial of service).

Key changes:
- plumbing.cpp: reject size == 0 for all plumbing types, and
  size < 2 for NT_SERIES (Series::Forward requires two or more
  networks and always aborts on one).
- tessedit.cpp: fail the language load gracefully when the LSTM model
  cannot be loaded, instead of aborting via ASSERT_HOST, so
  TessBaseAPI::Init returns -1 on corrupt traineddata.
- unittest: add plumbing_test, which builds a minimal traineddata
  with empty/undersized LSTM plumbing stacks and expects a graceful
  init failure. On unpatched code the tests die on the original
  SEGV in Series::CacheXScaleFactor / Plumbing::XScaleFactor.

Reported-by: Zhixi "Jace" Sun <[email protected]>
Assisted-by: OpenCode / qwen3.8-27b-thinking (Alibaba Cloud)
Signed-off-by: Stefan Weil <[email protected]>
---
# Backport note: test-only hunks (Makefile.am, unittest/) dropped -
# the OBS build runs no %%check. Source fix hunks are verbatim upstream.
#
 Makefile.am               |   5 ++
 src/ccmain/tessedit.cpp   |   7 +-
 src/lstm/plumbing.cpp     |   6 ++
 unittest/plumbing_test.cc | 165 ++++++++++++++++++++++++++++++++++++++
 4 files changed, 182 insertions(+), 1 deletion(-)
 create mode 100644 unittest/plumbing_test.cc

diff --git a/src/ccmain/tessedit.cpp b/src/ccmain/tessedit.cpp
index cfbeca6211..aaa1599bfc 100644
--- a/src/ccmain/tessedit.cpp
+++ b/src/ccmain/tessedit.cpp
@@ -169,7 +169,12 @@ bool Tesseract::init_tesseract_lang_data(const std::string 
&language, OcrEngineM
 #endif // ndef DISABLED_LEGACY_ENGINE
     if (mgr->IsComponentAvailable(TESSDATA_LSTM)) {
       lstm_recognizer_ = new LSTMRecognizer(language_data_path_prefix.c_str());
-      ASSERT_HOST(lstm_recognizer_->Load(this->params(), lstm_use_matrix ? 
language : "", mgr));
+      if (!lstm_recognizer_->Load(this->params(), lstm_use_matrix ? language : 
"", mgr)) {
+        delete lstm_recognizer_;
+        lstm_recognizer_ = nullptr;
+        tprintf("Error: Failed to load the LSTM model from %s\n", 
tessdata_path.c_str());
+        return false;
+      }
     } else {
 #ifdef DISABLED_LEGACY_ENGINE
       // The legacy engine is compiled out, so we cannot fall back to it.
diff --git a/src/lstm/plumbing.cpp b/src/lstm/plumbing.cpp
index 7288bd02b5..908a751381 100644
--- a/src/lstm/plumbing.cpp
+++ b/src/lstm/plumbing.cpp
@@ -228,6 +228,12 @@ bool Plumbing::DeSerialize(TFile *fp) {
   if (size > 10000) {
     return false;
   }
+  // Reject empty stacks: XScaleFactor, CacheXScaleFactor and other methods
+  // unconditionally dereference stack_[0] during network initialization.
+  // A Series needs at least two networks (see Series::Forward).
+  if (size == 0 || (type() == NT_SERIES && size == 1)) {
+    return false;
+  }
   for (uint32_t i = 0; i < size; ++i) {
     Network *network = CreateFromFile(fp);
     if (network == nullptr) {

Reply via email to