This is an automated email from the ASF dual-hosted git repository.

tballison pushed a commit to branch main
in repository https://gitbox.apache.org/repos/asf/tika.git


The following commit(s) were added to refs/heads/main by this push:
     new 7a303f2e50 update CHANGES.txt (#3255)
7a303f2e50 is described below

commit 7a303f2e500d73f9fc03d99d3f91c82d55f11f3a
Author: Tim Allison <[email protected]>
AuthorDate: Fri Sep 25 07:47:18 2026 -0400

    update CHANGES.txt (#3255)
---
 CHANGES.txt | 303 ++++++++++++++++++++++++++++++------------------------------
 1 file changed, 154 insertions(+), 149 deletions(-)

diff --git a/CHANGES.txt b/CHANGES.txt
index 90bc319f5e..3348b3049e 100644
--- a/CHANGES.txt
+++ b/CHANGES.txt
@@ -1,7 +1,143 @@
 Release 4.1.0 - 9/25/2026
 
-   * Revert jackcess to 5.0.0: 5.0.1 failed to open some Access databases that
-     5.0.0 reads ("Did not find required parent table id") (TIKA-4880).
+  HIGHLIGHTS
+
+   * OCR and enrichment engines are selected by name in a new
+     "text-recognizers" list ("[]" turns enrichment off; no list resolves one
+     engine per media type at startup); the image and PDF parsers invoke them
+     rather than the composite dispatching to them. The image/ocr-*
+     pseudo-types are retired; a third-party engine still advertising them is
+     a legacy text recognizer with a WARN. VLM parsers gain "textRecognizer"
+     (TIKA-4872, TIKA-4884).
+
+   * New "exception-reporting" parse-context config redacts and bounds
+     exception text in metadata, tika-server error bodies and pipes/grpc
+     messages; FileSystemEmitter writes atomically. Compat: tk:exception:*
+     values may now begin with the TikaException wrapper line (TIKA-4848).
+
+  PERFORMANCE
+
+   * Detection (magic ~35% faster on unmatched input, override
+     keys honored before magic, cached type->parser map), markdown output
+     ~30x faster, .doc cleanup, CSV sniffing, zip legacy-method rewinds, pipes
+     ACK overlap, raw UTF-8 content passback for tika-server
+     (content-bytes-config). Compat: with a Content-Type override set,
+     DefaultDetector no longer lets a more specific magic result overrule it
+     (TIKA-4868).
+
+   * PDF incremental-update scanning reads in blocks; ~4x less CPU (TIKA-4898).
+
+   * Embedded documents in Office files, PDFs, PST and the zip-family
+     containers are re-opened from their container on rewind instead of
+     cached (TIKA-4878).
+
+   * Improve spooling/decrease number of spills to disk. The pipes cache
+     memory budget defaults to a quarter of the fork heap
+     (-Dtika.pipes.cacheMemoryBudgetBytes overrides) (TIKA-4835).
+
+   * Digesting embedded documents no longer spools each to a temp file; a
+     process-wide CacheMemoryBudget keeps them in memory. Zip-family parsing
+     uses seekable channels, so hasFile() may be false afterward (TIKA-4828).
+
+  BREAKING CHANGES AND DEPRECATIONS
+
+   * Pipes IPC carries inline bytes as a raw binary field, not in the tuple;
+     4.0.0 tuples with an "inline-bytes" parse-context entry are rejected
+     (TIKA-4829).
+
+   * Pipes plugins no longer bundle Jackson; plugin config is parsed by a
+     shared strict mapper that rejects unknown and duplicate keys (TIKA-4840).
+
+   * OpenAIVLMParser no longer auto-registers via SPI; per-request config for
+     the embedding filters works and locks baseUrl/apiKey/model; inline PDF
+     page OCR accumulates tk:chunks from every page (TIKA-4871).
+
+   * Temp files follow -Djava.io.tmpdir on the parent JVM (Tika, its
+     libraries and forks); pipes.tempDirectory is deprecated. A fork whose
+     parent dies deletes its own temp dir; failure-path temp leaks fixed
+     (TIKA-4877).
+
+   * Kafka pipes iterator no longer stops on the first empty poll
+     (assignmentTimeoutMs, drainIdleMs); groupInitialRebalanceDelayMs is
+     deprecated (TIKA-4833).
+
+  INFERENCE (EXPERIMENTAL)
+
+   * tika-inference and tika-vlm are still experimental: classes, config keys 
and
+     output may change in minor releases. Stable: the "engines" map, the
+     "text-recognizers" list and pdf-parser "text" (TIKA-4884, TIKA-4895).
+
+   * Refactor OCR and inference to use engines for inference and
+     page rendering (TIKA-4889).
+
+   * Enable audio/video chunking with ffmpeg and inference; the default
+     segment grid is 25 s windows with 5 s overlap (TIKA-4900).
+
+  SERVER, PIPES AND OPERATIONS
+
+   * tika-server: named configuration presets, and catalog presets that are
+     inert until the config names them (TIKA-4856).
+
+   * Add Micrometer reporting and opt-in endpoint for tika-server (TIKA-4839).
+
+   * Enable environment variable interpolation in tika-config.json (TIKA-4909).
+
+   * tika-server and tika-async-cli accept // and /* */ comments in config
+     during override merging (TIKA-4834).
+
+   * New file-system-jsonl-reporter pipes reporter records a batch run's
+     crashes; tika-eval Profile/Compare read that ledger and the run-info
+     json to classify NO_EXTRACT_FILE by cause (TIKA-4846, TIKA-4847).
+
+  EXTRACTION AND FORMATS
+
+   * Improve extraction of tagged PDFs. New tika-eval-structure tool compares
+     two sets of XHTML extracts block by block (TIKA-4891).
+
+   * OneNote: document-order extraction, superseded revisions omitted,
+     embedded BLOBs extracted, bounded recursion; malformed files fall back
+     to the legacy string dump via Henry Lindeman (TIKA-4814).
+
+   * Raw camera formats (NEF/NRW, PEF/PTX, ARW/SRF/SR2, SRW, DNG, RAF, RW2,
+     MRW, ORF) are detected by content instead of as image/tiff via Dominik
+     Schmidt (TIKA-4861).
+
+   * AVIF images are parsed by HeifParser (dimensions, EXIF, XMP) via Dominik
+     Schmidt (TIKA-4870).
+
+   * New GeoGebraParser for *.ggb/*.ggs/*.ggt with content-based detection
+     and THUMBNAIL embedded documents via Dominik Schmidt (TIKA-4831).
+
+   * The video of a Google/Android motion photo is emitted as an ATTACHMENT
+     embedded document via Dominik Schmidt (TIKA-4869).
+
+   * RawTiffParser extracts embedded JPEG previews as THUMBNAIL embedded
+     documents; image/x-raw-* are now subtypes of image/tiff
+     ("extractPreviews": false disables) via Dominik Schmidt (TIKA-4824).
+
+   * Apple Mail emlx files are detected as message/x-emlx and parsed by
+     RFC822Parser; a file name no longer turns plain text into a message type
+     once the type's magic has rejected the bytes (TIKA-4890).
+
+   * Audio cover art (ID3/FLAC/Vorbis front cover, MP4 covr) is emitted as a
+     THUMBNAIL embedded document rather than INLINE via Dominik Schmidt
+     (TIKA-4850).
+
+   * EpubParser emits the OPF cover image as a THUMBNAIL embedded document
+     via Dominik Schmidt (TIKA-4852).
+
+   * DWGReadParser emits THUMBNAILIMAGE as THUMBNAIL instead of INLINE
+     via Dominik Schmidt (TIKA-4853).
+
+   * iWork '09 and '18 preview images are emitted as THUMBNAIL embedded
+     documents via Dominik Schmidt (TIKA-4854).
+
+   * New poi-metafile-renderer rasterizes EMF/WMF (opt-in "renderImage");
+     OfficeParser emits the OLE2 SummaryInformation thumbnail as a THUMBNAIL
+     embedded document ("extractThumbnail": false disables) via Dominik
+     Schmidt (TIKA-4855).
+
+  OTHER CHANGES
 
    * Ogg, Vorbis, Opus and FLAC comment fields are read independently of the
      JVM's default locale: on a Turkish JVM TITLE and ARTIST were silently
@@ -17,20 +153,9 @@ Release 4.1.0 - 9/25/2026
      documented way to disable stall detection, no longer logs a warning
      on every parse (TIKA-4919).
 
-   * Apple Mail emlx files are detected as message/x-emlx and parsed by
-     RFC822Parser; a file name no longer turns plain text into a message type
-     once the type's magic has rejected the bytes (TIKA-4890).
-
    * Retry a 429, 502, 503 or 504 answer from a hosted inference engine with
      a jittered backoff (TIKA-4912).
 
-   * Handle inference exceptions more robustly in PDFParser (TIKA-4911).
-
-   * Enable environment variable interpolation in tika-config.json (TIKA-4909).
-
-   * Fix a cache-budget leak: tar, 7z and streamed zip entries never released
-     their in-memory cache's reservation (TIKA-4908).
-
    * Mp3Parser no longer writes "null" (a missing album) or the duration into
      the body (TIKA-4907).
 
@@ -49,37 +174,11 @@ Release 4.1.0 - 9/25/2026
      them. Curve values are formatted independently of the JVM's default
      locale, and single-entry gamma curves decode correctly (TIKA-4902).
 
-   * Digesting an embedded document no longer inflates it twice (TIKA-4901).
-
-   * Enable audio/video chunking with ffmpeg and inference; the default
-     segment grid is 25 s windows with 5 s overlap (TIKA-4900).
-
-   * Refactor OCR and inference to use engines for inference and
-     page rendering (TIKA-4889).
-
-   * PDF incremental-update scanning reads in blocks; ~4x less CPU (TIKA-4898).
-
    * Metadata never stores an unpaired UTF-16 surrogate (TIKA-4897).
 
-   * openai-embedding-engine gains "requestParameters" (TIKA-4896).
-
-   * tika-server: named configuration presets, and catalog presets that are
-     inert until the config names them (TIKA-4856).
-
    * unpack-config gains includeMetadata; tika-server /rmeta gains
      config/{handlerType} (TIKA-4881).
 
-   * Improve extraction of tagged PDFs. New tika-eval-structure tool compares
-     two sets of XHTML extracts block by block (TIKA-4891).
-
-   * Inference results in the single-object outputs (/tika as JSON, tika-app
-     -j, pipes CONCATENATE) land on the one metadata object handed back, with
-     an embedded locator, instead of being discarded (TIKA-4895).
-
-   * tika-inference and tika-vlm are experimental: classes, config keys and
-     output may change in minor releases. Stable: the "engines" map, the
-     "text-recognizers" list and pdf-parser "text" (TIKA-4884, TIKA-4895).
-
    * PDF bookmark outlines are walked without recursion; lists nest at most
      50 deep (TIKA-4894).
 
@@ -90,15 +189,9 @@ Release 4.1.0 - 9/25/2026
    * A PDF page rendered as an embedded document is enriched once, by the
      PDF parser; NO_OCR now means no page OCR at all (TIKA-4887).
 
-   * OCR engines are named in "text-recognizers" ("[]" turns enrichment off;
-     no list resolves one engine per media type at startup). The image/ocr-*
-     pseudo-types are retired; a third-party engine still advertising them is
-     a legacy text recognizer with a WARN until 5.0. VLM parsers gain
-     "textRecognizer" (TIKA-4884).
-
    * PDF AUTO OCR no longer emits a page's extracted text next to the OCR
      output that superseded it; AUTO without a text recognizer behaves as
-     NO_OCR. New TextRecognizer capability for content enrichers (TIKA-4883).
+     NO_OCR. (TIKA-4883).
 
    * tika-app -m and --json no longer drop metadata keys added after
      endDocument (TIKA-4885).
@@ -106,62 +199,16 @@ Release 4.1.0 - 9/25/2026
    * A missing OOXML relationship target (xlsx threaded comments, vsdx pages)
      no longer aborts the whole file (TIKA-4879).
 
-   * Raw camera formats (NEF/NRW, PEF/PTX, ARW/SRF/SR2, SRW, DNG, RAF, RW2,
-     MRW, ORF) are detected by content instead of as image/tiff. Fix a
-     BigTIFF offset overflow in RawTiffDetector that failed detection for the
-     whole document (TIKA-4861).
-
-   * Embedded documents in Office files, PDFs, PST and the zip-family
-     containers are re-opened from their container on rewind instead of
-     cached (TIKA-4878).
-
-   * Temp files follow -Djava.io.tmpdir on the parent JVM (Tika, its
-     libraries and forks); pipes.tempDirectory is deprecated. A fork whose
-     parent dies deletes its own temp dir; failure-path temp leaks fixed
-     (TIKA-4877).
-
    * tika-eval Profile/Compare speedups: langdetect preprocessing, H2 page
      cache sizing, a per-interval rate in the status log (TIKA-4875).
 
-   * New "text-recognizers" config list selects OCR and enrichment engines by
-     name; enrichers are invoked by the image and PDF parsers rather than
-     dispatched to by the composite (TIKA-4872).
-
-   * OpenAIVLMParser no longer auto-registers via SPI; per-request config for
-     the embedding filters works and locks baseUrl/apiKey/model; inline PDF
-     page OCR accumulates tk:chunks from every page (TIKA-4871).
-
-   * Performance: detection (magic ~35% faster on unmatched input, override
-     keys honored before magic, cached type->parser map), markdown output
-     ~30x faster, .doc cleanup, CSV sniffing, zip legacy-method rewinds, pipes
-     ACK overlap, raw UTF-8 content passback for tika-server
-     (content-bytes-config). Compat: with a Content-Type override set,
-     DefaultDetector no longer lets a more specific magic result overrule it
-     (TIKA-4868).
-
    * Embedded documents carry Content-Length where the stream knows it
-     (TIKA-4873).
-
-   * AVIF images are parsed by HeifParser (dimensions, EXIF, XMP) (TIKA-4870).
-
-   * The video of a Google/Android motion photo is emitted as an ATTACHMENT
-     embedded document (TIKA-4869).
-
-   * New poi-metafile-renderer rasterizes EMF/WMF (opt-in "renderImage");
-     OfficeParser emits the OLE2 SummaryInformation thumbnail as a THUMBNAIL
-     embedded document ("extractThumbnail": false disables) (TIKA-4855).
-
-   * New "exception-reporting" parse-context config redacts and bounds
-     exception text in metadata, tika-server error bodies and pipes/grpc
-     messages; FileSystemEmitter writes atomically. Compat: tk:exception:*
-     values may now begin with the TikaException wrapper line (TIKA-4848).
-
-   * Audio cover art (ID3/FLAC/Vorbis front cover, MP4 covr) is emitted as a
-     THUMBNAIL embedded document rather than INLINE (TIKA-4850).
+     via Dominik Schmidt (TIKA-4873).
 
    * The default plugins directory (tika-server, PipesForkParser, async CLI)
      and tika-grpc's plugin-roots fallback are resolved against the install
-     layout instead of the working directory (TIKA-4864, TIKA-4865).
+     layout instead of the working directory via Dominik Schmidt (TIKA-4864,
+     TIKA-4865).
 
    * Allow image compression settings in PDFBox-based renderer (TIKA-4862).
 
@@ -170,38 +217,10 @@ Release 4.1.0 - 9/25/2026
      CPU under forked parse workers (TIKA-4910, TIKA-4866, TIKA-4863).
 
    * embedded-limits maxDepth counts embedding levels again; values above 1
-     used to stop one level early (TIKA-4857).
-
-   * New GeoGebraParser for *.ggb/*.ggs/*.ggt with content-based detection
-     and THUMBNAIL embedded documents (TIKA-4831).
+     used to stop one level early via Dominik Schmidt (TIKA-4857).
 
    * Enum values in JSON configuration are matched case-insensitively
-     (TIKA-4859).
-
-   * iWork '09 and '18 preview images are emitted as THUMBNAIL embedded
-     documents (TIKA-4854).
-
-   * DWGReadParser emits THUMBNAILIMAGE as THUMBNAIL instead of INLINE
-     (TIKA-4853).
-
-   * EpubParser emits the OPF cover image as a THUMBNAIL embedded document
-     (TIKA-4852).
-
-   * RawTiffParser marks only the largest JPEG preview as THUMBNAIL; smaller
-     previews are INLINE (TIKA-4851).
-
-   * New file-system-jsonl-reporter pipes reporter records a batch run's
-     crashes; tika-eval Profile/Compare read that ledger and the run-info
-     json to classify NO_EXTRACT_FILE by cause (TIKA-4846, TIKA-4847).
-
-   * Stop spooling OLE2 objects whose header over-reserves BAT capacity
-     (TIKA-4845).
-
-   * Add Micrometer reporting and opt-in endpoint for tika-server (TIKA-4839).
-
-   * Improve spooling/decrease number of spills to disk. The pipes cache
-     memory budget defaults to a quarter of the fork heap
-     (-Dtika.pipes.cacheMemoryBudgetBytes overrides) (TIKA-4835).
+     via Dominik Schmidt (TIKA-4859).
 
    * Per-request config for parsers that lock fields (Tess4J, VLM, OpenAI
      image-embedding) threw even when empty (TIKA-4843).
@@ -216,34 +235,20 @@ Release 4.1.0 - 9/25/2026
      extractFontNames NullPointerException on a PDF page with no /Resources
      (TIKA-4842).
 
-   * Pipes plugins no longer bundle Jackson; plugin config is parsed by a
-     shared strict mapper that rejects unknown and duplicate keys (TIKA-4840).
-
-   * tika-server and tika-async-cli accept // and /* */ comments in config
-     during override merging (TIKA-4834).
-
-   * Kafka pipes iterator no longer stops on the first empty poll
-     (assignmentTimeoutMs, drainIdleMs); groupInitialRebalanceDelayMs is
-     deprecated (TIKA-4833).
+   * Pipes carries the client Content-Type into the forked worker as a
+     detection hint for every forked endpoint via Dominik Schmidt (TIKA-4825).
 
-   * Pipes IPC carries inline bytes as a raw binary field, not in the tuple;
-     4.0.0 tuples with an "inline-bytes" parse-context entry are rejected
-     (TIKA-4829).
+   * MP4 audio and video track codecs are exposed as FourCCs (audio:fourcc,
+     video:fourcc) via Dominik Schmidt (TIKA-4838).
 
-   * Digesting embedded documents no longer spools each to a temp file; a
-     process-wide CacheMemoryBudget keeps them in memory. Zip-family parsing
-     uses seekable channels, so hasFile() may be false afterward (TIKA-4828).
+   * FilenameUtils returns the default value for an unknown extension via
+     Tim Grein (TIKA-4882).
 
-   * Pipes carries the client Content-Type into the forked worker as a
-     detection hint for every forked endpoint (TIKA-4825).
-
-   * OneNote: document-order extraction, superseded revisions omitted,
-     embedded BLOBs extracted, bounded recursion; malformed files fall back
-     to the legacy string dump (TIKA-4814).
+   * ProcessUtils does not start a subprocess when the granted timeout is
+     <= 0 via Tim Grein (TIKA-4886).
 
-   * RawTiffParser extracts embedded JPEG previews as THUMBNAIL embedded
-     documents; image/x-raw-* are now subtypes of image/tiff
-     ("extractPreviews": false disables) (TIKA-4824).
+   * Remove the unnecessary jaxb-runtime dependency from the pdf and
+     miscoffice modules via Thorsten Heit (TIKA-4849).
 
 Release 4.0.0 - 8/18/2026
 

Reply via email to