This is an automated email from the ASF dual-hosted git repository.

tballison pushed a commit to branch main
in repository https://gitbox.apache.org/repos/asf/tika.git


The following commit(s) were added to refs/heads/main by this push:
     new fd4837da79 prep for 4.1.0 release (#3220)
fd4837da79 is described below

commit fd4837da79c70004ba12435f2446d4b9c92da266
Author: Tim Allison <[email protected]>
AuthorDate: Mon Sep 21 09:34:15 2026 -0400

    prep for 4.1.0 release (#3220)
---
 .skills/devs/tika-eval-compare/SKILL.md |  21 +
 CHANGES.txt                             | 879 +++++++-------------------------
 2 files changed, 204 insertions(+), 696 deletions(-)

diff --git a/.skills/devs/tika-eval-compare/SKILL.md 
b/.skills/devs/tika-eval-compare/SKILL.md
index 95c69fc479..96fa6c6972 100644
--- a/.skills/devs/tika-eval-compare/SKILL.md
+++ b/.skills/devs/tika-eval-compare/SKILL.md
@@ -108,6 +108,27 @@ ships the jsonl reporter (TIKA-4846), 
`crashes-<run.id>.jsonl` into
 (`-ra/-rb`, `-pa/-pb` override). A baseline without the reporter gets
 run-info but no ledger. Needs python3.
 
+Launching tika-app directly (a standing run script with its own config, as on a
+corpus box) skips `run-batch.sh`, so add the reporter to that config once with 
the
+ledger path read from the environment, and `export 
TIKA_EXTRACTS=<extracts-dir>`
+in the same shell as the run. An unset variable fails the load naming it and 
the
+JSON path, so a run cannot silently proceed without its ledger. The file name 
must
+keep the `crashes-` prefix for Compare/Profile to find it by default:
+
+```json
+"pipes-reporters": {
+  "file-system-jsonl-reporter": {
+    "path": "${env:TIKA_EXTRACTS}/.run-info/crashes-batch.jsonl",
+    "includes": ["OOM", "TIMEOUT", "UNSPECIFIED_CRASH", "FAILED_TO_INITIALIZE",
+                 "FETCHER_INITIALIZATION_EXCEPTION", 
"EMITTER_INITIALIZATION_EXCEPTION",
+                 "CLIENT_UNAVAILABLE_WITHIN_MS", "FETCH_EXCEPTION", 
"EMIT_EXCEPTION",
+                 "FETCHER_NOT_FOUND", "EMITTER_NOT_FOUND", 
"PAYLOAD_LIMIT_EXCEEDED"],
+    "onExists": "EXCEPTION",
+    "maxMessageLength": 4096
+  }
+}
+```
+
 ### Notes
 
 - Do NOT pass `-n <N>` as a trailing argument — it confuses the
diff --git a/CHANGES.txt b/CHANGES.txt
index 225d306e66..2fc0b8c098 100644
--- a/CHANGES.txt
+++ b/CHANGES.txt
@@ -1,744 +1,231 @@
-Release 4.1.0 - unreleased
+Release 4.1.0 - 9/21/2026
 
-   * A file name no longer turns plain text into a message type (emlx, mbox,
-     rfc822, news, HTTP capture) when the type's magic has already rejected
-     the bytes (TIKA-4890).
+   * Apple Mail emlx files are detected as message/x-emlx and parsed by
+     RFC822Parser; a file name no longer turns plain text into a message type
+     once the type's magic has rejected the bytes (TIKA-4890).
 
-   * Apple Mail emlx files are detected as message/x-emlx and RFC822Parser 
-     now parses them, dropping the byte-count line and, 
-     when the count is intact, the trailing plist (TIKA-4890).
-
-   * The tika-server UAT gains an inference pass (TIKA-4913).
-     
-   * Calls to a hosted inference engine retry a 429, 502, 503 or 504 answer
-     with a jittered backoff (TIKA-4912).
+   * Retry a 429, 502, 503 or 504 answer from a hosted inference engine with
+     a jittered backoff (TIKA-4912).
 
    * Handle inference exceptions more robustly in PDFParser (TIKA-4911).
 
-   * Add ffmpeg to tika-server's "full" Docker image (TIKA-4910).
-
    * Enable environment variable interpolation in tika-config.json (TIKA-4909).
 
    * Fix a cache-budget leak: tar, 7z and streamed zip entries never released
      their in-memory cache's reservation (TIKA-4908).
-     
-   * Mp3Parser no longer writes the literal "null" (a missing album) or the
-     duration into the body; an untagged MP3 has an empty body, and a TEXT
-     inference binding no longer embeds it (TIKA-4907).
 
-   * VLM parsers: a raw HTML block in the model's markdown (jina-ocr-v1 and
-     DeepSeek-OCR write tables that way) is now emitted as XHTML table 
(TIKA-4906).
+   * Mp3Parser no longer writes "null" (a missing album) or the duration into
+     the body (TIKA-4907).
+
+   * VLM parsers emit a raw HTML block in the model's markdown as XHTML
+     (TIKA-4906).
 
-   * The forked-worker CPU cap (-XX:ActiveProcessorCount) is now clamped at its
+   * The forked-worker CPU cap (-XX:ActiveProcessorCount) is clamped at its
      floor of 2 (TIKA-4905).
-     
-   * Tesseract script models are accepted by their bare name ("language":
-     "Latin", "eng+Japanese_vert") (TIKA-4903).
-     
-   * Per-request tesseract-ocr-parser config no longer accepts
-     otherTesseractConfig (TIKA-4904).
-
-   * ICC tone-curve and lookup-table tags (rTRC and friends) are no longer
-     written to icc:* by default; "jpeg-parser": {"includeIccCurvesAndLuts":
-     true} (likewise tiff-, heif-, bpg-, web-p- and raw-tiff-parser) restores
+
+   * Tesseract script models are accepted by their bare name ("Latin",
+     "eng+Japanese_vert"); per-request tesseract-ocr-parser config no longer
+     accepts otherTesseractConfig (TIKA-4903, TIKA-4904).
+
+   * ICC tone-curve and lookup-table tags are no longer written to icc:* by
+     default; "includeIccCurvesAndLuts": true on the image parsers restores
      them (TIKA-4902).
 
-   * Digesting an embedded document no longer makes its source hand out the
-     bytes twice. (TIKA-4901).
+   * Digesting an embedded document no longer inflates it twice (TIKA-4901).
 
-   * Enable audio/video chunking with ffmpeg and inference (TIKA-4900).
-     The default segment grid is 25 s windows with 5 s overlap: hosted omni
-     engines cap an audio item at 30 s inclusive, so the earlier 30 s default
-     had every audio segment refused.
+   * Enable audio/video chunking with ffmpeg and inference; the default
+     segment grid is 25 s windows with 5 s overlap (TIKA-4900).
 
-   * Refactor OCR and inference to use engines for inference and 
+   * Refactor OCR and inference to use engines for inference and
      page rendering (TIKA-4889).
-     
-   * PDF incremental-update scanning reads the file in blocks instead of one
-     byte at a time; same output, about 4x less CPU (TIKA-4898).
+
+   * PDF incremental-update scanning reads in blocks; ~4x less CPU (TIKA-4898).
 
    * Metadata never stores an unpaired UTF-16 surrogate (TIKA-4897).
 
-   * openai-embedding-engine gains "requestParameters" to enable custom
-     parameters per vendor (TIKA-4896).
-     
-   * New catalog presets, inert until the config names them. (TIKA-4856).
-   
-   * unpack-config gains includeMetadata (TIKA-4881).
-
-   * tika-server: /rmeta gains config/{handlerType} (TIKA-4881).
-     
-   * Tagged PDFs: text drawn twice for a bold effect with one copy tagged
-     came out as letters with spaces between them; each copy now reads as
-     its words. A tree that puts every glyph in its own paragraph (some form
-     generators) is left to the text stripper under AUTO (TIKA-4891).
-     
-   * Inference results in the single-object outputs: with /tika as JSON,
-     tika-app -j, or pipes CONCATENATE, the vectors of an attachment's pages
-     and pictures now land on the one metadata object handed back, each
-     with an embedded locator naming the part, instead of being discarded
-     with the attachment's own metadata. InferenceUnit gains
-     getDestination(), where a task writes; ChunkTarget.of(unit) replaces
-     resolve(target, parent) (TIKA-4895).
-
-   * New tika-eval-structure: a standalone tool that compares two sets of
-     XHTML extracts block by block (paragraph order as Kendall tau, splits and
-     merges, words one side glues that the other splits, artifact and untagged
-     shares), per file and per producer, where token dice cannot see order
-     (TIKA-4891).
-
-   * Improve extraction of tagged PDFs(TIKA-4891).
-
-   * PDF: a bookmark outline is now walked without recursion, and lists nest at
-     most 50 deep with deeper items joining the deepest list (TIKA-4894).
+   * openai-embedding-engine gains "requestParameters" (TIKA-4896).
 
-   * The image embedder writes a picture's vector onto the document the
-     picture appears in (INLINE and RENDERING children) with an "embedded"
-     locator naming the picture; attachments keep their own, and
-     liftToParent: false keeps every picture's vector on the picture.
-     RecursiveParserWrapper exposes the parent's Metadata to a child's parse
-     as ParentMetadata, and ChunkTarget (tika-inference) resolves where a
-     child's chunks land (TIKA-4888).
-     
-   * A PDF page rendered as an embedded document (imageStrategy RENDER_PAGES_*)
-     is enriched once, by the PDF parser: the text recognizer where the OCR
-     strategy says so, every other text-recognizers entry on every page, 
results on
-     the PDF itself. The embedded copy is no longer enriched again by
-     image-parser, so an image embedder no longer yields two vectors for a
-     text-poor page and Tesseract no longer OCRs a page twice; NO_OCR now means
-     no page OCR at all. ContentEnrichers gains scoped suspend and
-     suspendRecognizers for containers that enrich their own renderings
-     (TIKA-4887).
-
-   * OCR engines and other content enrichers: name the engine in 
"text-recognizers"
-     and every default parser stays loaded; "text-recognizers": [] turns 
enrichment
-     off everywhere. With no list, the loader resolves one engine
-     per media type from the loaded parsers at startup (an engine configured 
under
-     "parsers" beats one default-parser found; last registered wins; an entry's
-     _mime-include/_mime-exclude applies) and logs one line per effective 
engine;
-     a collision is a WARN at startup, as are two text recognizers sharing a 
type
-     (both run, their text is concatenated). An engine under "parsers" is a 
parser
-     only for the types no other parser there claims; when every type is 
claimed,
-     startup logs that it acts only as the enricher, or warns that it never 
runs.
-     The image/ocr-* pseudo-types
-     are retired: Content-Type never carries the prefix (extracted images get 
their
-     real extension instead of .bin) and a _mime-include/_mime-exclude naming 
one
-     fails config load with the real type in the message. image-parser now owns
-     image/jp2, image/jpx and image/x-portable-pixmap and calls the OCR engine 
on
-     them. VLM parsers gain "textRecognizer" (default true); set it false for a
-     captioning prompt so AUTO OCR never replaces extracted text with the VLM's
-     output (TIKA-4884).
-
-   * For enricher authors: implement ContentEnricher (TextRecognizer for OCR) 
and
-     advertise the real image types; an enricher the classpath supplied is 
never
-     dispatched to, and one named under "parsers" only fills gaps. A 
third-party
-     engine still advertising image/ocr-* is treated as a legacy text 
recognizer
-     with a WARN until 5.0. An invoked enricher is recorded in tk:parsed-by 
(images)
-     and tk:parsed-by-full-set (PDF page renders) as a dispatched parser would 
be.
-     ContentEnrichers.resolve(Parser) is the load-time resolution.
-     EmbeddedDocumentUtil.normalizeMediaType is deprecated (TIKA-4884).
-
-   * The PDF parser's AUTO OCR strategy no longer emits a page's extracted text
-     next to the OCR output that superseded it. Text that trips the AUTO 
verdict
-     is held back and replaced by OCR, and released unchanged when OCR cannot 
run
-     on that page: no engine on the classpath or in text-recognizers,
-     maxPagesToOcr exhausted, or a timeout or error. AUTO without an engine now
-     behaves as NO_OCR (TIKA-4883).
-
-   * New TextRecognizer capability for content enrichers (tika-core): a parser
-     implements it to say its output is the document's text rather than an
-     annotation. The PDF parser's AUTO strategy supersedes extracted text only
-     when a text recognizer is configured (Tesseract, Tess4J and the VLM 
parsers
-     declare it; an image-embedding enricher does not) and only when the engine
-     actually wrote text. Other enrichers still run on the rendered page
-     (TIKA-4883).
-
-   * tika-app's -m and --json wrote metadata at the SAX endDocument event
-     thereby dropping keys added by a parser after the endDocument (TIKA-4885).
-     
-   * "text-recognizers": _mime-exclude now matches the real media type, so
-     excluding image/tiff also drops a legacy image/ocr-tiff engine (it was
-     a no-op). The zero-media-types check at config load asks the engine
-     itself, so a _mime-include list no longer masks an unreachable one.
-     VLM parsers implement Closeable and release their HTTP client
-     (TIKA-4884).
+   * tika-server: named configuration presets, and catalog presets that are
+     inert until the config names them (TIKA-4856).
+
+   * unpack-config gains includeMetadata; tika-server /rmeta gains
+     config/{handlerType} (TIKA-4881).
+
+   * Improve extraction of tagged PDFs. New tika-eval-structure tool compares
+     two sets of XHTML extracts block by block (TIKA-4891).
+
+   * Inference results in the single-object outputs (/tika as JSON, tika-app
+     -j, pipes CONCATENATE) land on the one metadata object handed back, with
+     an embedded locator, instead of being discarded (TIKA-4895).
 
    * tika-inference and tika-vlm are experimental: classes, config keys and
-     metadata output may change in minor releases without deprecation or
-     compatibility shims (TIKA-4884). So are the "inference" config section,
-     pdf-parser "inference", the org.apache.tika.parser.enricher, .inference
-     and .hook packages, and the tk:chunks and tk:inference-released output:
-     4.2 batches recognition per document, meant as an opt-in addition (an
-     engine that takes one image at a time keeps working unchanged), and adds
-     media inputs and document-level tasks.
-     Stable: the "engines" map, the "text-recognizers" list, and pdf-parser
-     "text"; a 4.1 spelling of those loads as an alias for at least one minor
-     release if it changes (TIKA-4895).
-
-   * A missing OOXML relationship target no longer aborts the whole file:
-     the threaded-comment and person lookups in xlsx and the page lookups in
-     vsdx went straight to POI's getRelatedPart, whose unchecked
-     IllegalArgumentException surfaced as "Error creating OOXML extractor" and
-     dropped the text already extracted. They route through
-     safeGetRelatedPart, as branch_3x already did (TIKA-4879).
-     
-   * RawTiffDetector rejects a BigTIFF directory offset near Long.MAX_VALUE
-     instead of letting the bounds check overflow. Adding the entry-count
-     size to such an offset wrapped negative and read as "already in the
-     prefix", so a 16-byte file threw ArrayIndexOutOfBoundsException out of
-     Detector.detect, which CompositeDetector does not catch: detection
-     failed for the document and the remaining detectors never ran. Raw
-     detection runs on every stream, so this was reachable from every entry
-     point (TIKA-4861).
-     
-   * Entries of ODF, EPUB, GeoGebra, WACZ, XLZ and iWork containers, mbox
-     messages and the AppleSingle data fork are re-opened from their
-     container on rewind instead of cached: digesting rewinds every
-     embedded document, and the cached copy cost heap for the whole entry
-     and, past the cache budget or the 1 MB floor, a temp file. AppleSingle
-     no longer spools its data fork to a temp file on every parse, and
-     GeoGebra no longer spools every embedded picture to detect it
-     (TIKA-4878).
-
-   * A declared Content-Length is no longer treated as a measurement: the
-     zip-bomb ratio counts only measured input bytes (a container-declared
-     size on an embedded document could inflate its denominator), and a
-     re-openable source no longer reserves cache budget or sizes its buffer
-     from the declared length (a lying one could push a small payload to
-     disk or churn the shared budget). Neither is in a release: the
-     exposure arrived with TIKA-4868 and TIKA-4873 (TIKA-4878).
-
-   * Embedded objects in Office documents are re-opened from their container
-     instead of cached: every OOXML part (pictures, media, attachments), the
-     OLE 2.0 package inside an OOXML part, the CONTENTS entry of an OLE 2.0
-     object in a binary Office file, embedded objects in .ppt, XPS page
-     images and Word EMF icons. Digesting an embedded document rewinds it;
-     the cached copy that made possible cost heap for the whole object and,
-     past the cache budget or the 1 MB floor, a temp file. The container
-     hands the bytes back on demand, so neither is needed (TIKA-4878).
-
-   * PDF attachments, PDF XMP packets, 3D on-instantiate scripts and PST
-     attachments are re-opened from their document instead of cached when
-     the embedded-document extractor rewinds them (digesting does, for every
-     embedded document). The cached copy cost heap for the whole attachment
-     and, past the cache budget or the 1 MB floor, a temp file; PDFBox and
-     java-libpst hand the bytes back on demand (TIKA-4878).
-
-   * tika-server: named configuration presets (TIKA-4856).
+     output may change in minor releases. Stable: the "engines" map, the
+     "text-recognizers" list and pdf-parser "text" (TIKA-4884, TIKA-4895).
+
+   * PDF bookmark outlines are walked without recursion; lists nest at most
+     50 deep (TIKA-4894).
+
+   * The image embedder writes a picture's vector onto the document the
+     picture appears in, with an "embedded" locator; liftToParent: false
+     keeps it on the picture (TIKA-4888).
+
+   * A PDF page rendered as an embedded document is enriched once, by the
+     PDF parser; NO_OCR now means no page OCR at all (TIKA-4887).
+
+   * OCR engines are named in "text-recognizers" ("[]" turns enrichment off;
+     no list resolves one engine per media type at startup). The image/ocr-*
+     pseudo-types are retired; a third-party engine still advertising them is
+     a legacy text recognizer with a WARN until 5.0. VLM parsers gain
+     "textRecognizer" (TIKA-4884).
+
+   * PDF AUTO OCR no longer emits a page's extracted text next to the OCR
+     output that superseded it; AUTO without a text recognizer behaves as
+     NO_OCR. New TextRecognizer capability for content enrichers (TIKA-4883).
+
+   * tika-app -m and --json no longer drop metadata keys added after
+     endDocument (TIKA-4885).
+
+   * A missing OOXML relationship target (xlsx threaded comments, vsdx pages)
+     no longer aborts the whole file (TIKA-4879).
+
+   * Raw camera formats (NEF/NRW, PEF/PTX, ARW/SRF/SR2, SRW, DNG, RAF, RW2,
+     MRW, ORF) are detected by content instead of as image/tiff. Fix a
+     BigTIFF offset overflow in RawTiffDetector that failed detection for the
+     whole document (TIKA-4861).
+
+   * Embedded documents in Office files, PDFs, PST and the zip-family
+     containers are re-opened from their container on rewind instead of
+     cached (TIKA-4878).
 
    * Temp files follow -Djava.io.tmpdir on the parent JVM (Tika, its
-     libraries, and forks all honor it); TikaLoader fails at config load
-     if it is unusable. pipes.tempDirectory is deprecated for removal in
-     5.0: it only covered the forks. Do not use tmpfs: spool size is
-     bounded by input, and an orphaned fork dir pins RAM (TIKA-4877).
-
-   * A fork whose parent dies deletes its own temp dir; the parent
-     surfaces a fork's hs_err log before every delete. Failure-path temp
-     file leaks fixed in PDFBoxRenderer, PopplerRenderer, truncated RTF,
-     and MarianTranslator (TIKA-4877).
-     
-   * tika-eval Profile/Compare speedups: single-pass URL/mail stripping
-     replaces the bounded regexes in langdetect preprocessing (same output,
-     17-290x faster on web text), the default H2 db URL sizes the page cache
-     at a quarter of the heap clamped to [64MB, 1GB] (override with
-     -Dtika.eval.h2.cacheSizeKb=<kb>), and the status log adds a last-interval
-     docs-per-sec rate next to the cumulative average (TIKA-4875).
-     
-   * New "text-recognizers" config list (TIKA-4872): select the OCR engine
-     ("tesseract-ocr-parser", "tess4j-parser", "openai-vlm-parser", ...) by
-     name instead of by classpath registration of the image/ocr-* pseudo
-     media types. Enrichers advertise real media types (legacy engines that
-     still advertise image/ocr-* are mapped to the real type, so all are
-     nameable) and are invoked by the image and PDF parsers rather than
-     dispatched to by the composite, so an enricher no longer displaces the
-     parser registered for the same type. Enricher selection uses the
-     detected media type, captured before a parser can refine Content-Type.
-     Every enricher matching a media type runs, in config order (e.g. an
-     OCR engine then a VLM tagger for the same image), best-effort: one
-     enricher's failure does not stop the others and is still reported;
-     timeouts abort the chain. The list is authoritative: a media type no
-     configured enricher matches gets no enrichment -- never a classpath
-     engine that was not named -- and a named engine that reports no media
-     types at load (missing binary, unreachable inference server) fails
-     config load instead of going silently inert. With no
-     "text-recognizers" configured, the legacy ocr-* dispatch applies
-     unchanged; a WARN at config load now names colliding OCR engines and
-     the winner. TesseractOCRParser's
-     component name is pinned as "tesseract-ocr-parser".
-
-   * Inference/OCR hardening (TIKA-4871): OpenAIVLMParser no longer
-     auto-registers via SPI, matching its Claude/Gemini siblings; select
-     it by name ("openai-vlm-parser") in config. Per-request parse-context
-     config for the embedding filters now works and is validated:
-     {"openai-embedding-filter": {"skipEmbedding": true}} (likewise
-     "jina-embedding-filter") merges over the server config, and
-     baseUrl/apiKey/model may not be changed at runtime. The embedding
-     filters release their HTTP client resources on close(). Inline PDF
-     page OCR now accumulates tk:chunks from every OCR'd page onto the
-     parent document instead of keeping only the first page's.
-
-   * Placeholder streams -- the empty stand-ins parsers hand parseEmbedded
-     for content that is never extracted -- report an unknown length rather
-     than their own zero, and the macro-failure entry is registered without
-     parsing its sentinel (TIKA-4874).
-
-   * TikaInputStream.hasReliableLength() distinguishes measured lengths
-     from declared Content-Length hints, and one-shot streams now carry a
-     declared length without spooling; detection sizes its magic read only
-     from measured lengths, so a lying declared length can no longer
-     truncate it (TIKA-4868).
-
-   * ParseContext entries holding per-parse runtime state (ParseRecord,
-     ParseTimeout, the parser-map cache) are skipped during serialization
-     instead of failing as unregistered components (TIKA-4868).
-
-   * Mojibuster's adaptive probe strips incrementally instead of
-     re-stripping the whole buffer on every read (quadratic on tag-heavy
-     pages); JunkDetector's Unicode block lookup uses a precomputed BMP
-     table. Output unchanged (TIKA-4868).
-
-   * Markdown rendering is another ~8x faster on large documents: a custom
-     Text-node renderer emits unescaped spans in bulk instead of
-     commonmark's per-character escape-check-and-append. Byte-identical
-     output, guarded by a fast-vs-stock differential test (TIKA-4868).
-
-   * ZipParser no longer re-decompresses an entry on every rewind when the
-     entry uses a legacy compression method (implode, shrink, bzip2, ...):
-     such entries replay from the budgeted cache instead of re-opening.
-     An imploded 606KB zip drops from 610ms to 102ms (TIKA-4868).
-
-   * More detection/dispatch savings: the message/rfc822 priority-45 magic
-     is gated behind a lossless ':' scan of the first 30 bytes;
-     CompositeParser caches the built type->parser map in the ParseContext
-     so embedded documents reuse the container's map; MagicDetector
-     precomputes a per-pattern first-byte table. Adds a RESOURCE_TIMING
-     log, silenced by default in the shipped log4j2 configs (TIKA-4868).
-
-   * WordExtractor (.doc) cleans each character run and tests paragraph
-     blankness in single passes; ToMarkdownContentHandler collapses line
-     breaks copy-free for clean runs. ~39% off a text-heavy 2MB .doc
+     libraries and forks); pipes.tempDirectory is deprecated. A fork whose
+     parent dies deletes its own temp dir; failure-path temp leaks fixed
+     (TIKA-4877).
+
+   * tika-eval Profile/Compare speedups: langdetect preprocessing, H2 page
+     cache sizing, a per-interval rate in the status log (TIKA-4875).
+
+   * New "text-recognizers" config list selects OCR and enrichment engines by
+     name; enrichers are invoked by the image and PDF parsers rather than
+     dispatched to by the composite (TIKA-4872).
+
+   * OpenAIVLMParser no longer auto-registers via SPI; per-request config for
+     the embedding filters works and locks baseUrl/apiKey/model; inline PDF
+     page OCR accumulates tk:chunks from every page (TIKA-4871).
+
+   * Performance: detection (magic ~35% faster on unmatched input, override
+     keys honored before magic, cached type->parser map), markdown output
+     ~30x faster, .doc cleanup, CSV sniffing, zip legacy-method rewinds, pipes
+     ACK overlap, raw UTF-8 content passback for tika-server
+     (content-bytes-config). Compat: with a Content-Type override set,
+     DefaultDetector no longer lets a more specific magic result overrule it
      (TIKA-4868).
 
-   * tika-server's raw-output endpoints (/tika, /tika/text, ...) carry the
-     extracted content as raw UTF-8 bytes from the pipes worker to the HTTP
-     response instead of a Smile-encoded string (9MB text: 115ms -> 75ms).
-     Opt-in via the new content-bytes-config parse-context component, which
-     moves CONTENT_ONLY passback content out of tk:content into
-     EmitData.getContentBytes(); results routed to a regular Emitter get
-     the content restored to the metadata (TIKA-4868).
-
-   * Detection hot-path cleanups: MagicMatch resolves its detector via
-     double-checked locking; glob patterns are compiled once at
-     registration; MimeTypes.forName reads a ConcurrentHashMap (fixing an
-     unsynchronized-read race) and indexes normalized keys; resource names
-     containing spaces skip the URI-parse-by-exception; the magic-header
-     buffer is sized by the stream's measured length instead of a fixed
-     64KB; the Adobe Illustrator ranged regex is gated behind a literal
-     scan. Detection results unchanged (TIKA-4868).
-
-   * Magic detection is ~35% faster on unmatched (e.g. plain-text) input:
-     range scans find first-byte candidates before running the full
-     masked/case-folded compare (TIKA-4868).
-
-   * CSVSniffer reads its detection window once into a shared buffer and
-     runs every delimiter hypothesis against it; windows with no delimiter
-     and no quote skip the scan outright. Results unchanged (TIKA-4868).
-
-   * Pipes workers no longer stall between pre-parse and parse waiting
-     for the client to acknowledge the intermediate-result frame; the
-     ACK round trip now overlaps the parse. Adds per-request timing logs
-     on org.apache.tika.pipes.timing.*, silenced by default in the shipped
-     log4j2 configs; raise that logger to info to enable (TIKA-4868).
-
-   * DefaultDetector honors CONTENT_TYPE_USER_OVERRIDE and
-     CONTENT_TYPE_PARSER_OVERRIDE before running magic detection, matching
-     CompositeDetector's contract. Removes the second full magic scan every
-     pipes parse paid per document. Compat note: with either override set,
-     DefaultDetector no longer lets a more specific magic result overrule
-     it; parts whose parser declares a type from container headers (e.g.
-     inline text/* mail parts) now report the declared type, and
-     CONTENT_TYPE_MAGIC_DETECTED is not recorded when an override short
-     circuits detection (TIKA-4868).
-
-   * Markdown output is ~4x faster on large documents:
-     ToMarkdownContentHandler now buffers the commonmark renderer's
-     per-character writes instead of paying the synchronized
-     Writer.write(int) cost for every character (TIKA-4868).
-   * Embedded documents carry their size: ParsingEmbeddedDocumentExtractor
-     sets Content-Length from the stream where the stream knows it and the
-     parser did not say, which never spools to measure one, and the raw
-     camera previews and the audio cover art set the length they read from
-     the file (TIKA-4873).
-
-   * AVIF images are parsed rather than only detected: HeifParser accepts
-     image/avif, which is the same ISO-BMFF container, so dimensions, EXIF
-     and XMP come out of it the way they do for HEIC (TIKA-4870).
-
-   * The video of a Google/Android motion photo, appended after the image and
-     described by the Motion Photo or MicroVideo XMP, is emitted as an
-     ATTACHMENT embedded document named after what the file declares it to
-     be. Nothing is emitted, and nothing is recorded, when the declared
-     length does not fit the file or the bytes there are not recognized,
-     which is what sharing a motion photo out of a gallery leaves behind
-     (TIKA-4869).
-
-   * Raster previews for the vector thumbnails of Office documents: the new
-     poi-metafile-renderer draws EMF and WMF images through POI (a PNG of
-     a configurable width; Word's bitmap-in-WMF thumbnails from the bitmap
-     directly), EMFParser and WMFParser are RenderingParsers that emit the
-     rendering as a RENDERING embedded document with "emf-parser" /
-     "wmf-parser": {"renderImage": true, "renderWidth": 800}, off by
-     default and restrictable to e.g. THUMBNAIL embedded documents with
-     "renderOnlyEmbeddedResourceTypes", and OfficeParser emits the
-     SummaryInformation thumbnail of the OLE2 formats (a WMF) as a THUMBNAIL
-     embedded document, as the OOXML parsers do with the docProps thumbnail,
-     switchable with "office-parser": {"extractThumbnail": false}
-     (TIKA-4855).
-
-   * Add "exception-reporting" parse-context config to redact and bound
+   * Embedded documents carry Content-Length where the stream knows it
+     (TIKA-4873).
+
+   * AVIF images are parsed by HeifParser (dimensions, EXIF, XMP) (TIKA-4870).
+
+   * The video of a Google/Android motion photo is emitted as an ATTACHMENT
+     embedded document (TIKA-4869).
+
+   * New poi-metafile-renderer rasterizes EMF/WMF (opt-in "renderImage");
+     OfficeParser emits the OLE2 SummaryInformation thumbnail as a THUMBNAIL
+     embedded document ("extractThumbnail": false disables) (TIKA-4855).
+
+   * New "exception-reporting" parse-context config redacts and bounds
      exception text in metadata, tika-server error bodies and pipes/grpc
-     messages; FileSystemEmitter writes atomically (TIKA-4848).
-     Compat notes: a truncated TSD envelope now records its read failure
-     under tk:exception:embedded-stream-exception rather than
-     tk:exception:embedded-exception; recordException and
-     recordEmbeddedStreamException no longer strip a bare TikaException
-     wrapper, so the first line of tk:exception:* values may name the
-     wrapper (affects consumers keyed on that line, e.g. eval cause
-     counts across the 4.1 boundary).
-     
-   * Audio cover art is emitted as a THUMBNAIL embedded document, like the
-     preview image of the document container formats: the front cover (ID3
-     APIC and FLAC/Vorbis picture type 3), else the first picture of type
-     "Other" or unknown, else the first picture, and the first covr image
-     of an MP4. Further pictures
-     stay INLINE. Clients that looked for cover art as INLINE need to
-     accept THUMBNAIL as well (TIKA-4850).
-
-   * tika-grpc resolves its plugin-roots fallback against the install
-     layout via DefaultPluginsDir instead of a working-directory-relative
-     pf4j default, and a WARN names the resolved directory when no plugins
-     directory exists (TIKA-4865).
-
-   * The tika-server full and tika-grpc Docker images install fonts-noto-cjk:
-     without any CJK face, PDFs using non-embedded CJK fonts render (and OCR)
-     as .notdef boxes in every renderer, even though the images ship Japanese
-     tesseract data (TIKA-4866).
-     
-   * The default plugins directory is resolved against the install layout
-     (next to the jar, or next to its lib/ directory) and always as an
-     absolute path, shared by tika-server, PipesForkParser and the async
-     CLI; it no longer depends on the working directory (TIKA-4864).
-
-   * tika-server error bodies (the 422/500 exception mapper, /meta/{field})
-     now honor the exception-reporting policy; /meta/{field} returns the
-     already-formatted container exception instead of re-wrapping it with
-     server frames. Completes TIKA-4848 (TIKA-4848).
-
-   * The exception-reporting policy now also governs the messages a pipes
-     worker returns (fetch/emit/crash) and the container exception it
-     records; part of TIKA-4848 step 3 (TIKA-4848).
-     
+     messages; FileSystemEmitter writes atomically. Compat: tk:exception:*
+     values may now begin with the TikaException wrapper line (TIKA-4848).
+
+   * Audio cover art (ID3/FLAC/Vorbis front cover, MP4 covr) is emitted as a
+     THUMBNAIL embedded document rather than INLINE (TIKA-4850).
+
+   * The default plugins directory (tika-server, PipesForkParser, async CLI)
+     and tika-grpc's plugin-roots fallback are resolved against the install
+     layout instead of the working directory (TIKA-4864, TIKA-4865).
+
    * Allow image compression settings in PDFBox-based renderer (TIKA-4862).
 
-   * The tika-server full and tika-grpc Docker images set OMP_THREAD_LIMIT=1:
-     to avoid oversubscribing the CPU under forked parse workers (TIKA-4863).
-
-   * embedded-limits maxDepth counts embedding levels again instead of the
-     parsers a parse passes through; with AutoDetectParser over DefaultParser
-     every value above 1 used to stop one level early (TIKA-4857).
-
-   * GeoGebraParser emits the icon of a tool (*.ggt, the macro's iconFile)
-     as its THUMBNAIL embedded document; tool files have no thumbnail of
-     their own (TIKA-4831).
-
-   * Enum values in JSON configuration are matched case-insensitively, so
-     "no_ocr" works as well as "NO_OCR"; the server docs used the lower-case
-     form in their examples (TIKA-4859).
-
-   * The preview image of iWork '09 packages (QuickLook/Thumbnail.jpg) and of
-     iWork '18 packages (preview.jpg) is emitted as a THUMBNAIL embedded
-     document, as it already was for iWork '13 (TIKA-4854).
-
-   * Raw camera formats are detected by content: RawTiffDetector tells
-     Nikon NEF/NRW, Pentax PEF/PTX, Sony ARW/SRF/SR2, Samsung SRW and Adobe
-     DNG from a plain TIFF by their image directory (DNGVersion, the vendor
-     Compression codes, or a CFA/LinearRaw image plus Make), and Fuji RAF,
-     Panasonic RW2, Minolta MRW and the remaining Olympus ORF byte orders
-     get magic entries. Streams without a file name used to be image/tiff.
-     image/x-raw-samsung (*.srw) is new and parsed by RawTiffParser
-     (TIKA-4861).
-
-   * DWGReadParser emits the drawing's THUMBNAILIMAGE as a THUMBNAIL embedded
-     document instead of INLINE (TIKA-4853).
-
-   * EpubParser emits the cover image named by the OPF (the EPUB 3
-     cover-image manifest property, or the EPUB 2 cover meta) as a THUMBNAIL
-     embedded document (TIKA-4852).
-
-   * RawTiffParser marks only the largest embedded JPEG preview as the
-     THUMBNAIL embedded document; the smaller previews of the same image are
-     INLINE images named image-N.jpg. Previously every preview was a
-     THUMBNAIL, so a client had to compare them to find the representative
-     one (TIKA-4851).
-
-   * tika-eval: Profile/Compare accept the batch run's jsonl crash ledger
-     (--pipesReport, -pa/-pb) and a run-info json (--runInfo, -ra/-rb), and
-     read both from <extracts>/.run-info/ by default (refusing an ambiguous
-     dir). containers gains pipes_status/pipes_message; a new run_info table
-     records eval and batch provenance; reports and
-     summary.md classify NO_EXTRACT_FILE by ledger status (CRASH, the raw
-     status, NO_PIPES_RECORD, BATCH_WITHOUT_LEDGER, NO_PIPES_REPORT_SUPPLIED).
-     Report on a db from an earlier tika-eval skips the reports it cannot run
-     instead of aborting (TIKA-4847).
-
-   * New file-system-jsonl-reporter pipes reporter (TIKA-4846).
-   
+   * The tika-server full and tika-grpc Docker images add ffmpeg and
+     fonts-noto-cjk, and set OMP_THREAD_LIMIT=1 to avoid oversubscribing the
+     CPU under forked parse workers (TIKA-4910, TIKA-4866, TIKA-4863).
+
+   * embedded-limits maxDepth counts embedding levels again; values above 1
+     used to stop one level early (TIKA-4857).
+
+   * New GeoGebraParser for *.ggb/*.ggs/*.ggt with content-based detection
+     and THUMBNAIL embedded documents (TIKA-4831).
+
+   * Enum values in JSON configuration are matched case-insensitively
+     (TIKA-4859).
+
+   * iWork '09 and '18 preview images are emitted as THUMBNAIL embedded
+     documents (TIKA-4854).
+
+   * DWGReadParser emits THUMBNAILIMAGE as THUMBNAIL instead of INLINE
+     (TIKA-4853).
+
+   * EpubParser emits the OPF cover image as a THUMBNAIL embedded document
+     (TIKA-4852).
+
+   * RawTiffParser marks only the largest JPEG preview as THUMBNAIL; smaller
+     previews are INLINE (TIKA-4851).
+
+   * New file-system-jsonl-reporter pipes reporter records a batch run's
+     crashes; tika-eval Profile/Compare read that ledger and the run-info
+     json to classify NO_EXTRACT_FILE by cause (TIKA-4846, TIKA-4847).
+
    * Stop spooling OLE2 objects whose header over-reserves BAT capacity
      (TIKA-4845).
 
    * Add Micrometer reporting and opt-in endpoint for tika-server (TIKA-4839).
-     
-   * Improve spooling/decrease number of spills to disk (TIKA-4835).
-
-   * Fixed a bug that made per-request (parse-context) configuration unusable
-     for parsers that lock some config fields against caller modification --
-     Tess4J, the VLM parsers and the OpenAI image-embedding parser. Any such
-     config threw, including an empty one: the defaults were deep-copied
-     through their own setters, which the runtime config overrides to reject
-     caller input, so the copy tripped the parser's own guards before the
-     caller's JSON was read. Locked fields are still rejected when a caller
-     actually sets them. Configuration supplied at initialization time (the
-     "parsers" section) was never affected (TIKA-4843).
-     
-   * OOXML parsers flag package parts that are unreachable through the OPC
-     relationship graph: msoffice:has-unreferenced-parts (boolean) and
-     msoffice:unreferenced-part-names. Purely structural (no bytes are
-     inspected; content types come from [Content_Types].xml by extension), so
-     expect false positives from tools that leave orphan parts behind. A hiding
-     place a raw-ZIP scanner can still see, not a statement about what Tika
-     parsed. Applies to Word, Excel, PowerPoint and Visio OOXML (including
-     macro-enabled variants); XPS links content by markup rather than
-     relationships and is not checked (TIKA-4837).
-     
-   * Shared pipes server (useSharedServer: true, not the default): a client 
whose
-     in-flight parse was killed by another client's restart could restart the
-     healthy replacement. ensureRunning holds its lock across the whole fork, 
so
-     siblings cannot report a dead worker until after the replacement is up, 
and
-     the pending-restart flag carried no process identity -- so a report about
-     the process that just died was applied to its successor, which was then
-     destroyed and re-forked. One worker death produced two restarts and a 
second
-     round of destroyed in-flight work; under sustained concurrent load it
-     sustained itself at one spurious restart per round, appearing as periodic
-     unexplained worker churn and intermittent parse failures that succeed on
-     retry. Each fork now carries a generation that clients capture when they
-     connect and hand back with every report, and reports about a superseded
-     process are dropped. Also fixed in shared mode: ensureRunning could fork a
-     replacement after shutdown() that nothing owned and nothing would ever
-     destroy, and an interrupt during process teardown left the process handle
-     pointing at a killed process and leaked the temp directory. Affects 4.0.0
-     and earlier (TIKA-4844).
-
-   * tika-pipes: the cache memory budget (how much rewindable content a forked
-     worker keeps in memory before spilling to disk; new since 4.0.0, which had
-     no budget at all) defaults to a quarter of the fork's heap, so raising
-     -Xmx raises it. It is one pool per forked JVM shared by all of its 
threads.
-     -Dtika.pipes.cacheMemoryBudgetBytes in forkedJvmArgs overrides it (below
-     the quarter-heap ceiling; <=0 disables); the fork logs the value and its
-     source at startup. TikaInputStream.hasFile() now also reports content the
-     stream cache spilled on its own, not only content a getPath() call put on
-     disk; note getPath() may still have to drain the rest of the source into
-     that file. TikaInputStream.toString() no longer forces a spill, so logging
-     or debugger-inspecting a stream is side-effect-free.
-     TikaInputStream.inMemoryContent(channel) gives a zero-copy read-only view
-     of cached content for consumers that need random access. Digester
-     gains digestSink(), a DigestSink that digests as it is written; nothing is
-     written to the metadata unless the producer calls commit(), so any failure
-     -- exception, Error, or a producer that closes the sink itself -- 
publishes
-     no digest rather than a digest of the bytes that happened to arrive. A
-     translator that claims a stream and writes nothing likewise publishes
-     nothing: embedded PST mail items, whose translator is still a stub, no
-     longer carry the digest of zero bytes (the same value for every one of
-     them) and instead carry no digest at all. DigestHelper uses it for
-     translated embedded streams, which no longer touch a temp file when the
-     digester implements digestSink (all of Tika's do; one that only implements
-     digest() still buffers).
-     TemporaryResources.closeAll(Closeable...) closes every argument even when
-     one throws unchecked; TemporaryResources, CachingSource, 
CachingInputStream
-     and CompositeDigester use it (TIKA-4835).
-
-   * Documentation: corrected a batch of pages and javadocs that contradicted
-     the code. Notably: the ES/OpenSearch attachmentStrategy has no default
-     (unset means embedded documents get neither the parent field nor the
-     parent/child relation); Kafka's connectionsMaxIdleMs is passed to the
-     producer, not ignored; jdbc queryTimeoutSeconds is applied only when > 0,
-     so 0 does not mean "no limit"; the Solr emitter/iterator support only
-     basic auth, not ntlm, and only when a userName is set; pipes-reporters
-     silently loads zero reporters when given a JSON array, and
-     pipes-iterator/pipes-reporters instances are built at config load rather
-     than lazily; under CONTENT_ONLY only a parse-context filter replaces the
-     built-in one, not the top-level metadata-filters chain;
-     _mime-include/_mime-exclude also accept a bare string; Tess4J locks
-     poolSize and maxImagePixels as well as the two paths; and pdf:trapped and
-     xmp:pdf:Trapped are new 4.x keys rather than renames (3.x captured the
-     flag only as pdf:docinfo:trapped and dropped the XMP value). Also
-     corrected the config nesting shown in every pipes-plugin fetcher/emitter
-     javadoc -- 23 of them, which had it inverted (the instance id is the
-     outer key, the component name the inner) -- and removed references to a
-     TesseractOCRConfig.properties file that 4.x does not load (TIKA-4842).
-
-   * Pipes plugins no longer bundle their own Jackson: jackson-core, -databind
-     and -annotations are provided by the host (tika-serialization) and the
-     plugins parent pom now bans bundling them, so a mapper can cross the
-     plugin boundary without a second copy of the Jackson classes (seven plugin
-     zips shipped one). Plugin configuration JSON is parsed by one shared
-     mapper, PluginJson (tika-plugins-core), which rejects unknown keys,
-     numbers for enums and duplicate keys, and accepts
-     // and /* */ comments; the 33 per-plugin *Config classes use it instead of
-     their own bare ObjectMapper (TIKA-4840).
-
-   * tika-server and tika-async-cli now start from a config that contains
-     // or /* */ comments, as the configuration docs have always said they
-     may. The main loader accepted them; the steps that re-read the user's
-     file to merge in server/CLI overrides (ConfigMerger, ensurePluginRoots)
-     used their own bare parser and refused the whole file; they now use the
-     shared TikaObjectMapperFactory mapper (TIKA-4834).
-
-   * The Kafka pipes iterator no longer stops at the first empty poll. A newly
-     subscribed consumer spends its first poll(s) joining the group and returns
-     empty even when the topic has a backlog, so the iterator could enqueue 
zero
-     files and report success. It now waits for a partition assignment (bounded
-     by the new assignmentTimeoutMs, default 30s) and requires a continuous 
quiet
-     window (drainIdleMs, default 1s) before concluding the topic is drained.
-     groupInitialRebalanceDelayMs is deprecated and no longer sent to the
-     consumer: it is a broker setting that Kafka has always ignored 
(TIKA-4833).
-
-   * Pipes IPC: carry inline document bytes as a raw binary field beside the
-     tuple in the request envelope -- never inside the tuple or its
-     ParseContext -- and disable Smile's 7-bit binary encoding. Tuple JSON
-     serialized by 4.0.0 with an "inline-bytes" parse-context entry no longer
-     loads; it is rejected with a tailored message (TIKA-4829).
-
-   * Digesting embedded documents no longer buffers each embedded object to a
-     temp file. Zip entries are re-read from the parent archive on rewind, and
-     a new process-wide CacheMemoryBudget (seeded by the pipes forked server;
-     default 256MB, clamped to a quarter of the fork's heap; tunable via
-     -Dtika.pipes.cacheMemoryBudgetBytes in the config's forkedJvmArgs, <=0
-     disables) lets embedded objects stay in memory past the per-object 1MB
-     threshold. New public API on TikaInputStream: get(IOSupplier,...),
-     enableRewind(CacheMemoryBudget), getSeekableByteChannel(). Zip/7z/epub/odf
-     parsing and zip container detection now read through seekable channels, so
-     after detection/parsing a TikaInputStream may no longer be file-backed
-     (hasFile() false); getPath()/getFile() still work and spool on demand
-     (TIKA-4828).
-     
-   * Pipes now carries the caller-supplied Content-Type across the worker's
-     fresh-metadata boundary as a soft detection hint, so every forked-parse
-     endpoint (/tika, /meta, /rmeta, /unpack, /async, /pipes, plus tika-grpc
-     and embedded PipesForkParser) can route on a client Content-Type, not
-     only on the filename. Detection keeps the hint only when it equals or
-     specializes the content-detected type (e.g. refining image/tiff to
-     image/x-canon-cr2); for bytes with no magic it can select any type,
-     matching the routing power the filename already had. The
-     CONTENT_TYPE_USER_OVERRIDE key is deliberately not carried, so the hint
-     cannot force an unrelated type (TIKA-4825).
-
-   * OneNote extraction now follows document order, omits superseded page
-     revisions, sorts author metadata, extracts embedded object BLOBs, and
-     bounds malformed-input recursion and file-derived allocations. Parse
-     warnings and embedded relationship IDs are exposed in metadata. Malformed
-     or truncated files that cannot be fully parsed, and files whose walk
-     yields no content, now fall back to the legacy string dump instead of
-     failing or returning empty output. The legacy MS-ONESTORE walker bounds
-     its recursion (depth caps plus file-node-list and fragment-chain cycle
-     guards) and now honors shouldParseEmbedded for embedded file data
-   * PDF: extractFontNames threw NullPointerException on a page with no
-     /Resources dictionary (TIKA-4842).
-
-   * tika-server: opt-in Micrometer metrics reporting and endpoint
-     (TIKA-4839).
-
-   * Per-request (parse-context) config for parsers that lock fields
-     (Tess4J, VLM, OpenAI image-embedding) threw even when empty; locked
-     fields are still rejected when actually set (TIKA-4843).
+
+   * Improve spooling/decrease number of spills to disk. The pipes cache
+     memory budget defaults to a quarter of the fork heap
+     (-Dtika.pipes.cacheMemoryBudgetBytes overrides) (TIKA-4835).
+
+   * Per-request config for parsers that lock fields (Tess4J, VLM, OpenAI
+     image-embedding) threw even when empty (TIKA-4843).
 
    * OOXML: new msoffice:has-unreferenced-parts and
-     msoffice:unreferenced-part-names flag package parts unreachable via the
-     OPC relationship graph. Structural only, expect false positives; not
-     applied to XPS (TIKA-4837).
-
-   * Shared pipes server (useSharedServer: true): a worker death could trigger
-     a second, spurious restart that killed the healthy replacement. Forks now
-     carry a generation; stale reports are dropped. Also fixed: a fork after
-     shutdown() that was never destroyed, and a temp-dir leak on interrupt
-     during teardown (TIKA-4844).
-
-   * tika-pipes cache memory budget defaults to a quarter of the fork heap;
-     override with -Dtika.pipes.cacheMemoryBudgetBytes in forkedJvmArgs
-     (<=0 disables). TikaInputStream: hasFile() also reports cache spills,
-     toString() no longer spills, new inMemoryContent(channel). Digester gains
-     digestSink(); a digest is published only on commit(), so failed or empty
-     translations (e.g. stub PST items) publish no digest. New
-     TemporaryResources.closeAll(Closeable...) (TIKA-4835).
-
-   * Docs/javadocs reconciled with the code: ES/OpenSearch attachmentStrategy
-     has no default; Kafka connectionsMaxIdleMs is honored; jdbc
-     queryTimeoutSeconds 0 is not "no limit"; Solr basic auth only; pipes
-     reporters/iterators are built at config load; Tess4J also locks poolSize
-     and maxImagePixels; pdf:trapped is new, not renamed; plugin config
-     nesting fixed in 23 javadocs (TIKA-4842).
-
-   * Pipes plugins no longer bundle Jackson; the host provides it. Plugin
-     config is parsed by a shared strict PluginJson mapper (rejects unknown
-     and duplicate keys; accepts comments) (TIKA-4840).
+     msoffice:unreferenced-part-names (structural only; not XPS) (TIKA-4837).
+
+   * Shared pipes server: a worker death could trigger a spurious second
+     restart; forks now carry a generation (TIKA-4844).
+
+   * Docs and javadocs reconciled with the code; fixes an
+     extractFontNames NullPointerException on a PDF page with no /Resources
+     (TIKA-4842).
+
+   * Pipes plugins no longer bundle Jackson; plugin config is parsed by a
+     shared strict mapper that rejects unknown and duplicate keys (TIKA-4840).
 
    * tika-server and tika-async-cli accept // and /* */ comments in config
-     during override merging, as documented (TIKA-4834).
+     during override merging (TIKA-4834).
 
-   * Kafka pipes iterator no longer stops on the first empty poll; waits for
-     partition assignment (assignmentTimeoutMs, 30s) and a quiet window
-     (drainIdleMs, 1s). groupInitialRebalanceDelayMs is deprecated
-     (TIKA-4833).
+   * Kafka pipes iterator no longer stops on the first empty poll
+     (assignmentTimeoutMs, drainIdleMs); groupInitialRebalanceDelayMs is
+     deprecated (TIKA-4833).
 
    * Pipes IPC carries inline bytes as a raw binary field, not in the tuple;
-     Smile 7-bit binary encoding disabled. 4.0.0 tuples with an "inline-bytes"
-     parse-context entry are rejected (TIKA-4829).
+     4.0.0 tuples with an "inline-bytes" parse-context entry are rejected
+     (TIKA-4829).
 
    * Digesting embedded documents no longer spools each to a temp file; a
-     process-wide CacheMemoryBudget (default 256MB) keeps them in memory. New
-     TikaInputStream API: get(IOSupplier,...), enableRewind(CacheMemoryBudget),
-     getSeekableByteChannel(). Zip-family parsing and detection use seekable
-     channels, so hasFile() may be false afterward; getPath() still spools on
-     demand (TIKA-4828).
+     process-wide CacheMemoryBudget keeps them in memory. Zip-family parsing
+     uses seekable channels, so hasFile() may be false afterward (TIKA-4828).
 
    * Pipes carries the client Content-Type into the forked worker as a
-     detection hint for all forked endpoints; honored only when it equals or
-     specializes the detected type, or when there is no magic. The
-     user-override key is not carried (TIKA-4825).
-
-   * OneNote: document-order extraction, superseded revisions omitted, embedded
-     BLOBs extracted, warnings and relationship IDs in metadata, bounded
-     recursion/allocation; malformed files fall back to the legacy string dump
-     (TIKA-4814).
-
-   * New GeoGebraParser for *.ggb/*.ggs/*.ggt: geogebra:* metadata, text and
-     the thumbnail as a THUMBNAIL embedded document, with content-based
-     detection. Previously typed application/zip with every entry as an
-     attachment. *.ggs and *.ggp are new mime types; *.ggp is glob-only
-     (TIKA-4831).
-
-   * RawTiffParser extracts the camera-generated JPEG previews embedded in
-     TIFF-based raw images (Nikon NEF/NRW, Sony ARW/SRF/SR2, Pentax PEF/PTX,
-     Adobe DNG and Canon CR2, including BigTIFF DNG containers) as thumbnail
-     embedded documents. image/x-raw-{nikon,sony,pentax,adobe} are now
-     sub-classes of image/tiff, so a named NEF/ARW/PEF/DNG that used to detect
-     as image/tiff (TiffParser, metadata only) now detects as image/x-raw-* and
-     emits thumbnail-N.jpg attachments in /rmeta and /unpack; CR2 keeps its
-     detection but also gains the attachments. Disable via
-     "raw-tiff-parser": {"extractPreviews": false} (TIKA-4824).
-   * RawTiffParser extracts embedded JPEG previews from NEF/NRW, ARW/SRF/SR2,
-     PEF/PTX, DNG and CR2 as thumbnail embedded documents. image/x-raw-* are
-     now subtypes of image/tiff, so named raw files detect as image/x-raw-*.
-     Disable with "raw-tiff-parser": {"extractPreviews": false} (TIKA-4824).
+     detection hint for every forked endpoint (TIKA-4825).
+
+   * OneNote: document-order extraction, superseded revisions omitted,
+     embedded BLOBs extracted, bounded recursion; malformed files fall back
+     to the legacy string dump (TIKA-4814).
+
+   * RawTiffParser extracts embedded JPEG previews as THUMBNAIL embedded
+     documents; image/x-raw-* are now subtypes of image/tiff
+     ("extractPreviews": false disables) (TIKA-4824).
 
 Release 4.0.0 - 8/18/2026
 

Reply via email to