This is an automated email from the ASF dual-hosted git repository.
tballison pushed a commit to branch main
in repository https://gitbox.apache.org/repos/asf/tika.git
The following commit(s) were added to refs/heads/main by this push:
new fd4837da79 prep for 4.1.0 release (#3220)
fd4837da79 is described below
commit fd4837da79c70004ba12435f2446d4b9c92da266
Author: Tim Allison <[email protected]>
AuthorDate: Mon Sep 21 09:34:15 2026 -0400
prep for 4.1.0 release (#3220)
---
.skills/devs/tika-eval-compare/SKILL.md | 21 +
CHANGES.txt | 879 +++++++-------------------------
2 files changed, 204 insertions(+), 696 deletions(-)
diff --git a/.skills/devs/tika-eval-compare/SKILL.md
b/.skills/devs/tika-eval-compare/SKILL.md
index 95c69fc479..96fa6c6972 100644
--- a/.skills/devs/tika-eval-compare/SKILL.md
+++ b/.skills/devs/tika-eval-compare/SKILL.md
@@ -108,6 +108,27 @@ ships the jsonl reporter (TIKA-4846),
`crashes-<run.id>.jsonl` into
(`-ra/-rb`, `-pa/-pb` override). A baseline without the reporter gets
run-info but no ledger. Needs python3.
+Launching tika-app directly (a standing run script with its own config, as on a
+corpus box) skips `run-batch.sh`, so add the reporter to that config once with
the
+ledger path read from the environment, and `export
TIKA_EXTRACTS=<extracts-dir>`
+in the same shell as the run. An unset variable fails the load naming it and
the
+JSON path, so a run cannot silently proceed without its ledger. The file name
must
+keep the `crashes-` prefix for Compare/Profile to find it by default:
+
+```json
+"pipes-reporters": {
+ "file-system-jsonl-reporter": {
+ "path": "${env:TIKA_EXTRACTS}/.run-info/crashes-batch.jsonl",
+ "includes": ["OOM", "TIMEOUT", "UNSPECIFIED_CRASH", "FAILED_TO_INITIALIZE",
+ "FETCHER_INITIALIZATION_EXCEPTION",
"EMITTER_INITIALIZATION_EXCEPTION",
+ "CLIENT_UNAVAILABLE_WITHIN_MS", "FETCH_EXCEPTION",
"EMIT_EXCEPTION",
+ "FETCHER_NOT_FOUND", "EMITTER_NOT_FOUND",
"PAYLOAD_LIMIT_EXCEEDED"],
+ "onExists": "EXCEPTION",
+ "maxMessageLength": 4096
+ }
+}
+```
+
### Notes
- Do NOT pass `-n <N>` as a trailing argument — it confuses the
diff --git a/CHANGES.txt b/CHANGES.txt
index 225d306e66..2fc0b8c098 100644
--- a/CHANGES.txt
+++ b/CHANGES.txt
@@ -1,744 +1,231 @@
-Release 4.1.0 - unreleased
+Release 4.1.0 - 9/21/2026
- * A file name no longer turns plain text into a message type (emlx, mbox,
- rfc822, news, HTTP capture) when the type's magic has already rejected
- the bytes (TIKA-4890).
+ * Apple Mail emlx files are detected as message/x-emlx and parsed by
+ RFC822Parser; a file name no longer turns plain text into a message type
+ once the type's magic has rejected the bytes (TIKA-4890).
- * Apple Mail emlx files are detected as message/x-emlx and RFC822Parser
- now parses them, dropping the byte-count line and,
- when the count is intact, the trailing plist (TIKA-4890).
-
- * The tika-server UAT gains an inference pass (TIKA-4913).
-
- * Calls to a hosted inference engine retry a 429, 502, 503 or 504 answer
- with a jittered backoff (TIKA-4912).
+ * Retry a 429, 502, 503 or 504 answer from a hosted inference engine with
+ a jittered backoff (TIKA-4912).
* Handle inference exceptions more robustly in PDFParser (TIKA-4911).
- * Add ffmpeg to tika-server's "full" Docker image (TIKA-4910).
-
* Enable environment variable interpolation in tika-config.json (TIKA-4909).
* Fix a cache-budget leak: tar, 7z and streamed zip entries never released
their in-memory cache's reservation (TIKA-4908).
-
- * Mp3Parser no longer writes the literal "null" (a missing album) or the
- duration into the body; an untagged MP3 has an empty body, and a TEXT
- inference binding no longer embeds it (TIKA-4907).
- * VLM parsers: a raw HTML block in the model's markdown (jina-ocr-v1 and
- DeepSeek-OCR write tables that way) is now emitted as XHTML table
(TIKA-4906).
+ * Mp3Parser no longer writes "null" (a missing album) or the duration into
+ the body (TIKA-4907).
+
+ * VLM parsers emit a raw HTML block in the model's markdown as XHTML
+ (TIKA-4906).
- * The forked-worker CPU cap (-XX:ActiveProcessorCount) is now clamped at its
+ * The forked-worker CPU cap (-XX:ActiveProcessorCount) is clamped at its
floor of 2 (TIKA-4905).
-
- * Tesseract script models are accepted by their bare name ("language":
- "Latin", "eng+Japanese_vert") (TIKA-4903).
-
- * Per-request tesseract-ocr-parser config no longer accepts
- otherTesseractConfig (TIKA-4904).
-
- * ICC tone-curve and lookup-table tags (rTRC and friends) are no longer
- written to icc:* by default; "jpeg-parser": {"includeIccCurvesAndLuts":
- true} (likewise tiff-, heif-, bpg-, web-p- and raw-tiff-parser) restores
+
+ * Tesseract script models are accepted by their bare name ("Latin",
+ "eng+Japanese_vert"); per-request tesseract-ocr-parser config no longer
+ accepts otherTesseractConfig (TIKA-4903, TIKA-4904).
+
+ * ICC tone-curve and lookup-table tags are no longer written to icc:* by
+ default; "includeIccCurvesAndLuts": true on the image parsers restores
them (TIKA-4902).
- * Digesting an embedded document no longer makes its source hand out the
- bytes twice. (TIKA-4901).
+ * Digesting an embedded document no longer inflates it twice (TIKA-4901).
- * Enable audio/video chunking with ffmpeg and inference (TIKA-4900).
- The default segment grid is 25 s windows with 5 s overlap: hosted omni
- engines cap an audio item at 30 s inclusive, so the earlier 30 s default
- had every audio segment refused.
+ * Enable audio/video chunking with ffmpeg and inference; the default
+ segment grid is 25 s windows with 5 s overlap (TIKA-4900).
- * Refactor OCR and inference to use engines for inference and
+ * Refactor OCR and inference to use engines for inference and
page rendering (TIKA-4889).
-
- * PDF incremental-update scanning reads the file in blocks instead of one
- byte at a time; same output, about 4x less CPU (TIKA-4898).
+
+ * PDF incremental-update scanning reads in blocks; ~4x less CPU (TIKA-4898).
* Metadata never stores an unpaired UTF-16 surrogate (TIKA-4897).
- * openai-embedding-engine gains "requestParameters" to enable custom
- parameters per vendor (TIKA-4896).
-
- * New catalog presets, inert until the config names them. (TIKA-4856).
-
- * unpack-config gains includeMetadata (TIKA-4881).
-
- * tika-server: /rmeta gains config/{handlerType} (TIKA-4881).
-
- * Tagged PDFs: text drawn twice for a bold effect with one copy tagged
- came out as letters with spaces between them; each copy now reads as
- its words. A tree that puts every glyph in its own paragraph (some form
- generators) is left to the text stripper under AUTO (TIKA-4891).
-
- * Inference results in the single-object outputs: with /tika as JSON,
- tika-app -j, or pipes CONCATENATE, the vectors of an attachment's pages
- and pictures now land on the one metadata object handed back, each
- with an embedded locator naming the part, instead of being discarded
- with the attachment's own metadata. InferenceUnit gains
- getDestination(), where a task writes; ChunkTarget.of(unit) replaces
- resolve(target, parent) (TIKA-4895).
-
- * New tika-eval-structure: a standalone tool that compares two sets of
- XHTML extracts block by block (paragraph order as Kendall tau, splits and
- merges, words one side glues that the other splits, artifact and untagged
- shares), per file and per producer, where token dice cannot see order
- (TIKA-4891).
-
- * Improve extraction of tagged PDFs(TIKA-4891).
-
- * PDF: a bookmark outline is now walked without recursion, and lists nest at
- most 50 deep with deeper items joining the deepest list (TIKA-4894).
+ * openai-embedding-engine gains "requestParameters" (TIKA-4896).
- * The image embedder writes a picture's vector onto the document the
- picture appears in (INLINE and RENDERING children) with an "embedded"
- locator naming the picture; attachments keep their own, and
- liftToParent: false keeps every picture's vector on the picture.
- RecursiveParserWrapper exposes the parent's Metadata to a child's parse
- as ParentMetadata, and ChunkTarget (tika-inference) resolves where a
- child's chunks land (TIKA-4888).
-
- * A PDF page rendered as an embedded document (imageStrategy RENDER_PAGES_*)
- is enriched once, by the PDF parser: the text recognizer where the OCR
- strategy says so, every other text-recognizers entry on every page,
results on
- the PDF itself. The embedded copy is no longer enriched again by
- image-parser, so an image embedder no longer yields two vectors for a
- text-poor page and Tesseract no longer OCRs a page twice; NO_OCR now means
- no page OCR at all. ContentEnrichers gains scoped suspend and
- suspendRecognizers for containers that enrich their own renderings
- (TIKA-4887).
-
- * OCR engines and other content enrichers: name the engine in
"text-recognizers"
- and every default parser stays loaded; "text-recognizers": [] turns
enrichment
- off everywhere. With no list, the loader resolves one engine
- per media type from the loaded parsers at startup (an engine configured
under
- "parsers" beats one default-parser found; last registered wins; an entry's
- _mime-include/_mime-exclude applies) and logs one line per effective
engine;
- a collision is a WARN at startup, as are two text recognizers sharing a
type
- (both run, their text is concatenated). An engine under "parsers" is a
parser
- only for the types no other parser there claims; when every type is
claimed,
- startup logs that it acts only as the enricher, or warns that it never
runs.
- The image/ocr-* pseudo-types
- are retired: Content-Type never carries the prefix (extracted images get
their
- real extension instead of .bin) and a _mime-include/_mime-exclude naming
one
- fails config load with the real type in the message. image-parser now owns
- image/jp2, image/jpx and image/x-portable-pixmap and calls the OCR engine
on
- them. VLM parsers gain "textRecognizer" (default true); set it false for a
- captioning prompt so AUTO OCR never replaces extracted text with the VLM's
- output (TIKA-4884).
-
- * For enricher authors: implement ContentEnricher (TextRecognizer for OCR)
and
- advertise the real image types; an enricher the classpath supplied is
never
- dispatched to, and one named under "parsers" only fills gaps. A
third-party
- engine still advertising image/ocr-* is treated as a legacy text
recognizer
- with a WARN until 5.0. An invoked enricher is recorded in tk:parsed-by
(images)
- and tk:parsed-by-full-set (PDF page renders) as a dispatched parser would
be.
- ContentEnrichers.resolve(Parser) is the load-time resolution.
- EmbeddedDocumentUtil.normalizeMediaType is deprecated (TIKA-4884).
-
- * The PDF parser's AUTO OCR strategy no longer emits a page's extracted text
- next to the OCR output that superseded it. Text that trips the AUTO
verdict
- is held back and replaced by OCR, and released unchanged when OCR cannot
run
- on that page: no engine on the classpath or in text-recognizers,
- maxPagesToOcr exhausted, or a timeout or error. AUTO without an engine now
- behaves as NO_OCR (TIKA-4883).
-
- * New TextRecognizer capability for content enrichers (tika-core): a parser
- implements it to say its output is the document's text rather than an
- annotation. The PDF parser's AUTO strategy supersedes extracted text only
- when a text recognizer is configured (Tesseract, Tess4J and the VLM
parsers
- declare it; an image-embedding enricher does not) and only when the engine
- actually wrote text. Other enrichers still run on the rendered page
- (TIKA-4883).
-
- * tika-app's -m and --json wrote metadata at the SAX endDocument event
- thereby dropping keys added by a parser after the endDocument (TIKA-4885).
-
- * "text-recognizers": _mime-exclude now matches the real media type, so
- excluding image/tiff also drops a legacy image/ocr-tiff engine (it was
- a no-op). The zero-media-types check at config load asks the engine
- itself, so a _mime-include list no longer masks an unreachable one.
- VLM parsers implement Closeable and release their HTTP client
- (TIKA-4884).
+ * tika-server: named configuration presets, and catalog presets that are
+ inert until the config names them (TIKA-4856).
+
+ * unpack-config gains includeMetadata; tika-server /rmeta gains
+ config/{handlerType} (TIKA-4881).
+
+ * Improve extraction of tagged PDFs. New tika-eval-structure tool compares
+ two sets of XHTML extracts block by block (TIKA-4891).
+
+ * Inference results in the single-object outputs (/tika as JSON, tika-app
+ -j, pipes CONCATENATE) land on the one metadata object handed back, with
+ an embedded locator, instead of being discarded (TIKA-4895).
* tika-inference and tika-vlm are experimental: classes, config keys and
- metadata output may change in minor releases without deprecation or
- compatibility shims (TIKA-4884). So are the "inference" config section,
- pdf-parser "inference", the org.apache.tika.parser.enricher, .inference
- and .hook packages, and the tk:chunks and tk:inference-released output:
- 4.2 batches recognition per document, meant as an opt-in addition (an
- engine that takes one image at a time keeps working unchanged), and adds
- media inputs and document-level tasks.
- Stable: the "engines" map, the "text-recognizers" list, and pdf-parser
- "text"; a 4.1 spelling of those loads as an alias for at least one minor
- release if it changes (TIKA-4895).
-
- * A missing OOXML relationship target no longer aborts the whole file:
- the threaded-comment and person lookups in xlsx and the page lookups in
- vsdx went straight to POI's getRelatedPart, whose unchecked
- IllegalArgumentException surfaced as "Error creating OOXML extractor" and
- dropped the text already extracted. They route through
- safeGetRelatedPart, as branch_3x already did (TIKA-4879).
-
- * RawTiffDetector rejects a BigTIFF directory offset near Long.MAX_VALUE
- instead of letting the bounds check overflow. Adding the entry-count
- size to such an offset wrapped negative and read as "already in the
- prefix", so a 16-byte file threw ArrayIndexOutOfBoundsException out of
- Detector.detect, which CompositeDetector does not catch: detection
- failed for the document and the remaining detectors never ran. Raw
- detection runs on every stream, so this was reachable from every entry
- point (TIKA-4861).
-
- * Entries of ODF, EPUB, GeoGebra, WACZ, XLZ and iWork containers, mbox
- messages and the AppleSingle data fork are re-opened from their
- container on rewind instead of cached: digesting rewinds every
- embedded document, and the cached copy cost heap for the whole entry
- and, past the cache budget or the 1 MB floor, a temp file. AppleSingle
- no longer spools its data fork to a temp file on every parse, and
- GeoGebra no longer spools every embedded picture to detect it
- (TIKA-4878).
-
- * A declared Content-Length is no longer treated as a measurement: the
- zip-bomb ratio counts only measured input bytes (a container-declared
- size on an embedded document could inflate its denominator), and a
- re-openable source no longer reserves cache budget or sizes its buffer
- from the declared length (a lying one could push a small payload to
- disk or churn the shared budget). Neither is in a release: the
- exposure arrived with TIKA-4868 and TIKA-4873 (TIKA-4878).
-
- * Embedded objects in Office documents are re-opened from their container
- instead of cached: every OOXML part (pictures, media, attachments), the
- OLE 2.0 package inside an OOXML part, the CONTENTS entry of an OLE 2.0
- object in a binary Office file, embedded objects in .ppt, XPS page
- images and Word EMF icons. Digesting an embedded document rewinds it;
- the cached copy that made possible cost heap for the whole object and,
- past the cache budget or the 1 MB floor, a temp file. The container
- hands the bytes back on demand, so neither is needed (TIKA-4878).
-
- * PDF attachments, PDF XMP packets, 3D on-instantiate scripts and PST
- attachments are re-opened from their document instead of cached when
- the embedded-document extractor rewinds them (digesting does, for every
- embedded document). The cached copy cost heap for the whole attachment
- and, past the cache budget or the 1 MB floor, a temp file; PDFBox and
- java-libpst hand the bytes back on demand (TIKA-4878).
-
- * tika-server: named configuration presets (TIKA-4856).
+ output may change in minor releases. Stable: the "engines" map, the
+ "text-recognizers" list and pdf-parser "text" (TIKA-4884, TIKA-4895).
+
+ * PDF bookmark outlines are walked without recursion; lists nest at most
+ 50 deep (TIKA-4894).
+
+ * The image embedder writes a picture's vector onto the document the
+ picture appears in, with an "embedded" locator; liftToParent: false
+ keeps it on the picture (TIKA-4888).
+
+ * A PDF page rendered as an embedded document is enriched once, by the
+ PDF parser; NO_OCR now means no page OCR at all (TIKA-4887).
+
+ * OCR engines are named in "text-recognizers" ("[]" turns enrichment off;
+ no list resolves one engine per media type at startup). The image/ocr-*
+ pseudo-types are retired; a third-party engine still advertising them is
+ a legacy text recognizer with a WARN until 5.0. VLM parsers gain
+ "textRecognizer" (TIKA-4884).
+
+ * PDF AUTO OCR no longer emits a page's extracted text next to the OCR
+ output that superseded it; AUTO without a text recognizer behaves as
+ NO_OCR. New TextRecognizer capability for content enrichers (TIKA-4883).
+
+ * tika-app -m and --json no longer drop metadata keys added after
+ endDocument (TIKA-4885).
+
+ * A missing OOXML relationship target (xlsx threaded comments, vsdx pages)
+ no longer aborts the whole file (TIKA-4879).
+
+ * Raw camera formats (NEF/NRW, PEF/PTX, ARW/SRF/SR2, SRW, DNG, RAF, RW2,
+ MRW, ORF) are detected by content instead of as image/tiff. Fix a
+ BigTIFF offset overflow in RawTiffDetector that failed detection for the
+ whole document (TIKA-4861).
+
+ * Embedded documents in Office files, PDFs, PST and the zip-family
+ containers are re-opened from their container on rewind instead of
+ cached (TIKA-4878).
* Temp files follow -Djava.io.tmpdir on the parent JVM (Tika, its
- libraries, and forks all honor it); TikaLoader fails at config load
- if it is unusable. pipes.tempDirectory is deprecated for removal in
- 5.0: it only covered the forks. Do not use tmpfs: spool size is
- bounded by input, and an orphaned fork dir pins RAM (TIKA-4877).
-
- * A fork whose parent dies deletes its own temp dir; the parent
- surfaces a fork's hs_err log before every delete. Failure-path temp
- file leaks fixed in PDFBoxRenderer, PopplerRenderer, truncated RTF,
- and MarianTranslator (TIKA-4877).
-
- * tika-eval Profile/Compare speedups: single-pass URL/mail stripping
- replaces the bounded regexes in langdetect preprocessing (same output,
- 17-290x faster on web text), the default H2 db URL sizes the page cache
- at a quarter of the heap clamped to [64MB, 1GB] (override with
- -Dtika.eval.h2.cacheSizeKb=<kb>), and the status log adds a last-interval
- docs-per-sec rate next to the cumulative average (TIKA-4875).
-
- * New "text-recognizers" config list (TIKA-4872): select the OCR engine
- ("tesseract-ocr-parser", "tess4j-parser", "openai-vlm-parser", ...) by
- name instead of by classpath registration of the image/ocr-* pseudo
- media types. Enrichers advertise real media types (legacy engines that
- still advertise image/ocr-* are mapped to the real type, so all are
- nameable) and are invoked by the image and PDF parsers rather than
- dispatched to by the composite, so an enricher no longer displaces the
- parser registered for the same type. Enricher selection uses the
- detected media type, captured before a parser can refine Content-Type.
- Every enricher matching a media type runs, in config order (e.g. an
- OCR engine then a VLM tagger for the same image), best-effort: one
- enricher's failure does not stop the others and is still reported;
- timeouts abort the chain. The list is authoritative: a media type no
- configured enricher matches gets no enrichment -- never a classpath
- engine that was not named -- and a named engine that reports no media
- types at load (missing binary, unreachable inference server) fails
- config load instead of going silently inert. With no
- "text-recognizers" configured, the legacy ocr-* dispatch applies
- unchanged; a WARN at config load now names colliding OCR engines and
- the winner. TesseractOCRParser's
- component name is pinned as "tesseract-ocr-parser".
-
- * Inference/OCR hardening (TIKA-4871): OpenAIVLMParser no longer
- auto-registers via SPI, matching its Claude/Gemini siblings; select
- it by name ("openai-vlm-parser") in config. Per-request parse-context
- config for the embedding filters now works and is validated:
- {"openai-embedding-filter": {"skipEmbedding": true}} (likewise
- "jina-embedding-filter") merges over the server config, and
- baseUrl/apiKey/model may not be changed at runtime. The embedding
- filters release their HTTP client resources on close(). Inline PDF
- page OCR now accumulates tk:chunks from every OCR'd page onto the
- parent document instead of keeping only the first page's.
-
- * Placeholder streams -- the empty stand-ins parsers hand parseEmbedded
- for content that is never extracted -- report an unknown length rather
- than their own zero, and the macro-failure entry is registered without
- parsing its sentinel (TIKA-4874).
-
- * TikaInputStream.hasReliableLength() distinguishes measured lengths
- from declared Content-Length hints, and one-shot streams now carry a
- declared length without spooling; detection sizes its magic read only
- from measured lengths, so a lying declared length can no longer
- truncate it (TIKA-4868).
-
- * ParseContext entries holding per-parse runtime state (ParseRecord,
- ParseTimeout, the parser-map cache) are skipped during serialization
- instead of failing as unregistered components (TIKA-4868).
-
- * Mojibuster's adaptive probe strips incrementally instead of
- re-stripping the whole buffer on every read (quadratic on tag-heavy
- pages); JunkDetector's Unicode block lookup uses a precomputed BMP
- table. Output unchanged (TIKA-4868).
-
- * Markdown rendering is another ~8x faster on large documents: a custom
- Text-node renderer emits unescaped spans in bulk instead of
- commonmark's per-character escape-check-and-append. Byte-identical
- output, guarded by a fast-vs-stock differential test (TIKA-4868).
-
- * ZipParser no longer re-decompresses an entry on every rewind when the
- entry uses a legacy compression method (implode, shrink, bzip2, ...):
- such entries replay from the budgeted cache instead of re-opening.
- An imploded 606KB zip drops from 610ms to 102ms (TIKA-4868).
-
- * More detection/dispatch savings: the message/rfc822 priority-45 magic
- is gated behind a lossless ':' scan of the first 30 bytes;
- CompositeParser caches the built type->parser map in the ParseContext
- so embedded documents reuse the container's map; MagicDetector
- precomputes a per-pattern first-byte table. Adds a RESOURCE_TIMING
- log, silenced by default in the shipped log4j2 configs (TIKA-4868).
-
- * WordExtractor (.doc) cleans each character run and tests paragraph
- blankness in single passes; ToMarkdownContentHandler collapses line
- breaks copy-free for clean runs. ~39% off a text-heavy 2MB .doc
+ libraries and forks); pipes.tempDirectory is deprecated. A fork whose
+ parent dies deletes its own temp dir; failure-path temp leaks fixed
+ (TIKA-4877).
+
+ * tika-eval Profile/Compare speedups: langdetect preprocessing, H2 page
+ cache sizing, a per-interval rate in the status log (TIKA-4875).
+
+ * New "text-recognizers" config list selects OCR and enrichment engines by
+ name; enrichers are invoked by the image and PDF parsers rather than
+ dispatched to by the composite (TIKA-4872).
+
+ * OpenAIVLMParser no longer auto-registers via SPI; per-request config for
+ the embedding filters works and locks baseUrl/apiKey/model; inline PDF
+ page OCR accumulates tk:chunks from every page (TIKA-4871).
+
+ * Performance: detection (magic ~35% faster on unmatched input, override
+ keys honored before magic, cached type->parser map), markdown output
+ ~30x faster, .doc cleanup, CSV sniffing, zip legacy-method rewinds, pipes
+ ACK overlap, raw UTF-8 content passback for tika-server
+ (content-bytes-config). Compat: with a Content-Type override set,
+ DefaultDetector no longer lets a more specific magic result overrule it
(TIKA-4868).
- * tika-server's raw-output endpoints (/tika, /tika/text, ...) carry the
- extracted content as raw UTF-8 bytes from the pipes worker to the HTTP
- response instead of a Smile-encoded string (9MB text: 115ms -> 75ms).
- Opt-in via the new content-bytes-config parse-context component, which
- moves CONTENT_ONLY passback content out of tk:content into
- EmitData.getContentBytes(); results routed to a regular Emitter get
- the content restored to the metadata (TIKA-4868).
-
- * Detection hot-path cleanups: MagicMatch resolves its detector via
- double-checked locking; glob patterns are compiled once at
- registration; MimeTypes.forName reads a ConcurrentHashMap (fixing an
- unsynchronized-read race) and indexes normalized keys; resource names
- containing spaces skip the URI-parse-by-exception; the magic-header
- buffer is sized by the stream's measured length instead of a fixed
- 64KB; the Adobe Illustrator ranged regex is gated behind a literal
- scan. Detection results unchanged (TIKA-4868).
-
- * Magic detection is ~35% faster on unmatched (e.g. plain-text) input:
- range scans find first-byte candidates before running the full
- masked/case-folded compare (TIKA-4868).
-
- * CSVSniffer reads its detection window once into a shared buffer and
- runs every delimiter hypothesis against it; windows with no delimiter
- and no quote skip the scan outright. Results unchanged (TIKA-4868).
-
- * Pipes workers no longer stall between pre-parse and parse waiting
- for the client to acknowledge the intermediate-result frame; the
- ACK round trip now overlaps the parse. Adds per-request timing logs
- on org.apache.tika.pipes.timing.*, silenced by default in the shipped
- log4j2 configs; raise that logger to info to enable (TIKA-4868).
-
- * DefaultDetector honors CONTENT_TYPE_USER_OVERRIDE and
- CONTENT_TYPE_PARSER_OVERRIDE before running magic detection, matching
- CompositeDetector's contract. Removes the second full magic scan every
- pipes parse paid per document. Compat note: with either override set,
- DefaultDetector no longer lets a more specific magic result overrule
- it; parts whose parser declares a type from container headers (e.g.
- inline text/* mail parts) now report the declared type, and
- CONTENT_TYPE_MAGIC_DETECTED is not recorded when an override short
- circuits detection (TIKA-4868).
-
- * Markdown output is ~4x faster on large documents:
- ToMarkdownContentHandler now buffers the commonmark renderer's
- per-character writes instead of paying the synchronized
- Writer.write(int) cost for every character (TIKA-4868).
- * Embedded documents carry their size: ParsingEmbeddedDocumentExtractor
- sets Content-Length from the stream where the stream knows it and the
- parser did not say, which never spools to measure one, and the raw
- camera previews and the audio cover art set the length they read from
- the file (TIKA-4873).
-
- * AVIF images are parsed rather than only detected: HeifParser accepts
- image/avif, which is the same ISO-BMFF container, so dimensions, EXIF
- and XMP come out of it the way they do for HEIC (TIKA-4870).
-
- * The video of a Google/Android motion photo, appended after the image and
- described by the Motion Photo or MicroVideo XMP, is emitted as an
- ATTACHMENT embedded document named after what the file declares it to
- be. Nothing is emitted, and nothing is recorded, when the declared
- length does not fit the file or the bytes there are not recognized,
- which is what sharing a motion photo out of a gallery leaves behind
- (TIKA-4869).
-
- * Raster previews for the vector thumbnails of Office documents: the new
- poi-metafile-renderer draws EMF and WMF images through POI (a PNG of
- a configurable width; Word's bitmap-in-WMF thumbnails from the bitmap
- directly), EMFParser and WMFParser are RenderingParsers that emit the
- rendering as a RENDERING embedded document with "emf-parser" /
- "wmf-parser": {"renderImage": true, "renderWidth": 800}, off by
- default and restrictable to e.g. THUMBNAIL embedded documents with
- "renderOnlyEmbeddedResourceTypes", and OfficeParser emits the
- SummaryInformation thumbnail of the OLE2 formats (a WMF) as a THUMBNAIL
- embedded document, as the OOXML parsers do with the docProps thumbnail,
- switchable with "office-parser": {"extractThumbnail": false}
- (TIKA-4855).
-
- * Add "exception-reporting" parse-context config to redact and bound
+ * Embedded documents carry Content-Length where the stream knows it
+ (TIKA-4873).
+
+ * AVIF images are parsed by HeifParser (dimensions, EXIF, XMP) (TIKA-4870).
+
+ * The video of a Google/Android motion photo is emitted as an ATTACHMENT
+ embedded document (TIKA-4869).
+
+ * New poi-metafile-renderer rasterizes EMF/WMF (opt-in "renderImage");
+ OfficeParser emits the OLE2 SummaryInformation thumbnail as a THUMBNAIL
+ embedded document ("extractThumbnail": false disables) (TIKA-4855).
+
+ * New "exception-reporting" parse-context config redacts and bounds
exception text in metadata, tika-server error bodies and pipes/grpc
- messages; FileSystemEmitter writes atomically (TIKA-4848).
- Compat notes: a truncated TSD envelope now records its read failure
- under tk:exception:embedded-stream-exception rather than
- tk:exception:embedded-exception; recordException and
- recordEmbeddedStreamException no longer strip a bare TikaException
- wrapper, so the first line of tk:exception:* values may name the
- wrapper (affects consumers keyed on that line, e.g. eval cause
- counts across the 4.1 boundary).
-
- * Audio cover art is emitted as a THUMBNAIL embedded document, like the
- preview image of the document container formats: the front cover (ID3
- APIC and FLAC/Vorbis picture type 3), else the first picture of type
- "Other" or unknown, else the first picture, and the first covr image
- of an MP4. Further pictures
- stay INLINE. Clients that looked for cover art as INLINE need to
- accept THUMBNAIL as well (TIKA-4850).
-
- * tika-grpc resolves its plugin-roots fallback against the install
- layout via DefaultPluginsDir instead of a working-directory-relative
- pf4j default, and a WARN names the resolved directory when no plugins
- directory exists (TIKA-4865).
-
- * The tika-server full and tika-grpc Docker images install fonts-noto-cjk:
- without any CJK face, PDFs using non-embedded CJK fonts render (and OCR)
- as .notdef boxes in every renderer, even though the images ship Japanese
- tesseract data (TIKA-4866).
-
- * The default plugins directory is resolved against the install layout
- (next to the jar, or next to its lib/ directory) and always as an
- absolute path, shared by tika-server, PipesForkParser and the async
- CLI; it no longer depends on the working directory (TIKA-4864).
-
- * tika-server error bodies (the 422/500 exception mapper, /meta/{field})
- now honor the exception-reporting policy; /meta/{field} returns the
- already-formatted container exception instead of re-wrapping it with
- server frames. Completes TIKA-4848 (TIKA-4848).
-
- * The exception-reporting policy now also governs the messages a pipes
- worker returns (fetch/emit/crash) and the container exception it
- records; part of TIKA-4848 step 3 (TIKA-4848).
-
+ messages; FileSystemEmitter writes atomically. Compat: tk:exception:*
+ values may now begin with the TikaException wrapper line (TIKA-4848).
+
+ * Audio cover art (ID3/FLAC/Vorbis front cover, MP4 covr) is emitted as a
+ THUMBNAIL embedded document rather than INLINE (TIKA-4850).
+
+ * The default plugins directory (tika-server, PipesForkParser, async CLI)
+ and tika-grpc's plugin-roots fallback are resolved against the install
+ layout instead of the working directory (TIKA-4864, TIKA-4865).
+
* Allow image compression settings in PDFBox-based renderer (TIKA-4862).
- * The tika-server full and tika-grpc Docker images set OMP_THREAD_LIMIT=1:
- to avoid oversubscribing the CPU under forked parse workers (TIKA-4863).
-
- * embedded-limits maxDepth counts embedding levels again instead of the
- parsers a parse passes through; with AutoDetectParser over DefaultParser
- every value above 1 used to stop one level early (TIKA-4857).
-
- * GeoGebraParser emits the icon of a tool (*.ggt, the macro's iconFile)
- as its THUMBNAIL embedded document; tool files have no thumbnail of
- their own (TIKA-4831).
-
- * Enum values in JSON configuration are matched case-insensitively, so
- "no_ocr" works as well as "NO_OCR"; the server docs used the lower-case
- form in their examples (TIKA-4859).
-
- * The preview image of iWork '09 packages (QuickLook/Thumbnail.jpg) and of
- iWork '18 packages (preview.jpg) is emitted as a THUMBNAIL embedded
- document, as it already was for iWork '13 (TIKA-4854).
-
- * Raw camera formats are detected by content: RawTiffDetector tells
- Nikon NEF/NRW, Pentax PEF/PTX, Sony ARW/SRF/SR2, Samsung SRW and Adobe
- DNG from a plain TIFF by their image directory (DNGVersion, the vendor
- Compression codes, or a CFA/LinearRaw image plus Make), and Fuji RAF,
- Panasonic RW2, Minolta MRW and the remaining Olympus ORF byte orders
- get magic entries. Streams without a file name used to be image/tiff.
- image/x-raw-samsung (*.srw) is new and parsed by RawTiffParser
- (TIKA-4861).
-
- * DWGReadParser emits the drawing's THUMBNAILIMAGE as a THUMBNAIL embedded
- document instead of INLINE (TIKA-4853).
-
- * EpubParser emits the cover image named by the OPF (the EPUB 3
- cover-image manifest property, or the EPUB 2 cover meta) as a THUMBNAIL
- embedded document (TIKA-4852).
-
- * RawTiffParser marks only the largest embedded JPEG preview as the
- THUMBNAIL embedded document; the smaller previews of the same image are
- INLINE images named image-N.jpg. Previously every preview was a
- THUMBNAIL, so a client had to compare them to find the representative
- one (TIKA-4851).
-
- * tika-eval: Profile/Compare accept the batch run's jsonl crash ledger
- (--pipesReport, -pa/-pb) and a run-info json (--runInfo, -ra/-rb), and
- read both from <extracts>/.run-info/ by default (refusing an ambiguous
- dir). containers gains pipes_status/pipes_message; a new run_info table
- records eval and batch provenance; reports and
- summary.md classify NO_EXTRACT_FILE by ledger status (CRASH, the raw
- status, NO_PIPES_RECORD, BATCH_WITHOUT_LEDGER, NO_PIPES_REPORT_SUPPLIED).
- Report on a db from an earlier tika-eval skips the reports it cannot run
- instead of aborting (TIKA-4847).
-
- * New file-system-jsonl-reporter pipes reporter (TIKA-4846).
-
+ * The tika-server full and tika-grpc Docker images add ffmpeg and
+ fonts-noto-cjk, and set OMP_THREAD_LIMIT=1 to avoid oversubscribing the
+ CPU under forked parse workers (TIKA-4910, TIKA-4866, TIKA-4863).
+
+ * embedded-limits maxDepth counts embedding levels again; values above 1
+ used to stop one level early (TIKA-4857).
+
+ * New GeoGebraParser for *.ggb/*.ggs/*.ggt with content-based detection
+ and THUMBNAIL embedded documents (TIKA-4831).
+
+ * Enum values in JSON configuration are matched case-insensitively
+ (TIKA-4859).
+
+ * iWork '09 and '18 preview images are emitted as THUMBNAIL embedded
+ documents (TIKA-4854).
+
+ * DWGReadParser emits THUMBNAILIMAGE as THUMBNAIL instead of INLINE
+ (TIKA-4853).
+
+ * EpubParser emits the OPF cover image as a THUMBNAIL embedded document
+ (TIKA-4852).
+
+ * RawTiffParser marks only the largest JPEG preview as THUMBNAIL; smaller
+ previews are INLINE (TIKA-4851).
+
+ * New file-system-jsonl-reporter pipes reporter records a batch run's
+ crashes; tika-eval Profile/Compare read that ledger and the run-info
+ json to classify NO_EXTRACT_FILE by cause (TIKA-4846, TIKA-4847).
+
* Stop spooling OLE2 objects whose header over-reserves BAT capacity
(TIKA-4845).
* Add Micrometer reporting and opt-in endpoint for tika-server (TIKA-4839).
-
- * Improve spooling/decrease number of spills to disk (TIKA-4835).
-
- * Fixed a bug that made per-request (parse-context) configuration unusable
- for parsers that lock some config fields against caller modification --
- Tess4J, the VLM parsers and the OpenAI image-embedding parser. Any such
- config threw, including an empty one: the defaults were deep-copied
- through their own setters, which the runtime config overrides to reject
- caller input, so the copy tripped the parser's own guards before the
- caller's JSON was read. Locked fields are still rejected when a caller
- actually sets them. Configuration supplied at initialization time (the
- "parsers" section) was never affected (TIKA-4843).
-
- * OOXML parsers flag package parts that are unreachable through the OPC
- relationship graph: msoffice:has-unreferenced-parts (boolean) and
- msoffice:unreferenced-part-names. Purely structural (no bytes are
- inspected; content types come from [Content_Types].xml by extension), so
- expect false positives from tools that leave orphan parts behind. A hiding
- place a raw-ZIP scanner can still see, not a statement about what Tika
- parsed. Applies to Word, Excel, PowerPoint and Visio OOXML (including
- macro-enabled variants); XPS links content by markup rather than
- relationships and is not checked (TIKA-4837).
-
- * Shared pipes server (useSharedServer: true, not the default): a client
whose
- in-flight parse was killed by another client's restart could restart the
- healthy replacement. ensureRunning holds its lock across the whole fork,
so
- siblings cannot report a dead worker until after the replacement is up,
and
- the pending-restart flag carried no process identity -- so a report about
- the process that just died was applied to its successor, which was then
- destroyed and re-forked. One worker death produced two restarts and a
second
- round of destroyed in-flight work; under sustained concurrent load it
- sustained itself at one spurious restart per round, appearing as periodic
- unexplained worker churn and intermittent parse failures that succeed on
- retry. Each fork now carries a generation that clients capture when they
- connect and hand back with every report, and reports about a superseded
- process are dropped. Also fixed in shared mode: ensureRunning could fork a
- replacement after shutdown() that nothing owned and nothing would ever
- destroy, and an interrupt during process teardown left the process handle
- pointing at a killed process and leaked the temp directory. Affects 4.0.0
- and earlier (TIKA-4844).
-
- * tika-pipes: the cache memory budget (how much rewindable content a forked
- worker keeps in memory before spilling to disk; new since 4.0.0, which had
- no budget at all) defaults to a quarter of the fork's heap, so raising
- -Xmx raises it. It is one pool per forked JVM shared by all of its
threads.
- -Dtika.pipes.cacheMemoryBudgetBytes in forkedJvmArgs overrides it (below
- the quarter-heap ceiling; <=0 disables); the fork logs the value and its
- source at startup. TikaInputStream.hasFile() now also reports content the
- stream cache spilled on its own, not only content a getPath() call put on
- disk; note getPath() may still have to drain the rest of the source into
- that file. TikaInputStream.toString() no longer forces a spill, so logging
- or debugger-inspecting a stream is side-effect-free.
- TikaInputStream.inMemoryContent(channel) gives a zero-copy read-only view
- of cached content for consumers that need random access. Digester
- gains digestSink(), a DigestSink that digests as it is written; nothing is
- written to the metadata unless the producer calls commit(), so any failure
- -- exception, Error, or a producer that closes the sink itself --
publishes
- no digest rather than a digest of the bytes that happened to arrive. A
- translator that claims a stream and writes nothing likewise publishes
- nothing: embedded PST mail items, whose translator is still a stub, no
- longer carry the digest of zero bytes (the same value for every one of
- them) and instead carry no digest at all. DigestHelper uses it for
- translated embedded streams, which no longer touch a temp file when the
- digester implements digestSink (all of Tika's do; one that only implements
- digest() still buffers).
- TemporaryResources.closeAll(Closeable...) closes every argument even when
- one throws unchecked; TemporaryResources, CachingSource,
CachingInputStream
- and CompositeDigester use it (TIKA-4835).
-
- * Documentation: corrected a batch of pages and javadocs that contradicted
- the code. Notably: the ES/OpenSearch attachmentStrategy has no default
- (unset means embedded documents get neither the parent field nor the
- parent/child relation); Kafka's connectionsMaxIdleMs is passed to the
- producer, not ignored; jdbc queryTimeoutSeconds is applied only when > 0,
- so 0 does not mean "no limit"; the Solr emitter/iterator support only
- basic auth, not ntlm, and only when a userName is set; pipes-reporters
- silently loads zero reporters when given a JSON array, and
- pipes-iterator/pipes-reporters instances are built at config load rather
- than lazily; under CONTENT_ONLY only a parse-context filter replaces the
- built-in one, not the top-level metadata-filters chain;
- _mime-include/_mime-exclude also accept a bare string; Tess4J locks
- poolSize and maxImagePixels as well as the two paths; and pdf:trapped and
- xmp:pdf:Trapped are new 4.x keys rather than renames (3.x captured the
- flag only as pdf:docinfo:trapped and dropped the XMP value). Also
- corrected the config nesting shown in every pipes-plugin fetcher/emitter
- javadoc -- 23 of them, which had it inverted (the instance id is the
- outer key, the component name the inner) -- and removed references to a
- TesseractOCRConfig.properties file that 4.x does not load (TIKA-4842).
-
- * Pipes plugins no longer bundle their own Jackson: jackson-core, -databind
- and -annotations are provided by the host (tika-serialization) and the
- plugins parent pom now bans bundling them, so a mapper can cross the
- plugin boundary without a second copy of the Jackson classes (seven plugin
- zips shipped one). Plugin configuration JSON is parsed by one shared
- mapper, PluginJson (tika-plugins-core), which rejects unknown keys,
- numbers for enums and duplicate keys, and accepts
- // and /* */ comments; the 33 per-plugin *Config classes use it instead of
- their own bare ObjectMapper (TIKA-4840).
-
- * tika-server and tika-async-cli now start from a config that contains
- // or /* */ comments, as the configuration docs have always said they
- may. The main loader accepted them; the steps that re-read the user's
- file to merge in server/CLI overrides (ConfigMerger, ensurePluginRoots)
- used their own bare parser and refused the whole file; they now use the
- shared TikaObjectMapperFactory mapper (TIKA-4834).
-
- * The Kafka pipes iterator no longer stops at the first empty poll. A newly
- subscribed consumer spends its first poll(s) joining the group and returns
- empty even when the topic has a backlog, so the iterator could enqueue
zero
- files and report success. It now waits for a partition assignment (bounded
- by the new assignmentTimeoutMs, default 30s) and requires a continuous
quiet
- window (drainIdleMs, default 1s) before concluding the topic is drained.
- groupInitialRebalanceDelayMs is deprecated and no longer sent to the
- consumer: it is a broker setting that Kafka has always ignored
(TIKA-4833).
-
- * Pipes IPC: carry inline document bytes as a raw binary field beside the
- tuple in the request envelope -- never inside the tuple or its
- ParseContext -- and disable Smile's 7-bit binary encoding. Tuple JSON
- serialized by 4.0.0 with an "inline-bytes" parse-context entry no longer
- loads; it is rejected with a tailored message (TIKA-4829).
-
- * Digesting embedded documents no longer buffers each embedded object to a
- temp file. Zip entries are re-read from the parent archive on rewind, and
- a new process-wide CacheMemoryBudget (seeded by the pipes forked server;
- default 256MB, clamped to a quarter of the fork's heap; tunable via
- -Dtika.pipes.cacheMemoryBudgetBytes in the config's forkedJvmArgs, <=0
- disables) lets embedded objects stay in memory past the per-object 1MB
- threshold. New public API on TikaInputStream: get(IOSupplier,...),
- enableRewind(CacheMemoryBudget), getSeekableByteChannel(). Zip/7z/epub/odf
- parsing and zip container detection now read through seekable channels, so
- after detection/parsing a TikaInputStream may no longer be file-backed
- (hasFile() false); getPath()/getFile() still work and spool on demand
- (TIKA-4828).
-
- * Pipes now carries the caller-supplied Content-Type across the worker's
- fresh-metadata boundary as a soft detection hint, so every forked-parse
- endpoint (/tika, /meta, /rmeta, /unpack, /async, /pipes, plus tika-grpc
- and embedded PipesForkParser) can route on a client Content-Type, not
- only on the filename. Detection keeps the hint only when it equals or
- specializes the content-detected type (e.g. refining image/tiff to
- image/x-canon-cr2); for bytes with no magic it can select any type,
- matching the routing power the filename already had. The
- CONTENT_TYPE_USER_OVERRIDE key is deliberately not carried, so the hint
- cannot force an unrelated type (TIKA-4825).
-
- * OneNote extraction now follows document order, omits superseded page
- revisions, sorts author metadata, extracts embedded object BLOBs, and
- bounds malformed-input recursion and file-derived allocations. Parse
- warnings and embedded relationship IDs are exposed in metadata. Malformed
- or truncated files that cannot be fully parsed, and files whose walk
- yields no content, now fall back to the legacy string dump instead of
- failing or returning empty output. The legacy MS-ONESTORE walker bounds
- its recursion (depth caps plus file-node-list and fragment-chain cycle
- guards) and now honors shouldParseEmbedded for embedded file data
- * PDF: extractFontNames threw NullPointerException on a page with no
- /Resources dictionary (TIKA-4842).
-
- * tika-server: opt-in Micrometer metrics reporting and endpoint
- (TIKA-4839).
-
- * Per-request (parse-context) config for parsers that lock fields
- (Tess4J, VLM, OpenAI image-embedding) threw even when empty; locked
- fields are still rejected when actually set (TIKA-4843).
+
+ * Improve spooling/decrease number of spills to disk. The pipes cache
+ memory budget defaults to a quarter of the fork heap
+ (-Dtika.pipes.cacheMemoryBudgetBytes overrides) (TIKA-4835).
+
+ * Per-request config for parsers that lock fields (Tess4J, VLM, OpenAI
+ image-embedding) threw even when empty (TIKA-4843).
* OOXML: new msoffice:has-unreferenced-parts and
- msoffice:unreferenced-part-names flag package parts unreachable via the
- OPC relationship graph. Structural only, expect false positives; not
- applied to XPS (TIKA-4837).
-
- * Shared pipes server (useSharedServer: true): a worker death could trigger
- a second, spurious restart that killed the healthy replacement. Forks now
- carry a generation; stale reports are dropped. Also fixed: a fork after
- shutdown() that was never destroyed, and a temp-dir leak on interrupt
- during teardown (TIKA-4844).
-
- * tika-pipes cache memory budget defaults to a quarter of the fork heap;
- override with -Dtika.pipes.cacheMemoryBudgetBytes in forkedJvmArgs
- (<=0 disables). TikaInputStream: hasFile() also reports cache spills,
- toString() no longer spills, new inMemoryContent(channel). Digester gains
- digestSink(); a digest is published only on commit(), so failed or empty
- translations (e.g. stub PST items) publish no digest. New
- TemporaryResources.closeAll(Closeable...) (TIKA-4835).
-
- * Docs/javadocs reconciled with the code: ES/OpenSearch attachmentStrategy
- has no default; Kafka connectionsMaxIdleMs is honored; jdbc
- queryTimeoutSeconds 0 is not "no limit"; Solr basic auth only; pipes
- reporters/iterators are built at config load; Tess4J also locks poolSize
- and maxImagePixels; pdf:trapped is new, not renamed; plugin config
- nesting fixed in 23 javadocs (TIKA-4842).
-
- * Pipes plugins no longer bundle Jackson; the host provides it. Plugin
- config is parsed by a shared strict PluginJson mapper (rejects unknown
- and duplicate keys; accepts comments) (TIKA-4840).
+ msoffice:unreferenced-part-names (structural only; not XPS) (TIKA-4837).
+
+ * Shared pipes server: a worker death could trigger a spurious second
+ restart; forks now carry a generation (TIKA-4844).
+
+ * Docs and javadocs reconciled with the code; fixes an
+ extractFontNames NullPointerException on a PDF page with no /Resources
+ (TIKA-4842).
+
+ * Pipes plugins no longer bundle Jackson; plugin config is parsed by a
+ shared strict mapper that rejects unknown and duplicate keys (TIKA-4840).
* tika-server and tika-async-cli accept // and /* */ comments in config
- during override merging, as documented (TIKA-4834).
+ during override merging (TIKA-4834).
- * Kafka pipes iterator no longer stops on the first empty poll; waits for
- partition assignment (assignmentTimeoutMs, 30s) and a quiet window
- (drainIdleMs, 1s). groupInitialRebalanceDelayMs is deprecated
- (TIKA-4833).
+ * Kafka pipes iterator no longer stops on the first empty poll
+ (assignmentTimeoutMs, drainIdleMs); groupInitialRebalanceDelayMs is
+ deprecated (TIKA-4833).
* Pipes IPC carries inline bytes as a raw binary field, not in the tuple;
- Smile 7-bit binary encoding disabled. 4.0.0 tuples with an "inline-bytes"
- parse-context entry are rejected (TIKA-4829).
+ 4.0.0 tuples with an "inline-bytes" parse-context entry are rejected
+ (TIKA-4829).
* Digesting embedded documents no longer spools each to a temp file; a
- process-wide CacheMemoryBudget (default 256MB) keeps them in memory. New
- TikaInputStream API: get(IOSupplier,...), enableRewind(CacheMemoryBudget),
- getSeekableByteChannel(). Zip-family parsing and detection use seekable
- channels, so hasFile() may be false afterward; getPath() still spools on
- demand (TIKA-4828).
+ process-wide CacheMemoryBudget keeps them in memory. Zip-family parsing
+ uses seekable channels, so hasFile() may be false afterward (TIKA-4828).
* Pipes carries the client Content-Type into the forked worker as a
- detection hint for all forked endpoints; honored only when it equals or
- specializes the detected type, or when there is no magic. The
- user-override key is not carried (TIKA-4825).
-
- * OneNote: document-order extraction, superseded revisions omitted, embedded
- BLOBs extracted, warnings and relationship IDs in metadata, bounded
- recursion/allocation; malformed files fall back to the legacy string dump
- (TIKA-4814).
-
- * New GeoGebraParser for *.ggb/*.ggs/*.ggt: geogebra:* metadata, text and
- the thumbnail as a THUMBNAIL embedded document, with content-based
- detection. Previously typed application/zip with every entry as an
- attachment. *.ggs and *.ggp are new mime types; *.ggp is glob-only
- (TIKA-4831).
-
- * RawTiffParser extracts the camera-generated JPEG previews embedded in
- TIFF-based raw images (Nikon NEF/NRW, Sony ARW/SRF/SR2, Pentax PEF/PTX,
- Adobe DNG and Canon CR2, including BigTIFF DNG containers) as thumbnail
- embedded documents. image/x-raw-{nikon,sony,pentax,adobe} are now
- sub-classes of image/tiff, so a named NEF/ARW/PEF/DNG that used to detect
- as image/tiff (TiffParser, metadata only) now detects as image/x-raw-* and
- emits thumbnail-N.jpg attachments in /rmeta and /unpack; CR2 keeps its
- detection but also gains the attachments. Disable via
- "raw-tiff-parser": {"extractPreviews": false} (TIKA-4824).
- * RawTiffParser extracts embedded JPEG previews from NEF/NRW, ARW/SRF/SR2,
- PEF/PTX, DNG and CR2 as thumbnail embedded documents. image/x-raw-* are
- now subtypes of image/tiff, so named raw files detect as image/x-raw-*.
- Disable with "raw-tiff-parser": {"extractPreviews": false} (TIKA-4824).
+ detection hint for every forked endpoint (TIKA-4825).
+
+ * OneNote: document-order extraction, superseded revisions omitted,
+ embedded BLOBs extracted, bounded recursion; malformed files fall back
+ to the legacy string dump (TIKA-4814).
+
+ * RawTiffParser extracts embedded JPEG previews as THUMBNAIL embedded
+ documents; image/x-raw-* are now subtypes of image/tiff
+ ("extractPreviews": false disables) (TIKA-4824).
Release 4.0.0 - 8/18/2026