This is an automated email from the ASF dual-hosted git repository.
tballison pushed a commit to branch main
in repository https://gitbox.apache.org/repos/asf/tika.git
The following commit(s) were added to refs/heads/main by this push:
new c705468dbe TIKA-4936: Extract icons from PE executables (EXE/DLL) as
embedded documents (#3263)
c705468dbe is described below
commit c705468dbed174861281ef5c5fe6cb8b9835241e
Author: Dominik Schmidt <[email protected]>
AuthorDate: Tue Oct 6 23:33:29 2026 +0200
TIKA-4936: Extract icons from PE executables (EXE/DLL) as embedded
documents (#3263)
* TIKA-4936: Extract icons from PE executables (EXE/DLL) as embedded
documents
ExecutableParser so far only read the COFF header of PE files. It now
continues through the optional header and section table to the resource
section and rebuilds every RT_GROUP_ICON together with its RT_ICON images
into a standalone .ico file, which is handed to the
EmbeddedDocumentExtractor as image/vnd.microsoft.icon. The first group in
resource order (the icon Windows shows for the file) is marked THUMBNAIL,
the others ATTACHMENT; the embedded relationship id carries
type/name/language of the resource.
The extractor reads strictly forward, bounds-checks every offset, caps the
tree depth, entry count and section size, detects cycles and treats a
truncated resource section as "fewer icons" rather than an error. Failures
in the resource section are recorded as an embedded stream exception and
never cost the header metadata. Extraction can be turned off with
setExtractIcons(false).
Test fixtures are a MinGW-built PE32 EXE and PE32+ DLL with two icon
groups (numeric id and named; BMP and PNG encoded images) plus the source
.ico files, so the tests can check that the rebuilt icons are byte
identical.
* TIKA-4936: harden PE icon extraction after review
Correctness
- Resource tree offsets are relative to the directory root, not the
section start; rebase subdirectory, name and data-entry offsets so a
tree that does not begin at the section's VirtualAddress is read
correctly. Covered by a synthetic PE that moves the real .rsrc 0x300
bytes into its section.
- Rethrow SecurityException and EmbeddedLimitReachedException instead of
recording them as a broken resource section, so throwOnMaxCount and
friends still reach the caller.
- Skip the MS-DOS stub with skipFully rather than InputStream.skip.
Robustness against hostile input
- A group may not list the same icon twice and a rebuilt icon may not
exceed the section size cap, closing a 256x amplification that ended in
an uncatchable OutOfMemoryError. The .ico is assembled into a
pre-sized buffer.
- Every directory entry visited counts against one budget; when it runs
out the walk stops and a warning is recorded, and the icons collected
so far are still emitted.
- The section buffer grows with what is actually read instead of being
allocated from the header's declared size.
Efficiency
- The resource section is read lazily and only as far as the icon data
reaches, so a file without icon types costs its directory, not its
whole section. File-backed input is read through a positioned channel
instead of skipping.
- Icons are indexed by id and language for O(1) lookup.
Idiom
- Metadata.newInstance(context), RESOURCE_NAME_EXTENSION_INFERRED,
EndianUtils.getUIntLE instead of a private copy, tree state passed as
parameters instead of a copy constructor plus depth switch.
- Icons are emitted with outputHtml=false through an
EmbeddedContentHandler, like Office thumbnails: the name is Tika's
invention, so the executable's text output stays empty.
Tests
- The truncation test cut at RVAs that lay beyond the file, so it never
truncated anything; it now cuts at real offsets and asserts the icons
that survive and when an exception is recorded.
- New coverage for the shifted root, duplicate ids, BytesInRes
disagreement, data outside the section, a cycle, the entry budget,
lazy reading, a declining extractor, embedded limits, file-backed
input, the empty body, and language variants with fallback via a new
multi-language DLL fixture. The recording extractor keeps the metadata
as handed over so content type and names are asserted before any
re-detection.
Docs: formats.adoc describes the icon extraction and the extractIcons
option.
* TIKA-4936: keep the released parsePE(InputStream) overload, deprecated
As requested in review: the pre-4.1.1 parsePE(XHTMLContentHandler,
Metadata, InputStream, byte[]) stays as a deprecated metadata-only entry
point pointing at the new overload, so released API keeps working. Both
share parsePEHeader(); only the new one goes on to extract icons.
* TIKA-4936: stop the resource walk below the language level
A subdirectory entry at the language level recursed with depth 2 again,
so the depth cap never applied and a chain of nested directories was
followed until the visited set or the budget stopped it, deep enough to
overflow the stack. Such an entry is malformed; the walk now ignores it.
Test with a 25000 directory chain.
* TIKA-4936: bound icon output across groups, grow the section buffer with
bytes read, address review
Images are deduplicated per group by their data, not their id: several
ids can name one data entry. All rebuilt icons together may take at most
four times the bytes read from the resource section, since any number of
groups can share images; beyond that extraction stops with a warning.
The earlier claim that the per-group id check closed the amplification
was wrong.
The section buffer now grows with the bytes that arrive instead of
jumping to an offset the file declares, so a small file cannot make the
parser allocate the 64 MB cap.
An e_lfanew of 0x3f is an MS-DOS executable again instead of an
IllegalArgumentException from skipFully. An IOException other than EOF
while reading the resource section is the source failing and fails the
parse instead of being recorded as an embedded problem.
Also: @deprecated since 4.2.0, CHANGES entry that names the default-on
behaviour, and the sources and commands the test executables were built
from.
* TIKA-4936: read file-backed resources where the tree points, groups
before icons, clean resource names
File-backed input is no longer buffered as a section prefix: directories
and icons are read at their offsets, so resources stored between the
tree and the icon groups (RCDATA payloads) are not read, nothing is read
for a group the extractor declines, and the 64 MB cap applies only to
the buffered part of a section read from a stream. Icons beyond that cap
are reported instead of the whole section being dropped.
The icon groups are walked before the icons, so a tree that exhausts the
entry budget among its icons still yields the groups whose icons were
reached. The tree-wide visited set is gone; the walk is three levels
deep by construction, and two names may share a language directory.
Resource names have path separators and control characters replaced
before they reach the file name and the relationship id, and a group
that cannot be told apart from an earlier one by name and language comes
out once.
The stream position is derived from the header instead of the stream's
own counter, so the public parsePE overload works with a stream that
starts after the first four bytes. A file that ends before its resource
section is reported for file-backed input as it is for streams.
Adds tests for the icons-per-group and icon size limits.
* TIKA-4936: config class for extractIcons, separate embedded failures from
section failures, split Section
extractIcons moves from a bean setter to ExecutableParserConfig, with a
JsonConfig constructor and a ParseContext override for one parse, as
PSDParser has it. Tests cover the constructor, the context and a JSON
configuration.
Only locating the resource section and walking its tree are guarded
now: a file cut before the section or a broken tree is recorded and
leaves the header metadata alone. What the embedded document extractor
throws while an icon is handed on reaches the caller like any other
embedded document's failure, and extract no longer declares a
TikaException it never threw.
Images of a group may not share bytes at all, which covers overlapping
ranges as well as the same image twice. The language suffix of an icon
name depends on the languages a group name comes in, not on how many
groups carry the name.
Section becomes FileSection and BufferedSection, the tree walk moves
into ResourceTree with one method per level. A parity test runs hostile
and truncated inputs through both the file and the stream path.
* TIKA-4936: honor a per-request JSON config for executable-parser
parsePE looked its config up by class only, which finds an object set
in the same JVM. A request to tika-server or through pipes carries
{"executable-parser": {"extractIcons": false}} as JSON under the
parser's name, which that lookup never saw, so the icons came out
anyway. The config now goes through ParseContextConfig, which resolves
both.
---------
Co-authored-by: Tilman Hausherr <[email protected]>
---
CHANGES.txt | 8 +
docs/modules/ROOT/pages/formats.adoc | 6 +-
.../tika/parser/executable/ExecutableParser.java | 104 +-
.../tika/parser/executable/PEIconExtractor.java | 766 ++++++++++++++
.../parser/executable/ExecutableParserTest.java | 22 +
.../parser/executable/PEIconExtractorTest.java | 1085 ++++++++++++++++++++
.../configs/tika-config-executable-no-icons.json | 10 +
.../test-documents/testWindows-icons-app.ico | Bin 0 -> 7065 bytes
.../test-documents/testWindows-icons-doc.ico | Bin 0 -> 10806 bytes
.../test-documents/testWindows-x86-32-icons.exe | Bin 0 -> 32768 bytes
.../testWindows-x86-64-icons-lang.dll | Bin 0 -> 22528 bytes
.../test-documents/testWindows-x86-64-icons.dll | Bin 0 -> 22528 bytes
12 files changed, 1993 insertions(+), 8 deletions(-)
diff --git a/CHANGES.txt b/CHANGES.txt
index d7bcc18b01..ef89983799 100644
--- a/CHANGES.txt
+++ b/CHANGES.txt
@@ -72,6 +72,14 @@ Release 4.2.0 - unreleased
image/x-win-bitmap (*.cur). Compat: image/vnd.microsoft.icon moves
from ImageParser to ICOParser (TIKA-4937).
+ * ExecutableParser extracts the icon groups of PE files (EXE/DLL) as
+ embedded .ico documents, the first one typed THUMBNAIL. Compat: this is
+ on by default, so every EXE/DLL with icons now yields embedded documents
+ (extra RMETA entries, extra unpacked files); "extractIcons": false on
+ the executable-parser turns it off. parsePE(XHTMLContentHandler,
+ Metadata, InputStream, byte[]) is deprecated in favour of the overload
+ that takes a ParseContext (TIKA-4936).
+
* tika-core's OSGi manifest no longer requires a Service Loader Mediator or
providers for Parser, Detector, EncodingDetector, LanguageDetector and
MetadataFilter; 4.1.0 failed to install on its own in Equinox/p2
diff --git a/docs/modules/ROOT/pages/formats.adoc
b/docs/modules/ROOT/pages/formats.adoc
index 6595fc5129..a4ff75f230 100644
--- a/docs/modules/ROOT/pages/formats.adoc
+++ b/docs/modules/ROOT/pages/formats.adoc
@@ -331,7 +331,11 @@ pull in large native and scientific dependencies and are
not in `tika-app`.
link:{api}/org/apache/tika/parser/executable/ExecutableParser.html[ExecutableParser]
extracts platform, architecture and type metadata from a range of executable
-and library formats, such as Windows PE and Linux/BSD ELF binaries;
+and library formats, such as Windows PE and Linux/BSD ELF binaries. For
+Windows PE files (EXE, DLL) it also emits each icon group as an embedded
+`image/vnd.microsoft.icon` document, the first one, which Windows shows for
+the file, typed as a thumbnail; the text output stays empty. The
+`extractIcons` option (on by default) turns this off.
link:{api}/org/apache/tika/parser/executable/UniversalExecutableParser.html[UniversalExecutableParser]
unpacks macOS universal (fat) binaries.
diff --git
a/tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-code-module/src/main/java/org/apache/tika/parser/executable/ExecutableParser.java
b/tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-code-module/src/main/java/org/apache/tika/parser/executable/ExecutableParser.java
index e23e8193d9..0b9855a8e9 100644
---
a/tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-code-module/src/main/java/org/apache/tika/parser/executable/ExecutableParser.java
+++
b/tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-code-module/src/main/java/org/apache/tika/parser/executable/ExecutableParser.java
@@ -18,6 +18,7 @@ package org.apache.tika.parser.executable;
import java.io.IOException;
import java.io.InputStream;
+import java.io.Serializable;
import java.sql.Date;
import java.util.Arrays;
import java.util.Collections;
@@ -29,6 +30,9 @@ import org.xml.sax.ContentHandler;
import org.xml.sax.SAXException;
import org.apache.tika.annotation.TikaComponent;
+import org.apache.tika.config.ConfigDeserializer;
+import org.apache.tika.config.JsonConfig;
+import org.apache.tika.config.ParseContextConfig;
import org.apache.tika.exception.TikaException;
import org.apache.tika.io.EndianUtils;
import org.apache.tika.io.TikaInputStream;
@@ -79,13 +83,32 @@ public class ExecutableParser implements Parser,
MachineMetadata {
MACH_O_DYLINKER, MACH_O_BUNDLE, MACH_O_DYLIB_STUB,
MACH_O_DSYM,
MACH_O_KEXT_BUNDLE)));
+ private static final String CONFIG_KEY = "executable-parser";
+
+ private ExecutableParserConfig defaultConfig = new
ExecutableParserConfig();
+
+ public ExecutableParser() {
+ }
+
+ public ExecutableParser(ExecutableParserConfig config) {
+ this.defaultConfig = config;
+ }
+
+ public ExecutableParser(JsonConfig jsonConfig) {
+ this(ConfigDeserializer.buildConfig(jsonConfig,
ExecutableParserConfig.class));
+ }
+
public Set<MediaType> getSupportedTypes(ParseContext context) {
return SUPPORTED_TYPES;
}
+ public ExecutableParserConfig getDefaultConfig() {
+ return defaultConfig;
+ }
+
public void parse(TikaInputStream tis, ContentHandler handler, Metadata
metadata,
ParseContext context) throws IOException, SAXException,
TikaException {
- // We only do metadata, for now
+ // We only do metadata (plus icons for PE files), for now
XHTMLContentHandler xhtml = new XHTMLContentHandler(handler, metadata,
context);
xhtml.startDocument();
// What kind is it?
@@ -93,7 +116,7 @@ public class ExecutableParser implements Parser,
MachineMetadata {
IOUtils.readFully(tis, first4);
if (first4[0] == (byte) 'M' && first4[1] == (byte) 'Z') {
- parsePE(xhtml, metadata, tis, first4);
+ parsePE(xhtml, metadata, tis, first4, context);
} else if (first4[0] == (byte) 0x7f && first4[1] == (byte) 'E' &&
first4[2] == (byte) 'L' &&
first4[3] == (byte) 'F') {
parseELF(xhtml, metadata, tis, first4);
@@ -112,10 +135,46 @@ public class ExecutableParser implements Parser,
MachineMetadata {
}
/**
- * Parses a DOS or Windows PE file
+ * Parses a DOS or Windows PE file, extracting metadata only.
+ *
+ * @deprecated since 4.2.0, use
+ * {@link #parsePE(XHTMLContentHandler, Metadata, TikaInputStream, byte[],
ParseContext)},
+ * which also extracts the icons as embedded documents. That is the one
+ * {@link #parse} calls, so overriding this method no longer changes what
+ * a parse does.
*/
+ @Deprecated
public void parsePE(XHTMLContentHandler xhtml, Metadata metadata,
InputStream tis,
byte[] first4) throws TikaException, IOException {
+ parsePEHeader(metadata, tis);
+ }
+
+ /**
+ * Parses a DOS or Windows PE file, extracting metadata and, unless
+ * {@link ExecutableParserConfig#isExtractIcons()} says otherwise, the icon
+ * resources as embedded documents. A configuration in the context, an
+ * {@link ExecutableParserConfig} or JSON under "executable-parser", takes
+ * precedence over the parser's own.
+ */
+ public void parsePE(XHTMLContentHandler xhtml, Metadata metadata,
TikaInputStream tis,
+ byte[] first4, ParseContext context)
+ throws TikaException, IOException, SAXException {
+ CoffHeader header = parsePEHeader(metadata, tis);
+ ExecutableParserConfig config = ParseContextConfig.getConfig(context,
CONFIG_KEY,
+ ExecutableParserConfig.class, defaultConfig);
+ if (header != null && config.isExtractIcons()) {
+ PEIconExtractor.extract(tis, header, xhtml, metadata, context);
+ }
+ }
+
+ /**
+ * Reads the MS-DOS stub and the COFF header into metadata.
+ *
+ * @return the header fields the resource walk needs, or null if this is
+ * not a PE file
+ */
+ private CoffHeader parsePEHeader(Metadata metadata, InputStream tis)
+ throws TikaException, IOException {
metadata.set(HttpHeaders.CONTENT_TYPE, PE_EXE.toString());
metadata.set(PLATFORM, PLATFORM_WINDOWS);
@@ -127,13 +186,13 @@ public class ExecutableParser implements Parser,
MachineMetadata {
int peOffset = EndianUtils.readIntLE(tis);
// Reasonability check - while it may go anywhere, it's normally in
the first few kb
- if (peOffset > 4096 || peOffset < 0x3f) {
- return;
+ if (peOffset > 4096 || peOffset < 0x40) {
+ return null;
}
// Skip the rest of the MS-DOS stub (if PE), until we reach what should
// be the PE header (if this is a PE executable)
- tis.skip(peOffset - 0x40);
+ IOUtils.skipFully(tis, peOffset - 0x40);
// Read the PE header
byte[] pe = new byte[24];
@@ -144,7 +203,7 @@ public class ExecutableParser implements Parser,
MachineMetadata {
// Good, has a valid PE signature
} else {
// Old style MS-DOS
- return;
+ return null;
}
// Read the header values
@@ -255,6 +314,37 @@ public class ExecutableParser implements Parser,
MachineMetadata {
metadata.set(MACHINE_TYPE, MACHINE_UNKNOWN);
break;
}
+ return new CoffHeader(peOffset + pe.length, sizeOptHdrs, numSectors);
+ }
+
+ /**
+ * @param end the file offset right after the COFF header
+ */
+ record CoffHeader(long end, int sizeOptHdrs, int numSections) {
+ }
+
+ /**
+ * Configuration of {@link ExecutableParser}. One set on the
+ * {@link ParseContext}, as an object or as JSON under "executable-parser",
+ * takes the parser's place for that parse.
+ */
+ public static class ExecutableParserConfig implements Serializable {
+
+ private static final long serialVersionUID = 5210478935624115573L;
+
+ private boolean extractIcons = true;
+
+ /**
+ * Whether icons of PE files (EXE/DLL) are extracted as embedded
+ * <code>.ico</code> documents. Defaults to true.
+ */
+ public boolean isExtractIcons() {
+ return extractIcons;
+ }
+
+ public void setExtractIcons(boolean extractIcons) {
+ this.extractIcons = extractIcons;
+ }
}
/**
diff --git
a/tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-code-module/src/main/java/org/apache/tika/parser/executable/PEIconExtractor.java
b/tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-code-module/src/main/java/org/apache/tika/parser/executable/PEIconExtractor.java
new file mode 100644
index 0000000000..f3747046f3
--- /dev/null
+++
b/tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-code-module/src/main/java/org/apache/tika/parser/executable/PEIconExtractor.java
@@ -0,0 +1,766 @@
+/*
+ * Licensed to the Apache Software Foundation (ASF) under one or more
+ * contributor license agreements. See the NOTICE file distributed with
+ * this work for additional information regarding copyright ownership.
+ * The ASF licenses this file to You under the Apache License, Version 2.0
+ * (the "License"); you may not use this file except in compliance with
+ * the License. You may obtain a copy of the License at
+ *
+ * http://www.apache.org/licenses/LICENSE-2.0
+ *
+ * Unless required by applicable law or agreed to in writing, software
+ * distributed under the License is distributed on an "AS IS" BASIS,
+ * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+ * See the License for the specific language governing permissions and
+ * limitations under the License.
+ */
+package org.apache.tika.parser.executable;
+
+import java.io.EOFException;
+import java.io.IOException;
+import java.io.InputStream;
+import java.nio.ByteBuffer;
+import java.nio.ByteOrder;
+import java.nio.channels.FileChannel;
+import java.nio.charset.StandardCharsets;
+import java.util.ArrayList;
+import java.util.Arrays;
+import java.util.Comparator;
+import java.util.HashMap;
+import java.util.HashSet;
+import java.util.LinkedHashMap;
+import java.util.List;
+import java.util.Map;
+import java.util.Set;
+
+import org.apache.commons.io.IOUtils;
+import org.xml.sax.SAXException;
+
+import org.apache.tika.exception.TikaException;
+import org.apache.tika.extractor.EmbeddedDocumentExtractor;
+import org.apache.tika.extractor.EmbeddedDocumentUtil;
+import org.apache.tika.io.EndianUtils;
+import org.apache.tika.io.TikaInputStream;
+import org.apache.tika.metadata.HttpHeaders;
+import org.apache.tika.metadata.Metadata;
+import org.apache.tika.metadata.TikaCoreProperties;
+import org.apache.tika.parser.ParseContext;
+import org.apache.tika.parser.executable.ExecutableParser.CoffHeader;
+import org.apache.tika.sax.EmbeddedContentHandler;
+import org.apache.tika.sax.XHTMLContentHandler;
+
+/**
+ * Extracts the icons of a PE file (EXE/DLL) from its resource section and
+ * hands each icon group to the {@link EmbeddedDocumentExtractor} as a
+ * standalone <code>.ico</code> file.
+ * <p>
+ * Windows stores an icon as a group resource ({@code RT_GROUP_ICON}) that
+ * lists the individual images, which are stored as {@code RT_ICON} resources.
+ * The {@code .ico} file format is nearly identical to the group resource;
+ * the only difference is that the group refers to its images by resource id
+ * whereas the file refers to them by file offset. This class rebuilds the
+ * file from the two resource types.
+ * <p>
+ * File-backed input is read where the resource tree points, so only the
+ * directories and the icons themselves are touched. Anything else has to be
+ * read from the start: up to the resource section, then the section as far
+ * as its icons reach.
+ * Hostile input is contained by bounds checking every offset and by capping
+ * the buffered part of a section, the number of directory entries visited,
+ * the icons per group, the size of a rebuilt icon and the size of all
+ * rebuilt icons together. The tree is walked to its three levels and no
+ * deeper, so a cycle cannot loop.
+ */
+class PEIconExtractor {
+
+ static final String ICON_MIME_TYPE = "image/vnd.microsoft.icon";
+
+ private static final int RT_ICON = 3;
+ private static final int RT_GROUP_ICON = 14;
+
+ private static final int IMAGE_DIRECTORY_ENTRY_RESOURCE = 2;
+ private static final int PE32_MAGIC = 0x10b;
+ private static final int PE32PLUS_MAGIC = 0x20b;
+ private static final int SECTION_HEADER_SIZE = 40;
+ private static final int RESOURCE_DIRECTORY_SIZE = 16;
+ private static final int RESOURCE_DIRECTORY_ENTRY_SIZE = 8;
+ private static final int RESOURCE_DATA_ENTRY_SIZE = 16;
+ private static final int GRP_ICON_DIR_SIZE = 6;
+ private static final int GRP_ICON_DIR_ENTRY_SIZE = 14;
+ private static final int ICON_DIR_ENTRY_SIZE = 16;
+ private static final long HIGH_BIT = 0x80000000L;
+
+ // Sanity limits for hostile input
+ private static final int MAX_SECTIONS = 96; // the PE spec's own limit
+ private static final int MAX_BUFFERED_SECTION_MB = 64;
+ private static final int MAX_DIRECTORY_ENTRIES = 20000; // visited across
the whole tree
+ private static final int MAX_RESOURCE_NAME_LENGTH = 256;
+ private static final int MAX_ICONS_PER_GROUP = 256;
+ private static final long MAX_ICO_SIZE = 64 * 1024 * 1024;
+ // Groups may share images, so all icons together may outgrow the section,
but not by much
+ private static final int MAX_OUTPUT_FACTOR = 4;
+ private static final int SECTION_BUFFER_FLOOR = 8192;
+
+ private PEIconExtractor() {
+ }
+
+ /**
+ * Continues reading the PE file directly after the COFF file header and
+ * emits every icon group as an embedded document.
+ *
+ * @param stream the input positioned right after the COFF header
+ * @param metadata the PE file's own metadata, receives a note if the
+ * resource section is cut or broken and a warning if not
+ * all of the tree or of the icons could be handled
+ */
+ static void extract(TikaInputStream stream, CoffHeader header,
XHTMLContentHandler xhtml,
+ Metadata metadata, ParseContext context) throws
IOException, SAXException {
+ ResourceTree tree;
+ try {
+ tree = ResourceTree.read(stream, header);
+ } catch (SecurityException e) {
+ throw e;
+ } catch (EOFException | RuntimeException e) {
+ // A cut or broken resource section must not cost the caller the
+ // header metadata; any other IOException is the source failing
+ EmbeddedDocumentUtil.recordEmbeddedStreamException(e, metadata,
context);
+ return;
+ }
+ if (tree == null) {
+ return;
+ }
+ if (tree.budgetExceeded) {
+ EmbeddedDocumentUtil.recordException(new TikaException(
+ "PE resource directory has more than " +
MAX_DIRECTORY_ENTRIES +
+ " entries; icon extraction stopped early"),
metadata, context);
+ }
+ emitIcons(tree, xhtml, metadata, context);
+ if (tree.section.beyondLimit()) {
+ EmbeddedDocumentUtil.recordException(new TikaException(
+ "PE resource section is larger than " +
MAX_BUFFERED_SECTION_MB +
+ " MB and the input is not a file; icons beyond
that were not" +
+ " extracted"), metadata, context);
+ }
+ }
+
+ /**
+ * Rebuilds an <code>.ico</code> file for every icon group and passes it on
+ * as an embedded document. The first usable group in resource order is the
+ * one Windows shows for the file itself. The icon adds nothing to the text
+ * output; its name is Tika's invention, not content of the file.
+ */
+ private static void emitIcons(ResourceTree tree, XHTMLContentHandler xhtml,
+ Metadata parentMetadata, ParseContext
context)
+ throws IOException, SAXException {
+ if (tree.groups.isEmpty()) {
+ return;
+ }
+ EmbeddedDocumentExtractor extractor =
+ EmbeddedDocumentUtil.getEmbeddedDocumentExtractor(context);
+ // A group that comes in several languages carries the language in its
name
+ Map<String, Set<Integer>> languages = new HashMap<>();
+ for (Resource group : tree.groups) {
+ languages.computeIfAbsent(group.displayName(), k -> new
HashSet<>())
+ .add(group.language);
+ }
+ Set<String> relationshipIds = new HashSet<>();
+ long emitted = 0;
+ boolean first = true;
+ for (Resource group : tree.groups) {
+ Icon icon = resolveGroup(group, tree);
+ String relationshipId =
+ RT_GROUP_ICON + "/" + group.displayName() + "/" +
group.language;
+ // A name can spell an id, or two names can clean up to the same
one;
+ // what cannot be told apart comes out once
+ if (icon == null || !relationshipIds.add(relationshipId)) {
+ continue;
+ }
+ String name = "icon_" + group.displayName();
+ if (languages.get(group.displayName()).size() > 1) {
+ name += "_" + group.language;
+ }
+ Metadata metadata = Metadata.newInstance(context);
+ metadata.set(TikaCoreProperties.RESOURCE_NAME_KEY, name + ".ico");
+ metadata.set(TikaCoreProperties.RESOURCE_NAME_EXTENSION_INFERRED,
true);
+ metadata.set(HttpHeaders.CONTENT_TYPE, ICON_MIME_TYPE);
+ metadata.set(TikaCoreProperties.EMBEDDED_RELATIONSHIP_ID,
relationshipId);
+ metadata.set(TikaCoreProperties.EMBEDDED_RESOURCE_TYPE, first ?
+
TikaCoreProperties.EmbeddedResourceType.THUMBNAIL.toString() :
+
TikaCoreProperties.EmbeddedResourceType.ATTACHMENT.toString());
+ first = false;
+ if (!extractor.shouldParseEmbedded(metadata, context)) {
+ continue;
+ }
+ emitted += icoSize(icon.images);
+ if (emitted > MAX_OUTPUT_FACTOR * tree.section.available()) {
+ EmbeddedDocumentUtil.recordException(new TikaException(
+ "PE icons add up to more than " + MAX_OUTPUT_FACTOR +
+ " times the resource section they were read
from; icon" +
+ " extraction stopped early"), parentMetadata,
context);
+ return;
+ }
+ byte[] ico = buildIco(icon, tree.section);
+ if (ico == null) {
+ continue;
+ }
+ try (TikaInputStream tis = TikaInputStream.get(ico)) {
+ extractor.parseEmbedded(tis, new
EmbeddedContentHandler(xhtml), metadata, context,
+ false);
+ }
+ }
+ }
+
+ /**
+ * Checks a {@code GRPICONDIR} and looks up the images it references.
+ *
+ * @return the directory and its images in directory order, or null if the
+ * group is unusable
+ */
+ private static Icon resolveGroup(Resource group, ResourceTree tree) throws
IOException {
+ Section section = tree.section;
+ byte[] dir = section.read(group.offset, (int) Math.min(group.size,
+ GRP_ICON_DIR_SIZE + MAX_ICONS_PER_GROUP *
GRP_ICON_DIR_ENTRY_SIZE));
+ if (dir == null || dir.length < GRP_ICON_DIR_SIZE ||
EndianUtils.getUShortLE(dir, 0) != 0 ||
+ EndianUtils.getUShortLE(dir, 2) != 1) {
+ return null;
+ }
+ int count = EndianUtils.getUShortLE(dir, 4);
+ if (count == 0 || count > MAX_ICONS_PER_GROUP ||
+ GRP_ICON_DIR_SIZE + count * GRP_ICON_DIR_ENTRY_SIZE >
group.size) {
+ return null;
+ }
+ List<Resource> images = new ArrayList<>(count);
+ for (int i = 0; i < count; i++) {
+ int id = EndianUtils.getUShortLE(dir,
+ GRP_ICON_DIR_SIZE + i * GRP_ICON_DIR_ENTRY_SIZE + 12);
+ Resource image = tree.findIcon(id, group.language);
+ if (image == null) {
+ return null;
+ }
+ images.add(image);
+ }
+ if (shareBytes(images) || icoSize(images) > MAX_ICO_SIZE) {
+ return null;
+ }
+ for (Resource image : images) {
+ if (!section.canRead(image.offset, image.size)) {
+ return null;
+ }
+ }
+ return new Icon(dir, images);
+ }
+
+ /**
+ * Images that share bytes, the same one listed twice being the plain
+ * case, would let a tiny group inflate into a huge file.
+ */
+ private static boolean shareBytes(List<Resource> images) {
+ List<Resource> sorted = new ArrayList<>(images);
+ sorted.sort(Comparator.comparingLong(image -> image.offset));
+ for (int i = 1; i < sorted.size(); i++) {
+ Resource previous = sorted.get(i - 1);
+ if (sorted.get(i).offset < previous.offset + previous.size) {
+ return true;
+ }
+ }
+ return false;
+ }
+
+ private static long icoSize(List<Resource> images) {
+ long size = GRP_ICON_DIR_SIZE + (long) images.size() *
ICON_DIR_ENTRY_SIZE;
+ for (Resource image : images) {
+ size += image.size;
+ }
+ return size;
+ }
+
+ /**
+ * Converts a {@code GRPICONDIR} plus its {@code RT_ICON} images into an
+ * {@code ICONDIR} based <code>.ico</code> file.
+ *
+ * @return the file, or null if an image could not be read after all
+ */
+ private static byte[] buildIco(Icon icon, Section section) throws
IOException {
+ int count = icon.images.size();
+ int imageOffset = GRP_ICON_DIR_SIZE + count * ICON_DIR_ENTRY_SIZE;
+ ByteBuffer ico = ByteBuffer.allocate((int) icoSize(icon.images))
+ .order(ByteOrder.LITTLE_ENDIAN);
+ // ICONDIR: reserved, type, count - identical to the GRPICONDIR
+ ico.put(icon.dir, 0, GRP_ICON_DIR_SIZE);
+ for (int i = 0; i < count; i++) {
+ int size = (int) icon.images.get(i).size;
+ // width, height, colours, reserved, planes and bit count are
shared
+ ico.put(icon.dir, GRP_ICON_DIR_SIZE + i * GRP_ICON_DIR_ENTRY_SIZE,
8);
+ // the group's BytesInRes may disagree with the actual resource;
trust the resource
+ ico.putInt(size);
+ ico.putInt(imageOffset);
+ imageOffset += size;
+ }
+ for (Resource image : icon.images) {
+ if (!section.read(image.offset, ico.array(), ico.position(), (int)
image.size)) {
+ return null;
+ }
+ ico.position(ico.position() + (int) image.size);
+ }
+ return ico.array();
+ }
+
+ /**
+ * The three level resource tree (type / name / language), reduced to its
+ * icons and icon groups.
+ */
+ private static final class ResourceTree {
+ final Section section;
+ private final long sectionVa;
+ private final long rootOffset;
+ // icon id -> language -> icon, in directory order
+ private final Map<Integer, Map<Integer, Resource>> icons = new
HashMap<>();
+ final List<Resource> groups = new ArrayList<>();
+ private int budget = MAX_DIRECTORY_ENTRIES;
+ boolean budgetExceeded;
+
+ private ResourceTree(Section section, long sectionVa, long rootOffset)
{
+ this.section = section;
+ this.sectionVa = sectionVa;
+ this.rootOffset = rootOffset;
+ }
+
+ /**
+ * Finds the resource section through the optional header and the
+ * section table, which follow the COFF header, and walks its tree.
+ *
+ * @return the tree, or null if the file has no resource section
+ */
+ static ResourceTree read(TikaInputStream stream, CoffHeader header)
throws IOException {
+ if (header.numSections() <= 0 || header.numSections() >
MAX_SECTIONS) {
+ return null;
+ }
+ // The optional header holds the data directories, of which we
+ // need the one pointing at the resource tree
+ byte[] optHdr = new byte[header.sizeOptHdrs()];
+ IOUtils.readFully(stream, optHdr);
+ int dataDirOffset;
+ switch (optHdr.length >= 2 ? EndianUtils.getUShortLE(optHdr, 0) :
0) {
+ case PE32_MAGIC:
+ dataDirOffset = 96;
+ break;
+ case PE32PLUS_MAGIC:
+ dataDirOffset = 112;
+ break;
+ default:
+ return null;
+ }
+ int rsrcEntry = dataDirOffset + IMAGE_DIRECTORY_ENTRY_RESOURCE * 8;
+ if (rsrcEntry + 8 > optHdr.length) {
+ return null;
+ }
+ long numDataDirs = EndianUtils.getUIntLE(optHdr, dataDirOffset -
4);
+ if (numDataDirs <= IMAGE_DIRECTORY_ENTRY_RESOURCE) {
+ return null;
+ }
+ long rsrcRva = EndianUtils.getUIntLE(optHdr, rsrcEntry);
+ long rsrcSize = EndianUtils.getUIntLE(optHdr, rsrcEntry + 4);
+ if (rsrcRva == 0 || rsrcSize == 0) {
+ return null;
+ }
+
+ // The section table tells us where in the file the resource RVA
lives
+ byte[] sections = new byte[header.numSections() *
SECTION_HEADER_SIZE];
+ IOUtils.readFully(stream, sections);
+ long sectionVa = -1;
+ long sectionRawPtr = -1;
+ long sectionRawSize = -1;
+ for (int off = 0; off < sections.length; off +=
SECTION_HEADER_SIZE) {
+ long va = EndianUtils.getUIntLE(sections, off + 12);
+ long rawSize = EndianUtils.getUIntLE(sections, off + 16);
+ long rawPtr = EndianUtils.getUIntLE(sections, off + 20);
+ if (rsrcRva >= va && rsrcRva < va + rawSize) {
+ sectionVa = va;
+ sectionRawPtr = rawPtr;
+ sectionRawSize = rawSize;
+ break;
+ }
+ }
+ if (sectionVa < 0) {
+ return null;
+ }
+
+ Section section;
+ if (stream.hasFile()) {
+ section = new FileSection(stream.getFileChannel(),
sectionRawPtr, sectionRawSize);
+ } else {
+ long position = header.end() + optHdr.length + sections.length;
+ if (sectionRawPtr < position) {
+ return null;
+ }
+ IOUtils.skipFully(stream, sectionRawPtr - position);
+ section = new BufferedSection(stream, sectionRawSize);
+ }
+ // Offsets inside the tree are relative to its root, which normally
+ // but not necessarily sits at the start of the section
+ ResourceTree tree = new ResourceTree(section, sectionVa, rsrcRva -
sectionVa);
+ tree.readTypes();
+ return tree;
+ }
+
+ private void readTypes() throws IOException {
+ long groupDir = -1;
+ long iconDir = -1;
+ for (Entry entry : readEntries(rootOffset)) {
+ // Only icons are interesting, and a well-formed tree lists
each type once
+ if (entry.subdirectory && entry.nameField == RT_GROUP_ICON &&
groupDir < 0) {
+ groupDir = entry.target;
+ } else if (entry.subdirectory && entry.nameField == RT_ICON &&
iconDir < 0) {
+ iconDir = entry.target;
+ }
+ }
+ // Groups first: should the entry budget run out among the icons,
+ // the groups whose icons were reached still come out
+ if (groupDir >= 0) {
+ readNames(groupDir, RT_GROUP_ICON);
+ }
+ if (iconDir >= 0) {
+ readNames(iconDir, RT_ICON);
+ }
+ }
+
+ private void readNames(long dirOffset, int type) throws IOException {
+ for (Entry entry : readEntries(dirOffset)) {
+ if (!entry.subdirectory) {
+ continue;
+ }
+ String name = null;
+ if (entry.named()) {
+ name = readName(rootOffset + (entry.nameField &
~HIGH_BIT));
+ if (name == null) {
+ continue;
+ }
+ }
+ readLanguages(entry.target, type, entry.id(), name);
+ }
+ }
+
+ private void readLanguages(long dirOffset, int type, int id, String
name)
+ throws IOException {
+ for (Entry entry : readEntries(dirOffset)) {
+ // Language ids are always numeric, and a subdirectory below
the
+ // language level is malformed; nothing to find there
+ if (!entry.subdirectory && !entry.named()) {
+ readDataEntry(entry.target, new Resource(type, id, name,
entry.id()));
+ }
+ }
+ }
+
+ /**
+ * @return the entries of the directory, as many as the entry budget
still allows
+ */
+ private List<Entry> readEntries(long dirOffset) throws IOException {
+ byte[] dir = section.read(dirOffset, RESOURCE_DIRECTORY_SIZE);
+ if (dir == null) {
+ return List.of();
+ }
+ int numEntries = EndianUtils.getUShortLE(dir, 12) +
EndianUtils.getUShortLE(dir, 14);
+ int count = Math.min(numEntries, budget);
+ budget -= count;
+ budgetExceeded |= count < numEntries;
+ byte[] bytes = section.read(dirOffset + RESOURCE_DIRECTORY_SIZE,
+ count * RESOURCE_DIRECTORY_ENTRY_SIZE);
+ if (bytes == null) {
+ return List.of();
+ }
+ List<Entry> entries = new ArrayList<>(count);
+ for (int i = 0; i < bytes.length; i +=
RESOURCE_DIRECTORY_ENTRY_SIZE) {
+ long dataField = EndianUtils.getUIntLE(bytes, i + 4);
+ entries.add(new Entry(EndianUtils.getUIntLE(bytes, i),
+ (dataField & HIGH_BIT) != 0, rootOffset + (dataField &
~HIGH_BIT)));
+ }
+ return entries;
+ }
+
+ private void readDataEntry(long offset, Resource resource) throws
IOException {
+ byte[] entry = section.read(offset, RESOURCE_DATA_ENTRY_SIZE);
+ if (entry == null) {
+ return;
+ }
+ // Resource data normally lives in the same section as the tree;
+ // data elsewhere is out of reach
+ resource.offset = EndianUtils.getUIntLE(entry, 0) - sectionVa;
+ resource.size = EndianUtils.getUIntLE(entry, 4);
+ if (!section.contains(resource.offset, resource.size)) {
+ return;
+ }
+ if (resource.type == RT_GROUP_ICON) {
+ groups.add(resource);
+ } else if (resource.name == null) {
+ // Groups reference icons by numeric id, so a named icon is
unreachable
+ icons.computeIfAbsent(resource.id, k -> new LinkedHashMap<>())
+ .putIfAbsent(resource.language, resource);
+ }
+ }
+
+ /**
+ * @return the name with path separators and control characters
+ * replaced: it ends up in a file name and in an id whose parts are
+ * separated by slashes
+ */
+ private String readName(long offset) throws IOException {
+ byte[] prefix = section.read(offset, 2);
+ if (prefix == null) {
+ return null;
+ }
+ int length = EndianUtils.getUShortLE(prefix, 0);
+ if (length == 0 || length > MAX_RESOURCE_NAME_LENGTH) {
+ return null;
+ }
+ byte[] chars = section.read(offset + 2, length * 2);
+ if (chars == null) {
+ return null;
+ }
+ StringBuilder name = new StringBuilder(new String(chars,
StandardCharsets.UTF_16LE));
+ for (int i = 0; i < name.length(); i++) {
+ char c = name.charAt(i);
+ if (c == '/' || c == '\\' || Character.isISOControl(c)) {
+ name.setCharAt(i, '_');
+ }
+ }
+ return name.toString();
+ }
+
+ /**
+ * @return the icon with the given id, preferring the group's language
+ */
+ Resource findIcon(int id, int language) {
+ Map<Integer, Resource> byLanguage = icons.get(id);
+ if (byLanguage == null) {
+ return null;
+ }
+ Resource icon = byLanguage.get(language);
+ return icon != null ? icon : byLanguage.values().iterator().next();
+ }
+ }
+
+ /**
+ * One entry of a resource directory.
+ *
+ * @param nameField the id, or with the high bit set the offset of the
name
+ * @param subdirectory whether the target is a directory rather than a
data entry
+ * @param target the section offset the entry points to
+ */
+ private record Entry(long nameField, boolean subdirectory, long target) {
+
+ boolean named() {
+ return (nameField & HIGH_BIT) != 0;
+ }
+
+ int id() {
+ return (int) (nameField & 0xffff);
+ }
+ }
+
+ /**
+ * A leaf of the resource tree: type, name (or id) and language, plus the
+ * location of its data within the resource section.
+ */
+ private static final class Resource {
+ final int type;
+ final int id;
+ final String name;
+ final int language;
+ long offset;
+ long size;
+
+ Resource(int type, int id, String name, int language) {
+ this.type = type;
+ this.id = id;
+ this.name = name;
+ this.language = language;
+ }
+
+ String displayName() {
+ return name != null ? name : Integer.toString(id);
+ }
+ }
+
+ /**
+ * A usable icon group: its {@code GRPICONDIR} and the images it lists.
+ */
+ private record Icon(byte[] dir, List<Resource> images) {
+ }
+
+ /**
+ * The bytes of the resource section, addressed by their offset in it.
+ */
+ abstract static class Section {
+ private final long declaredSize;
+
+ Section(long declaredSize) {
+ this.declaredSize = declaredSize;
+ }
+
+ /**
+ * @return whether the range lies inside the section as declared
+ */
+ final boolean contains(long offset, long length) {
+ return offset >= 0 && length >= 0 && offset + length <=
declaredSize;
+ }
+
+ /**
+ * @return the number of bytes of the section known to exist
+ */
+ abstract long available();
+
+ /**
+ * @return whether the range can be read; finding out may read the
input that far
+ */
+ abstract boolean canRead(long offset, long length) throws IOException;
+
+ /**
+ * Copies a range that {@link #canRead(long, long)} has confirmed.
+ *
+ * @return false if the bytes are not there after all
+ */
+ abstract boolean copy(long offset, byte[] target, int targetOffset,
int length)
+ throws IOException;
+
+ /**
+ * @return whether a range was refused that the section declares but
+ * this class does not reach
+ */
+ boolean beyondLimit() {
+ return false;
+ }
+
+ /**
+ * @return false if the range cannot be read
+ */
+ final boolean read(long offset, byte[] target, int targetOffset, int
length)
+ throws IOException {
+ return canRead(offset, length) && copy(offset, target,
targetOffset, length);
+ }
+
+ /**
+ * @return the bytes, or null if the range cannot be read
+ */
+ final byte[] read(long offset, int length) throws IOException {
+ if (!canRead(offset, length)) {
+ return null;
+ }
+ byte[] bytes = new byte[length];
+ return copy(offset, bytes, 0, length) ? bytes : null;
+ }
+ }
+
+ /**
+ * A section of a file, read where the walk points.
+ */
+ static final class FileSection extends Section {
+ private final FileChannel channel;
+ private final long start;
+ private final long present;
+
+ /**
+ * @param start the file offset of the section
+ */
+ FileSection(FileChannel channel, long start, long declaredSize) throws
IOException {
+ super(declaredSize);
+ if (start > channel.size()) {
+ throw new EOFException("The file ends before its resource
section");
+ }
+ this.channel = channel;
+ this.start = start;
+ this.present = Math.min(declaredSize, channel.size() - start);
+ }
+
+ @Override
+ long available() {
+ return present;
+ }
+
+ @Override
+ boolean canRead(long offset, long length) {
+ return contains(offset, length) && offset + length <= present;
+ }
+
+ @Override
+ boolean copy(long offset, byte[] target, int targetOffset, int length)
+ throws IOException {
+ ByteBuffer buffer = ByteBuffer.wrap(target, targetOffset, length);
+ long position = start + offset;
+ while (buffer.hasRemaining()) {
+ int n = channel.read(buffer, position);
+ if (n < 0) {
+ return false;
+ }
+ position += n;
+ }
+ return true;
+ }
+ }
+
+ /**
+ * A section that can only be read forward. It is buffered from its start
+ * as far as the walk has reached; the buffer grows with the bytes that
+ * arrive, never with a size the file merely declares.
+ */
+ static final class BufferedSection extends Section {
+ private final InputStream source;
+ private final long limit;
+ byte[] buf = new byte[0];
+ private int length;
+ private boolean eof;
+ private boolean beyondLimit;
+
+ /**
+ * @param source a stream positioned at the start of the section
+ */
+ BufferedSection(InputStream source, long declaredSize) {
+ super(declaredSize);
+ this.source = source;
+ this.limit = Math.min(declaredSize, MAX_BUFFERED_SECTION_MB *
1024L * 1024L);
+ }
+
+ @Override
+ long available() {
+ return length;
+ }
+
+ @Override
+ boolean canRead(long offset, long length) throws IOException {
+ if (!contains(offset, length)) {
+ return false;
+ }
+ if (offset + length > limit) {
+ beyondLimit = true;
+ return false;
+ }
+ return ensure((int) (offset + length));
+ }
+
+ @Override
+ boolean copy(long offset, byte[] target, int targetOffset, int length)
{
+ System.arraycopy(buf, (int) offset, target, targetOffset, length);
+ return true;
+ }
+
+ @Override
+ boolean beyondLimit() {
+ return beyondLimit;
+ }
+
+ private boolean ensure(int end) throws IOException {
+ while (length < end && !eof) {
+ if (length == buf.length) {
+ buf = Arrays.copyOf(buf, (int) Math.min(limit,
+ Math.max(SECTION_BUFFER_FLOOR, 2L * buf.length)));
+ }
+ int n = source.read(buf, length, Math.min(end, buf.length) -
length);
+ if (n < 0) {
+ eof = true;
+ } else {
+ length += n;
+ }
+ }
+ return end <= length;
+ }
+ }
+}
diff --git
a/tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-code-module/src/test/java/org/apache/tika/parser/executable/ExecutableParserTest.java
b/tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-code-module/src/test/java/org/apache/tika/parser/executable/ExecutableParserTest.java
index 88e0c84a96..a412c509bb 100644
---
a/tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-code-module/src/test/java/org/apache/tika/parser/executable/ExecutableParserTest.java
+++
b/tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-code-module/src/test/java/org/apache/tika/parser/executable/ExecutableParserTest.java
@@ -17,14 +17,18 @@
package org.apache.tika.parser.executable;
import static org.junit.jupiter.api.Assertions.assertEquals;
+import static org.junit.jupiter.api.Assertions.assertNull;
import org.junit.jupiter.api.Test;
+import org.xml.sax.helpers.DefaultHandler;
import org.apache.tika.TikaTest;
+import org.apache.tika.io.TikaInputStream;
import org.apache.tika.metadata.HttpHeaders;
import org.apache.tika.metadata.MachineMetadata.Endian;
import org.apache.tika.metadata.Metadata;
import org.apache.tika.metadata.TikaCoreProperties;
+import org.apache.tika.parser.ParseContext;
public class ExecutableParserTest extends TikaTest {
@@ -44,6 +48,24 @@ public class ExecutableParserTest extends TikaTest {
}
+ /**
+ * A PE header cannot start inside the field that holds its offset; such a
+ * file is an MS-DOS executable, not a failure.
+ */
+ @Test
+ public void testPeOffsetInsideDosHeader() throws Exception {
+ byte[] mz = new byte[0x100];
+ mz[0] = 'M';
+ mz[1] = 'Z';
+ mz[0x3c] = 0x3f;
+ Metadata metadata = new Metadata();
+ try (TikaInputStream tis = TikaInputStream.get(mz)) {
+ new ExecutableParser().parse(tis, new DefaultHandler(), metadata,
new ParseContext());
+ }
+ assertEquals("Windows", metadata.get(ExecutableParser.PLATFORM));
+ assertNull(metadata.get(ExecutableParser.MACHINE_TYPE));
+ }
+
@Test
public void testElfParser_x86_32() throws Exception {
XMLResult r = getXML("testLinux-x86-32");
diff --git
a/tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-code-module/src/test/java/org/apache/tika/parser/executable/PEIconExtractorTest.java
b/tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-code-module/src/test/java/org/apache/tika/parser/executable/PEIconExtractorTest.java
new file mode 100644
index 0000000000..c5d66be4bf
--- /dev/null
+++
b/tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-code-module/src/test/java/org/apache/tika/parser/executable/PEIconExtractorTest.java
@@ -0,0 +1,1085 @@
+/*
+ * Licensed to the Apache Software Foundation (ASF) under one or more
+ * contributor license agreements. See the NOTICE file distributed with
+ * this work for additional information regarding copyright ownership.
+ * The ASF licenses this file to You under the Apache License, Version 2.0
+ * (the "License"); you may not use this file except in compliance with
+ * the License. You may obtain a copy of the License at
+ *
+ * http://www.apache.org/licenses/LICENSE-2.0
+ *
+ * Unless required by applicable law or agreed to in writing, software
+ * distributed under the License is distributed on an "AS IS" BASIS,
+ * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+ * See the License for the specific language governing permissions and
+ * limitations under the License.
+ */
+package org.apache.tika.parser.executable;
+
+import static org.junit.jupiter.api.Assertions.assertArrayEquals;
+import static org.junit.jupiter.api.Assertions.assertEquals;
+import static org.junit.jupiter.api.Assertions.assertNotNull;
+import static org.junit.jupiter.api.Assertions.assertNull;
+import static org.junit.jupiter.api.Assertions.assertThrows;
+import static org.junit.jupiter.api.Assertions.assertTrue;
+
+import java.io.ByteArrayInputStream;
+import java.io.EOFException;
+import java.io.FilterInputStream;
+import java.io.IOException;
+import java.io.InputStream;
+import java.nio.ByteBuffer;
+import java.nio.ByteOrder;
+import java.nio.channels.FileChannel;
+import java.nio.charset.StandardCharsets;
+import java.nio.file.Files;
+import java.nio.file.Path;
+import java.nio.file.StandardOpenOption;
+import java.util.ArrayList;
+import java.util.Arrays;
+import java.util.LinkedHashMap;
+import java.util.List;
+import java.util.Map;
+
+import org.apache.commons.io.input.CountingInputStream;
+import org.junit.jupiter.api.Test;
+import org.junit.jupiter.api.io.TempDir;
+import org.xml.sax.ContentHandler;
+
+import org.apache.tika.TikaTest;
+import org.apache.tika.config.EmbeddedLimits;
+import org.apache.tika.config.ParseTimeout;
+import org.apache.tika.config.TimeoutLimits;
+import org.apache.tika.config.loader.TikaLoader;
+import org.apache.tika.exception.EmbeddedLimitReachedException;
+import org.apache.tika.extractor.EmbeddedDocumentExtractor;
+import org.apache.tika.io.EndianUtils;
+import org.apache.tika.io.TikaInputStream;
+import org.apache.tika.metadata.HttpHeaders;
+import org.apache.tika.metadata.Metadata;
+import org.apache.tika.metadata.TikaCoreProperties;
+import org.apache.tika.parser.ParseContext;
+import org.apache.tika.parser.ParseRecord;
+import org.apache.tika.parser.Parser;
+import org.apache.tika.sax.BodyContentHandler;
+import org.apache.tika.sax.XHTMLContentHandler;
+
+/**
+ * The test executables were built with MinGW from an empty {@code main()}
+ * respectively an empty DLL plus a resource script that embeds two icon
+ * groups: id 1 (from testWindows-icons-app.ico, 16/32 px BMP + 256 px PNG)
+ * and the named group DOCICON (from testWindows-icons-doc.ico, 16/48 px BMP).
+ * testWindows-x86-64-icons-lang.dll instead carries group 1 twice, in
+ * language 1033 (app.ico) and 1031 (doc.ico). The two .ico files were drawn
+ * for these tests, a blue circle and a red square.
+ * <p>
+ * To rebuild them, on Ubuntu 24.04 with gcc-mingw-w64 13.2.0 (13-win32) and
+ * binutils-mingw-w64 2.41.90; the result differs from the files here only in
+ * the link timestamp (0x88, in the DLLs also 0xc04) and the checksum (0xd8):
+ * <pre>
+ * main.c: int main(void) { return 0; }
+ * dll.c: __declspec(dllexport) int tika_answer(void) { return 42; }
+ * icons.rc: 1 ICON "testWindows-icons-app.ico"
+ * DOCICON ICON "testWindows-icons-doc.ico"
+ * icons-lang.rc: LANGUAGE 9, 1
+ * 1 ICON "testWindows-icons-app.ico"
+ * LANGUAGE 7, 1
+ * 1 ICON "testWindows-icons-doc.ico"
+ *
+ * i686-w64-mingw32-windres icons.rc icons32.o
+ * i686-w64-mingw32-gcc -Os -s -o testWindows-x86-32-icons.exe main.c icons32.o
+ * x86_64-w64-mingw32-windres icons.rc icons64.o
+ * x86_64-w64-mingw32-gcc -Os -s -shared -nostdlib -o
testWindows-x86-64-icons.dll \
+ * dll.c icons64.o
+ * x86_64-w64-mingw32-windres icons-lang.rc icons-lang64.o
+ * x86_64-w64-mingw32-gcc -Os -s -shared -nostdlib -o
testWindows-x86-64-icons-lang.dll \
+ * dll.c icons-lang64.o
+ * </pre>
+ * Layout of testWindows-x86-32-icons.exe (file offsets, from pefile):
+ * .rsrc raw 0x3400-0x7c00 at RVA 0xa000; root directory 0x3400 with the
+ * type entries at 0x3410 (RT_ICON) and 0x3418 (RT_GROUP_ICON); data entries
+ * 0x3530-0x3590 (icons 1-5, DOCICON, group 1); icon data 0x35a0-0x7b18
+ * (icon 2 spans 0x3a08-0x4ab0); GRPICONDIR DOCICON 0x7b18-0x7b3a and
+ * group 1 0x7b40-0x7b70.
+ */
+public class PEIconExtractorTest extends TikaTest {
+
+ private static final String EXE = "testWindows-x86-32-icons.exe";
+ private static final String DLL = "testWindows-x86-64-icons.dll";
+ private static final String LANG_DLL = "testWindows-x86-64-icons-lang.dll";
+ private static final String APP_ICO = "testWindows-icons-app.ico";
+ private static final String DOC_ICO = "testWindows-icons-doc.ico";
+
+ private static final String THUMBNAIL =
+ TikaCoreProperties.EmbeddedResourceType.THUMBNAIL.toString();
+ private static final String ATTACHMENT =
+ TikaCoreProperties.EmbeddedResourceType.ATTACHMENT.toString();
+
+ @Test
+ public void testIconsFromExe() throws Exception {
+ List<Metadata> metadataList = getRecursiveMetadata(EXE);
+ assertEquals(3, metadataList.size());
+
+ Metadata exe = metadataList.get(0);
+ assertEquals("application/x-msdownload",
exe.get(HttpHeaders.CONTENT_TYPE));
+ assertEquals(ExecutableParser.MACHINE_x86_32,
exe.get(ExecutableParser.MACHINE_TYPE));
+ assertEquals("32", exe.get(ExecutableParser.ARCHITECTURE_BITS));
+
assertNull(exe.get(TikaCoreProperties.TIKA_META_EXCEPTION_EMBEDDED_STREAM));
+ assertNull(exe.get(TikaCoreProperties.TIKA_META_EXCEPTION_WARNING));
+
+ // Named entries precede numeric ids in the resource directory, so the
+ // DOCICON group is the first one and therefore the file's own icon
+ assertIcon(metadataList.get(1), "icon_DOCICON.ico", "14/DOCICON/1033",
THUMBNAIL);
+ assertIcon(metadataList.get(2), "icon_1.ico", "14/1/1033", ATTACHMENT);
+ }
+
+ @Test
+ public void testIconsFromDll() throws Exception {
+ List<Metadata> metadataList = getRecursiveMetadata(DLL);
+ assertEquals(3, metadataList.size());
+
+ Metadata dll = metadataList.get(0);
+ assertEquals("application/x-msdownload",
dll.get(HttpHeaders.CONTENT_TYPE));
+ assertEquals(ExecutableParser.MACHINE_x86_64,
dll.get(ExecutableParser.MACHINE_TYPE));
+ assertEquals("64", dll.get(ExecutableParser.ARCHITECTURE_BITS));
+
+ assertIcon(metadataList.get(1), "icon_DOCICON.ico", "14/DOCICON/1033",
THUMBNAIL);
+ assertIcon(metadataList.get(2), "icon_1.ico", "14/1/1033", ATTACHMENT);
+ }
+
+ /**
+ * The icon's name is Tika's invention, so it must not become text of the
file.
+ */
+ @Test
+ public void testIconsAddNoText() throws Exception {
+ assertContains("<body />", getXML(EXE).xml);
+ }
+
+ /**
+ * The resource compiler copies the images verbatim, so the rebuilt
+ * .ico files must be identical to the ones that went in. Checks the
+ * metadata as handed to the extractor, before any re-detection.
+ */
+ @Test
+ public void testReconstructedIcoIsByteIdentical() throws Exception {
+ for (String file : new String[]{EXE, DLL}) {
+ RecordingExtractor extractor = parse(readTestResource(file));
+ assertEquals(2, extractor.contents.size(), file);
+ assertArrayEquals(readTestResource(DOC_ICO),
extractor.contents.get(0), file);
+ assertArrayEquals(readTestResource(APP_ICO),
extractor.contents.get(1), file);
+ assertIcon(extractor.metadata.get(0), "icon_DOCICON.ico",
"14/DOCICON/1033", THUMBNAIL);
+ assertIcon(extractor.metadata.get(1), "icon_1.ico", "14/1/1033",
ATTACHMENT);
+ }
+ }
+
+ /**
+ * File-backed input is read through a positioned channel rather than by
+ * skipping; the result must not differ.
+ */
+ @Test
+ public void testFileBackedInput() throws Exception {
+ RecordingExtractor extractor = new RecordingExtractor();
+ try (TikaInputStream tis =
TikaInputStream.get(getResourceAsFile("/test-documents/" + EXE)
+ .toPath())) {
+ assertTrue(tis.hasFile());
+ new ExecutableParser().parse(tis, new BodyContentHandler(), new
Metadata(),
+ extractor.context());
+ }
+ assertEquals(2, extractor.contents.size());
+ assertArrayEquals(readTestResource(DOC_ICO),
extractor.contents.get(0));
+ assertArrayEquals(readTestResource(APP_ICO),
extractor.contents.get(1));
+ }
+
+ @Test
+ public void testLanguageVariants() throws Exception {
+ RecordingExtractor extractor = parse(readTestResource(LANG_DLL));
+ assertEquals(2, extractor.contents.size());
+ // 1031 sorts before 1033 in the language directory
+ assertIcon(extractor.metadata.get(0), "icon_1_1031.ico", "14/1/1031",
THUMBNAIL);
+ assertIcon(extractor.metadata.get(1), "icon_1_1033.ico", "14/1/1033",
ATTACHMENT);
+ assertArrayEquals(readTestResource(DOC_ICO),
extractor.contents.get(0));
+ assertArrayEquals(readTestResource(APP_ICO),
extractor.contents.get(1));
+ }
+
+ /**
+ * A group whose language has no matching icons falls back to whatever
+ * language the icons come in: icons 4 and 5 are relabelled from 1031 to
+ * 1033 (their language entries at 0x10b0 and 0x10c8), the 1031 group
+ * must still find them.
+ */
+ @Test
+ public void testLanguageFallback() throws Exception {
+ byte[] dll = readTestResource(LANG_DLL);
+ putIntLE(dll, 0x10b0, 1033);
+ putIntLE(dll, 0x10c8, 1033);
+ RecordingExtractor extractor = parse(dll);
+ assertEquals(2, extractor.contents.size());
+ assertIcon(extractor.metadata.get(0), "icon_1_1031.ico", "14/1/1031",
THUMBNAIL);
+ assertArrayEquals(readTestResource(DOC_ICO),
extractor.contents.get(0));
+ }
+
+ /**
+ * The GRPICONDIR's BytesInRes is not trusted; the icon resource's own
+ * size is. Corrupting BytesInRes of DOCICON's first entry (0x7b18 + 6 + 8)
+ * must not change the output.
+ */
+ @Test
+ public void testGroupSizeFieldIsIgnored() throws Exception {
+ byte[] exe = readTestResource(EXE);
+ putIntLE(exe, 0x7b26, 0x00ffffff);
+ RecordingExtractor extractor = parse(exe);
+ assertEquals(2, extractor.contents.size());
+ assertArrayEquals(readTestResource(DOC_ICO),
extractor.contents.get(0));
+ }
+
+ /**
+ * A group that lists the same icon twice could inflate to 256 times the
+ * icon; it is dropped and the next group takes the thumbnail slot.
+ */
+ @Test
+ public void testDuplicateIconIdInGroupIsRejected() throws Exception {
+ byte[] exe = readTestResource(EXE);
+ // nID of DOCICON's second entry (0x7b18 + 6 + 14 + 12) := id of the
first (4)
+ exe[0x7b38] = 4;
+ exe[0x7b39] = 0;
+ RecordingExtractor extractor = parse(exe);
+ assertEquals(1, extractor.contents.size());
+ assertIcon(extractor.metadata.get(0), "icon_1.ico", "14/1/1033",
THUMBNAIL);
+ assertArrayEquals(readTestResource(APP_ICO),
extractor.contents.get(0));
+ }
+
+ /**
+ * Two ids can name one image: icon 5's data entry (0x3570) is pointed at
+ * icon 4's data (0x3560), so DOCICON lists the same bytes twice under
+ * different ids. It is dropped like a repeated id.
+ */
+ @Test
+ public void testSharedImageDataInGroupIsRejected() throws Exception {
+ byte[] exe = readTestResource(EXE);
+ System.arraycopy(exe, 0x3560, exe, 0x3570, 8);
+ RecordingExtractor extractor = parse(exe);
+ assertEquals(1, extractor.contents.size());
+ assertIcon(extractor.metadata.get(0), "icon_1.ico", "14/1/1033",
THUMBNAIL);
+ assertArrayEquals(readTestResource(APP_ICO),
extractor.contents.get(0));
+ }
+
+ /**
+ * Images may not share bytes at all: icon 5's data entry (0x3570) starts
+ * one byte into icon 4's data.
+ */
+ @Test
+ public void testOverlappingImagesInGroupAreRejected() throws Exception {
+ byte[] exe = readTestResource(EXE);
+ putIntLE(exe, 0x3570, EndianUtils.getIntLE(exe, 0x3560) + 1);
+ putIntLE(exe, 0x3574, EndianUtils.getIntLE(exe, 0x3564));
+ RecordingExtractor extractor = parse(exe);
+ assertEquals(1, extractor.contents.size());
+ assertIcon(extractor.metadata.get(0), "icon_1.ico", "14/1/1033",
THUMBNAIL);
+ }
+
+ /**
+ * What goes wrong while an icon is handed on is not a broken resource
+ * section: it reaches the caller like any other embedded document's
failure.
+ */
+ @Test
+ public void testEmbeddedFailureSurfaces() throws Exception {
+ RecordingExtractor extractor = new RecordingExtractor() {
+ @Override
+ public void parseEmbedded(TikaInputStream stream, ContentHandler
handler,
+ Metadata metadata, ParseContext context,
+ boolean outputHtml) throws IOException {
+ throw new EOFException("embedded failed");
+ }
+ };
+ assertThrows(EOFException.class, () -> parse(readTestResource(EXE),
extractor));
+ assertNull(extractor.parentMetadata.get(
+ TikaCoreProperties.TIKA_META_EXCEPTION_EMBEDDED_STREAM));
+ }
+
+ /**
+ * Any number of groups can share one image, each rebuilding it into a file
+ * of its own. Together they may not outgrow a small multiple of the
+ * section they were read from; extraction stops there with a warning.
+ */
+ @Test
+ public void testOutputBudgetAcrossGroups() throws Exception {
+ int groups = 40;
+ int imageSize = 4096;
+ byte[] pe = sharedImagePe(groups, imageSize);
+ RecordingExtractor extractor = parse(pe);
+ long emitted = 0;
+ for (byte[] ico : extractor.contents) {
+ emitted += ico.length;
+ }
+ assertTrue(emitted > imageSize, "emitted " + emitted);
+ assertTrue(emitted <= 4L * (pe.length - SyntheticPE.SECTION_RAW_PTR),
"emitted " + emitted);
+ assertTrue(extractor.contents.size() < groups);
+ assertContains("stopped early",
+
extractor.parentMetadata.get(TikaCoreProperties.TIKA_META_EXCEPTION_WARNING));
+ }
+
+ /**
+ * The entry budget is spent on the groups first: 12000 icons take more
+ * than the 20000 entries allowed, yet the group whose icon was reached
+ * before the budget ran out still comes out.
+ */
+ @Test
+ public void testGroupsAreReadBeforeIcons() throws Exception {
+ RecordingExtractor extractor = parse(manyIconsPe());
+ assertEquals(1, extractor.contents.size());
+ assertIcon(extractor.metadata.get(0), "icon_1.ico", "14/1/1033",
THUMBNAIL);
+ assertContains("20000 entries",
+
extractor.parentMetadata.get(TikaCoreProperties.TIKA_META_EXCEPTION_WARNING));
+ }
+
+ /**
+ * A group may list 256 images, not 257.
+ */
+ @Test
+ public void testIconsPerGroupLimit() throws Exception {
+ int[] ids = new int[257];
+ for (int i = 0; i < ids.length; i++) {
+ ids[i] = i + 1;
+ }
+ byte[] tooMany = grpIconDir(ids);
+ byte[] allowed = grpIconDir(Arrays.copyOf(ids, 256));
+ byte[] data = new byte[tooMany.length + allowed.length + ids.length];
+ System.arraycopy(tooMany, 0, data, 0, tooMany.length);
+ System.arraycopy(allowed, 0, data, tooMany.length, allowed.length);
+ SyntheticResources resources = new SyntheticResources(data);
+ resources.group(0, tooMany.length);
+ resources.group(tooMany.length, allowed.length);
+ for (int i = 0; i < ids.length; i++) {
+ resources.icon(tooMany.length + allowed.length + i, 1);
+ }
+ RecordingExtractor extractor =
parse(SyntheticPE.build(resources.build(), 0));
+ assertEquals(1, extractor.contents.size());
+ assertIcon(extractor.metadata.get(0), "icon_2.ico", "14/2/1033",
THUMBNAIL);
+ assertEquals(6 + 256 * 16 + 256, extractor.contents.get(0).length);
+ }
+
+ /**
+ * A group whose images add up to more than 64 MB is skipped on its own;
+ * the group after it still comes out and nothing is reported. The three
+ * 23 MB images lie in the hole of a sparse file, which only a file-backed
+ * parse can reach.
+ */
+ @Test
+ public void testOversizedGroupIsSkipped(@TempDir Path tmp) throws
Exception {
+ int imageSize = 23 * 1024 * 1024;
+ byte[] oversized = grpIconDir(1, 2, 3);
+ byte[] small = grpIconDir(4);
+ int images = oversized.length + small.length;
+ byte[] data = new byte[images + 16];
+ System.arraycopy(oversized, 0, data, 0, oversized.length);
+ System.arraycopy(small, 0, data, oversized.length, small.length);
+ SyntheticResources resources = new SyntheticResources(data);
+ resources.group(0, oversized.length);
+ resources.group(oversized.length, small.length);
+ for (int i = 0; i < 3; i++) {
+ resources.icon(data.length + i * imageSize, imageSize);
+ }
+ resources.icon(images, 16);
+ byte[] section = resources.build();
+ int declared = section.length + 3 * imageSize;
+ Path file = Files.write(tmp.resolve("oversized.exe"),
+ SyntheticPE.build(section, 0, declared));
+ try (FileChannel channel = FileChannel.open(file,
StandardOpenOption.WRITE)) {
+ channel.write(ByteBuffer.allocate(1), SyntheticPE.SECTION_RAW_PTR
+ declared - 1);
+ }
+ RecordingExtractor extractor = parseFile(file);
+ assertEquals(1, extractor.contents.size());
+ assertIcon(extractor.metadata.get(0), "icon_2.ico", "14/2/1033",
THUMBNAIL);
+
assertNull(extractor.parentMetadata.get(TikaCoreProperties.TIKA_META_EXCEPTION_WARNING));
+ }
+
+ /**
+ * A resource name ends up in a file name and in an id that is separated
+ * by slashes: DOCICON's I (0x3528) becomes a slash.
+ */
+ @Test
+ public void testResourceNameIsCleaned() throws Exception {
+ byte[] exe = readTestResource(EXE);
+ exe[0x3528] = '/';
+ RecordingExtractor extractor = parse(exe);
+ assertEquals(2, extractor.contents.size());
+ assertIcon(extractor.metadata.get(0), "icon_DOC_CON.ico",
"14/DOC_CON/1033", THUMBNAIL);
+ }
+
+ /**
+ * A name can spell an id: DOCICON is renamed to "1" (length at 0x3520,
+ * characters from 0x3522) next to the group with id 1. Only one of the
+ * two can be told apart by name and id, so only one comes out.
+ */
+ @Test
+ public void testNameThatSpellsAnId() throws Exception {
+ byte[] exe = readTestResource(EXE);
+ exe[0x3520] = 1;
+ exe[0x3522] = '1';
+ RecordingExtractor extractor = parse(exe);
+ assertEquals(1, extractor.contents.size());
+ assertIcon(extractor.metadata.get(0), "icon_1.ico", "14/1/1033",
THUMBNAIL);
+ assertArrayEquals(readTestResource(DOC_ICO),
extractor.contents.get(0));
+ }
+
+ /**
+ * A stream that starts after the four bytes the caller has already read
+ * is all the public entry point promises; the section must still be found.
+ */
+ @Test
+ public void testStreamStartingAfterFirstFourBytes() throws Exception {
+ byte[] exe = readTestResource(EXE);
+ RecordingExtractor extractor = new RecordingExtractor();
+ ParseContext context = extractor.context();
+ try (TikaInputStream tis = TikaInputStream.get(Arrays.copyOfRange(exe,
4, exe.length))) {
+ XHTMLContentHandler xhtml = new XHTMLContentHandler(new
BodyContentHandler(),
+ extractor.parentMetadata, context);
+ new ExecutableParser().parsePE(xhtml, extractor.parentMetadata,
tis,
+ Arrays.copyOf(exe, 4), context);
+ }
+ assertEquals(2, extractor.contents.size());
+ assertArrayEquals(readTestResource(DOC_ICO),
extractor.contents.get(0));
+ }
+
+ /**
+ * A section larger than what is buffered from a stream still yields the
+ * icons within reach; one beyond it is reported. A file has no such limit.
+ */
+ @Test
+ public void testSectionLargerThanBuffer(@TempDir Path tmp) throws
Exception {
+ int declared = 100 * 1024 * 1024;
+ byte[] rsrc = resourceSectionAt(SyntheticPE.SECTION_VA);
+ byte[] pe = SyntheticPE.build(rsrc, 0, declared);
+ RecordingExtractor extractor = parse(pe);
+ assertEquals(2, extractor.contents.size());
+
assertNull(extractor.parentMetadata.get(TikaCoreProperties.TIKA_META_EXCEPTION_WARNING));
+
+ extractor = parseFile(Files.write(tmp.resolve("large-section.exe"),
pe));
+ assertEquals(2, extractor.contents.size());
+
+ // icon 1 (data entry 0x130) now lies 70 MB into the section
+ putIntLE(rsrc, 0x130, SyntheticPE.SECTION_VA + 70 * 1024 * 1024);
+ extractor = parse(SyntheticPE.build(rsrc, 0, declared));
+ assertEquals(1, extractor.contents.size());
+ assertIcon(extractor.metadata.get(0), "icon_DOCICON.ico",
"14/DOCICON/1033", THUMBNAIL);
+ assertContains("64 MB",
+
extractor.parentMetadata.get(TikaCoreProperties.TIKA_META_EXCEPTION_WARNING));
+ }
+
+ /**
+ * A file that ends before its resource section is reported the same way
+ * whether it is read from a file or from a stream.
+ */
+ @Test
+ public void testFileBackedTruncation(@TempDir Path tmp) throws Exception {
+ RecordingExtractor extractor =
parseFile(Files.write(tmp.resolve("truncated.exe"),
+ Arrays.copyOf(readTestResource(EXE), 0x3000)));
+ assertEquals(0, extractor.contents.size());
+ assertEquals(ExecutableParser.MACHINE_x86_32,
+ extractor.parentMetadata.get(ExecutableParser.MACHINE_TYPE));
+ assertContains("EOFException", extractor.parentMetadata.get(
+ TikaCoreProperties.TIKA_META_EXCEPTION_EMBEDDED_STREAM));
+ }
+
+ /**
+ * A file and a stream take different paths through the resource section.
+ * Whatever the input, they must come to the same icons and the same notes.
+ */
+ @Test
+ public void testFileBackedInputBehavesLikeStream(@TempDir Path tmp) throws
Exception {
+ byte[] exe = readTestResource(EXE);
+ Map<String, byte[]> inputs = new LinkedHashMap<>();
+ inputs.put("exe", exe);
+ inputs.put("dll", readTestResource(DLL));
+ inputs.put("language variants", readTestResource(LANG_DLL));
+ inputs.put("cycle", patched(exe, 0x3414, 0x80000000));
+ inputs.put("data outside the section", patched(exe, 0x3530, 0x1000));
+ inputs.put("images sharing bytes",
+ patched(exe, 0x3570, EndianUtils.getIntLE(exe, 0x3560) + 1));
+ inputs.put("name with a slash", patched(exe, 0x3528, '/'));
+ for (int length : new int[]{0x200, 0x3000, 0x3400, 0x4000, 0x7b20,
0x7b50, 0x7fff}) {
+ inputs.put("cut at 0x" + Integer.toHexString(length),
Arrays.copyOf(exe, length));
+ }
+ inputs.put("output budget", sharedImagePe(40, 4096));
+ inputs.put("entry budget", manyIconsPe());
+
+ for (Map.Entry<String, byte[]> input : inputs.entrySet()) {
+ String label = input.getKey();
+ RecordingExtractor stream = parse(input.getValue());
+ RecordingExtractor file = parseFile(
+ Files.write(Files.createTempFile(tmp, "pe", ".exe"),
input.getValue()));
+ assertEquals(stream.contents.size(), file.contents.size(), label);
+ for (int i = 0; i < stream.contents.size(); i++) {
+ assertArrayEquals(stream.contents.get(i),
file.contents.get(i), label);
+ }
+ assertEquals(stream.metadata.size(), file.metadata.size(), label);
+ for (int i = 0; i < stream.metadata.size(); i++) {
+ assertEquals(stream.metadata.get(i), file.metadata.get(i),
label);
+ }
+ assertEquals(warning(stream), warning(file), label);
+ // the messages differ, what was noticed must not
+ assertEquals(stream.parentMetadata.get(
+
TikaCoreProperties.TIKA_META_EXCEPTION_EMBEDDED_STREAM) == null,
+ file.parentMetadata.get(
+
TikaCoreProperties.TIKA_META_EXCEPTION_EMBEDDED_STREAM) == null,
+ label);
+ }
+ }
+
+ /**
+ * The section buffer follows what was read, not what the header declares:
+ * 64 MB declared, 100 bytes there.
+ */
+ @Test
+ public void testSectionBufferFollowsBytesRead() throws Exception {
+ int declared = 64 * 1024 * 1024;
+ PEIconExtractor.BufferedSection section = new
PEIconExtractor.BufferedSection(
+ new ByteArrayInputStream(new byte[100]), declared);
+ assertNull(section.read(declared - 16, 16));
+ assertTrue(section.buf.length <= 64 * 1024, "allocated " +
section.buf.length);
+ assertEquals(90, section.read(10, 90).length);
+ }
+
+ /**
+ * A failing source is not a broken resource section: it fails the parse
+ * instead of being noted as an embedded problem.
+ */
+ @Test
+ public void testSourceFailureSurfaces() throws Exception {
+ InputStream failing = new FilterInputStream(
+ new ByteArrayInputStream(readTestResource(EXE))) {
+ // gives out before the resource section at 0x3400
+ private long remaining = 0x3000;
+
+ @Override
+ public int read() throws IOException {
+ byte[] one = new byte[1];
+ return read(one, 0, 1) < 0 ? -1 : one[0] & 0xff;
+ }
+
+ @Override
+ public int read(byte[] b, int off, int len) throws IOException {
+ if (remaining <= 0) {
+ throw new IOException("source failed");
+ }
+ int n = super.read(b, off, (int) Math.min(len, remaining));
+ remaining -= Math.max(n, 0);
+ return n;
+ }
+
+ @Override
+ public long skip(long n) throws IOException {
+ if (remaining <= 0) {
+ throw new IOException("source failed");
+ }
+ long skipped = super.skip(Math.min(n, remaining));
+ remaining -= skipped;
+ return skipped;
+ }
+ };
+ RecordingExtractor extractor = new RecordingExtractor();
+ try (TikaInputStream tis = TikaInputStream.get(failing)) {
+ assertThrows(IOException.class, () -> new
ExecutableParser().parse(tis,
+ new BodyContentHandler(), extractor.parentMetadata,
extractor.context()));
+ }
+ assertNull(extractor.parentMetadata.get(
+ TikaCoreProperties.TIKA_META_EXCEPTION_EMBEDDED_STREAM));
+ }
+
+ /**
+ * Icon 1's data entry (0x3530) is pointed outside the resource section;
+ * group 1 loses an image and is dropped, DOCICON is unaffected.
+ */
+ @Test
+ public void testDataOutsideSectionIsSkipped() throws Exception {
+ byte[] exe = readTestResource(EXE);
+ putIntLE(exe, 0x3530, 0x1000);
+ RecordingExtractor extractor = parse(exe);
+ assertEquals(1, extractor.contents.size());
+ assertIcon(extractor.metadata.get(0), "icon_DOCICON.ico",
"14/DOCICON/1033", THUMBNAIL);
+
assertNull(extractor.parentMetadata.get(TikaCoreProperties.TIKA_META_EXCEPTION_EMBEDDED_STREAM));
+ }
+
+ /**
+ * The RT_ICON type entry (0x3414) is pointed back at the root directory.
+ */
+ @Test
+ public void testCycleInResourceTree() throws Exception {
+ byte[] exe = readTestResource(EXE);
+ putIntLE(exe, 0x3414, 0x80000000);
+ RecordingExtractor extractor = parse(exe);
+ assertEquals(0, extractor.contents.size());
+
assertNull(extractor.parentMetadata.get(TikaCoreProperties.TIKA_META_EXCEPTION_EMBEDDED_STREAM));
+ assertEquals(ExecutableParser.MACHINE_x86_32,
+ extractor.parentMetadata.get(ExecutableParser.MACHINE_TYPE));
+ }
+
+ /**
+ * Resource tree offsets are relative to the tree's root, which does not
+ * have to be the start of the section. The real .rsrc is moved 0x300
+ * bytes into a synthetic section.
+ */
+ @Test
+ public void testResourceRootNotAtSectionStart() throws Exception {
+ for (int shift : new int[]{0, 0x300}) {
+ byte[] rsrc = resourceSectionAt(SyntheticPE.SECTION_VA + shift);
+ byte[] section = new byte[shift + rsrc.length];
+ System.arraycopy(rsrc, 0, section, shift, rsrc.length);
+ RecordingExtractor extractor = parse(SyntheticPE.build(section,
shift));
+ assertEquals(2, extractor.contents.size(), "shift " + shift);
+ assertArrayEquals(readTestResource(DOC_ICO),
extractor.contents.get(0));
+ assertArrayEquals(readTestResource(APP_ICO),
extractor.contents.get(1));
+ }
+ }
+
+ /**
+ * A directory with more entries than the budget allows stops the walk
+ * with a warning instead of an exception.
+ */
+ @Test
+ public void testDirectoryEntryBudget() throws Exception {
+ int entries = 30000;
+ ByteBuffer rsrc = ByteBuffer.allocate(16 + 8 + 16 + 8 * entries)
+ .order(ByteOrder.LITTLE_ENDIAN);
+ // root: one id entry, type RT_ICON, pointing at the subdirectory at 24
+ rsrc.position(12).putShort((short) 0).putShort((short) 1);
+ rsrc.putInt(3).putInt(0x80000000 | 24);
+ // subdirectory with too many entries, none of them leading anywhere
+ rsrc.position(24 + 12).putShort((short) 0).putShort((short) entries);
+ for (int i = 0; i < entries; i++) {
+ rsrc.putInt(i).putInt(0);
+ }
+ RecordingExtractor extractor = parse(SyntheticPE.build(rsrc.array(),
0));
+ assertEquals(0, extractor.contents.size());
+ assertContains("20000 entries",
+
extractor.parentMetadata.get(TikaCoreProperties.TIKA_META_EXCEPTION_WARNING));
+ }
+
+ /**
+ * Subdirectories below the language level are malformed and must stop the
+ * walk there: a chain of 25000 nested directories under one language entry
+ * would otherwise be followed to the end and exhaust the entry budget.
+ */
+ @Test
+ public void testDepthCapBelowLanguageLevel() throws Exception {
+ int chain = 25000;
+ int dirSize = 16 + 8;
+ ByteBuffer rsrc = ByteBuffer.allocate(dirSize * (3 +
chain)).order(ByteOrder.LITTLE_ENDIAN);
+ // root -> type RT_ICON -> name 1 -> language dir, then the chain
+ for (int i = 0; i < 3 + chain; i++) {
+ int dir = i * dirSize;
+ rsrc.position(dir + 12).putShort((short) 0).putShort((short) 1);
+ rsrc.putInt(i == 0 ? 3 : 1).putInt(0x80000000 | (dir + dirSize));
+ }
+ RecordingExtractor extractor = parse(SyntheticPE.build(rsrc.array(),
0));
+ assertEquals(0, extractor.contents.size());
+
assertNull(extractor.parentMetadata.get(TikaCoreProperties.TIKA_META_EXCEPTION_WARNING));
+
assertNull(extractor.parentMetadata.get(TikaCoreProperties.TIKA_META_EXCEPTION_EMBEDDED_STREAM));
+ }
+
+ /**
+ * A file without icon types must not have its resource section read; the
+ * synthetic section declares 1 MB behind a directory that only lists
bitmaps.
+ */
+ @Test
+ public void testSectionIsReadLazily() throws Exception {
+ ByteBuffer rsrc = ByteBuffer.allocate(1024 *
1024).order(ByteOrder.LITTLE_ENDIAN);
+ rsrc.position(12).putShort((short) 0).putShort((short) 1);
+ rsrc.putInt(2).putInt(0x80000000 | 24); // RT_BITMAP
+ byte[] pe = SyntheticPE.build(rsrc.array(), 0);
+ CountingInputStream counting = new CountingInputStream(new
ByteArrayInputStream(pe));
+ RecordingExtractor extractor = new RecordingExtractor();
+ try (TikaInputStream tis = TikaInputStream.get(counting)) {
+ new ExecutableParser().parse(tis, new BodyContentHandler(), new
Metadata(),
+ extractor.context());
+ }
+ assertEquals(0, extractor.contents.size());
+ // TikaInputStream buffers in 8 KB blocks; the point is that the
megabyte stays unread
+ assertTrue(counting.getByteCount() < SyntheticPE.SECTION_RAW_PTR + 64
* 1024,
+ "read " + counting.getByteCount() + " bytes");
+ }
+
+ @Test
+ public void testExtractorMayDecline() throws Exception {
+ RecordingExtractor extractor = new RecordingExtractor();
+ extractor.accept = false;
+ parse(readTestResource(EXE), extractor);
+ assertEquals(2, extractor.metadata.size(), "asked about both groups");
+ assertEquals(0, extractor.contents.size());
+ }
+
+ /**
+ * Limits configured by the caller must surface, not be swallowed as a
+ * broken resource section.
+ */
+ @Test
+ public void testEmbeddedLimitSurfaces() throws Exception {
+ ParseContext context = new ParseContext();
+ context.set(EmbeddedLimits.class, new EmbeddedLimits(-1, false, 1,
true));
+ context.set(ParseRecord.class, ParseRecord.newInstance(context));
+ context.set(ParseTimeout.class, ParseTimeout.start(new
TimeoutLimits(60_000, 60_000)));
+ try (TikaInputStream tis = getResourceAsStream("/test-documents/" +
EXE)) {
+ assertThrows(EmbeddedLimitReachedException.class, () ->
AUTO_DETECT_PARSER
+ .parse(tis, new BodyContentHandler(), new Metadata(),
context));
+ }
+ }
+
+ @Test
+ public void testExeWithoutIcons() throws Exception {
+ List<Metadata> metadataList =
getRecursiveMetadata("testWindows-x86-32.exe");
+ assertEquals(1, metadataList.size());
+
assertNull(metadataList.get(0).get(TikaCoreProperties.TIKA_META_EXCEPTION_EMBEDDED_STREAM));
+ }
+
+ @Test
+ public void testExtractIconsDisabled() throws Exception {
+ ExecutableParser.ExecutableParserConfig noIcons =
+ new ExecutableParser.ExecutableParserConfig();
+ noIcons.setExtractIcons(false);
+
+ List<Metadata> metadataList = getRecursiveMetadata(EXE, new
ExecutableParser(noIcons));
+ assertEquals(1, metadataList.size());
+ assertEquals(ExecutableParser.MACHINE_x86_32,
+ metadataList.get(0).get(ExecutableParser.MACHINE_TYPE));
+
+ // for one parse, through the context
+ RecordingExtractor extractor = new RecordingExtractor();
+ ParseContext context = extractor.context();
+ context.set(ExecutableParser.ExecutableParserConfig.class, noIcons);
+ try (TikaInputStream tis = getResourceAsStream("/test-documents/" +
EXE)) {
+ new ExecutableParser().parse(tis, new BodyContentHandler(), new
Metadata(), context);
+ }
+ assertEquals(0, extractor.metadata.size());
+
+ // from a JSON configuration
+ Parser loaded =
TikaLoader.load(getConfigPath(PEIconExtractorTest.class,
+ "tika-config-executable-no-icons.json")).loadParsers();
+ assertEquals(1, getRecursiveMetadata(EXE, loaded).size());
+ assertEquals(3, getRecursiveMetadata(EXE).size());
+ }
+
+ /**
+ * A request to tika-server or through pipes carries its configuration as
+ * JSON under the parser's name, not as an object under its class.
+ */
+ @Test
+ public void testExtractIconsPerRequestJson() throws Exception {
+ RecordingExtractor extractor = new RecordingExtractor();
+ ParseContext context = extractor.context();
+ context.setJsonConfig("executable-parser", "{\"extractIcons\":
false}");
+ try (TikaInputStream tis = getResourceAsStream("/test-documents/" +
EXE)) {
+ new ExecutableParser().parse(tis, new BodyContentHandler(), new
Metadata(), context);
+ }
+ assertEquals(0, extractor.metadata.size());
+
+ // and the other way round, over a parser that was configured without
icons
+ ExecutableParser.ExecutableParserConfig noIcons =
+ new ExecutableParser.ExecutableParserConfig();
+ noIcons.setExtractIcons(false);
+ extractor = new RecordingExtractor();
+ context = extractor.context();
+ context.setJsonConfig("executable-parser", "{\"extractIcons\": true}");
+ try (TikaInputStream tis = getResourceAsStream("/test-documents/" +
EXE)) {
+ new ExecutableParser(noIcons).parse(tis, new BodyContentHandler(),
new Metadata(),
+ context);
+ }
+ assertEquals(2, extractor.contents.size());
+ }
+
+ /**
+ * The pre-4.2.0 entry point still yields the metadata, just no icons.
+ */
+ @Test
+ @SuppressWarnings("deprecation")
+ public void testLegacyParsePE() throws Exception {
+ Metadata metadata = new Metadata();
+ byte[] exe = readTestResource(EXE);
+ try (ByteArrayInputStream is = new ByteArrayInputStream(exe, 4,
exe.length - 4)) {
+ new ExecutableParser().parsePE(null, metadata, is,
Arrays.copyOf(exe, 4));
+ }
+ assertEquals(ExecutableParser.MACHINE_x86_32,
metadata.get(ExecutableParser.MACHINE_TYPE));
+ assertEquals("Windows", metadata.get(ExecutableParser.PLATFORM));
+ }
+
+ /**
+ * Cutting the file inside the resource section yields the icons that are
+ * still complete, silently; cutting before it is reported, but never
+ * costs the header metadata.
+ */
+ @Test
+ public void testTruncatedFile() throws Exception {
+ byte[] full = readTestResource(EXE);
+ byte[] docIco = readTestResource(DOC_ICO);
+ // length -> expected number of icons
+ int[][] cases = {
+ {0x200, 0}, // inside the section table
+ {0x3400, 0}, // resource section entirely missing
+ {0x4000, 0}, // inside icon 2, no group is complete
+ {0x7b20, 0}, // inside DOCICON's GRPICONDIR
+ {0x7b50, 1}, // inside group 1's GRPICONDIR, DOCICON is
complete
+ {0x7fff, 2}, // inside .reloc, after the resource section
+ };
+ for (int[] c : cases) {
+ int length = c[0];
+ RecordingExtractor extractor = parse(Arrays.copyOf(full, length));
+ String msg = "length 0x" + Integer.toHexString(length);
+ assertEquals(c[1], extractor.contents.size(), msg);
+ assertEquals(ExecutableParser.MACHINE_x86_32,
+
extractor.parentMetadata.get(ExecutableParser.MACHINE_TYPE), msg);
+ String recorded = extractor.parentMetadata.get(
+ TikaCoreProperties.TIKA_META_EXCEPTION_EMBEDDED_STREAM);
+ if (length < 0x3400) {
+ assertNotNull(recorded, msg);
+ assertContains("EOFException", recorded);
+ } else {
+ assertNull(recorded, msg);
+ }
+ if (c[1] >= 1) {
+ assertArrayEquals(docIco, extractor.contents.get(0), msg);
+ assertIcon(extractor.metadata.get(0), "icon_DOCICON.ico",
"14/DOCICON/1033",
+ THUMBNAIL);
+ }
+ }
+ }
+
+ private static void assertIcon(Metadata metadata, String name, String
relationshipId,
+ String resourceType) {
+ assertEquals(name, metadata.get(TikaCoreProperties.RESOURCE_NAME_KEY));
+ assertEquals("true",
metadata.get(TikaCoreProperties.RESOURCE_NAME_EXTENSION_INFERRED));
+ assertEquals(PEIconExtractor.ICON_MIME_TYPE,
metadata.get(HttpHeaders.CONTENT_TYPE));
+ assertEquals(relationshipId,
metadata.get(TikaCoreProperties.EMBEDDED_RELATIONSHIP_ID));
+ assertEquals(resourceType,
metadata.get(TikaCoreProperties.EMBEDDED_RESOURCE_TYPE));
+ }
+
+ /**
+ * @return a GRPICONDIR that lists the icons with the given ids
+ */
+ private static byte[] grpIconDir(int... ids) {
+ ByteBuffer dir = ByteBuffer.allocate(6 + 14 *
ids.length).order(ByteOrder.LITTLE_ENDIAN);
+ dir.putShort((short) 0).putShort((short) 1).putShort((short)
ids.length);
+ for (int i = 0; i < ids.length; i++) {
+ dir.putShort(6 + 14 * i + 12, (short) ids[i]);
+ }
+ return dir.array();
+ }
+
+ /**
+ * @return a file in which the given number of groups all list one image
+ */
+ private static byte[] sharedImagePe(int groups, int imageSize) {
+ byte[] data = new byte[20 + imageSize];
+ System.arraycopy(grpIconDir(1), 0, data, 0, 20);
+ SyntheticResources resources = new SyntheticResources(data);
+ for (int i = 0; i < groups; i++) {
+ resources.group(0, 20);
+ }
+ resources.icon(20, imageSize);
+ return SyntheticPE.build(resources.build(), 0);
+ }
+
+ /**
+ * @return a file with one group and 12000 icons, more than the entry
+ * budget of 20000 lets the walk visit
+ */
+ private static byte[] manyIconsPe() {
+ byte[] data = new byte[20 + 16];
+ System.arraycopy(grpIconDir(1), 0, data, 0, 20);
+ SyntheticResources resources = new SyntheticResources(data);
+ resources.group(0, 20);
+ for (int i = 0; i < 12000; i++) {
+ resources.icon(20, 16);
+ }
+ return SyntheticPE.build(resources.build(), 0);
+ }
+
+ /**
+ * @return the resource section of the test EXE, its seven data entries
+ * rebased from RVA 0xa000 to the given one
+ */
+ private byte[] resourceSectionAt(int rva) throws IOException {
+ byte[] rsrc = Arrays.copyOfRange(readTestResource(EXE), 0x3400,
0x7c00);
+ for (int entry = 0x130; entry <= 0x190; entry += 16) {
+ putIntLE(rsrc, entry, EndianUtils.getIntLE(rsrc, entry) - 0xa000 +
rva);
+ }
+ return rsrc;
+ }
+
+ /**
+ * @return the warning's message without the stack trace that follows it
+ */
+ private static String warning(RecordingExtractor extractor) {
+ String warning = extractor.parentMetadata.get(
+ TikaCoreProperties.TIKA_META_EXCEPTION_WARNING);
+ return warning == null ? null : warning.lines().findFirst().orElse("");
+ }
+
+ private static byte[] patched(byte[] data, int offset, int value) {
+ byte[] patched = data.clone();
+ putIntLE(patched, offset, value);
+ return patched;
+ }
+
+ private static void putIntLE(byte[] data, int offset, int value) {
+ ByteBuffer.wrap(data, offset,
4).order(ByteOrder.LITTLE_ENDIAN).putInt(value);
+ }
+
+ private byte[] readTestResource(String name) throws IOException {
+ try (TikaInputStream is = getResourceAsStream("/test-documents/" +
name)) {
+ return is.readAllBytes();
+ }
+ }
+
+ private static RecordingExtractor parse(byte[] pe) throws Exception {
+ return parse(pe, new RecordingExtractor());
+ }
+
+ private static RecordingExtractor parse(byte[] pe, RecordingExtractor
extractor)
+ throws Exception {
+ try (TikaInputStream tis = TikaInputStream.get(pe)) {
+ new ExecutableParser().parse(tis, new BodyContentHandler(),
extractor.parentMetadata,
+ extractor.context());
+ }
+ return extractor;
+ }
+
+ private static RecordingExtractor parseFile(Path file) throws Exception {
+ RecordingExtractor extractor = new RecordingExtractor();
+ try (TikaInputStream tis = TikaInputStream.get(file)) {
+ new ExecutableParser().parse(tis, new BodyContentHandler(),
extractor.parentMetadata,
+ extractor.context());
+ }
+ return extractor;
+ }
+
+ /**
+ * Records what the parser hands over, before any embedded parsing.
+ */
+ private static class RecordingExtractor implements
EmbeddedDocumentExtractor {
+ final List<byte[]> contents = new ArrayList<>();
+ final List<Metadata> metadata = new ArrayList<>();
+ final Metadata parentMetadata = new Metadata();
+ boolean accept = true;
+
+ ParseContext context() {
+ ParseContext context = new ParseContext();
+ context.set(EmbeddedDocumentExtractor.class, this);
+ return context;
+ }
+
+ @Override
+ public boolean shouldParseEmbedded(Metadata metadata, ParseContext
context) {
+ this.metadata.add(metadata);
+ return accept;
+ }
+
+ @Override
+ public void parseEmbedded(TikaInputStream stream, ContentHandler
handler,
+ Metadata metadata, ParseContext context,
boolean outputHtml)
+ throws IOException {
+ contents.add(stream.readAllBytes());
+ }
+ }
+
+ /**
+ * A resource section whose icons and icon groups are numbered from 1 in
+ * the order they are added, all in language 1033, their data being ranges
+ * of one block that follows the tree.
+ */
+ private static final class SyntheticResources {
+ private static final int DIR_SIZE = 16 + 8;
+ private final List<int[]> icons = new ArrayList<>();
+ private final List<int[]> groups = new ArrayList<>();
+ private final byte[] data;
+
+ SyntheticResources(byte[] data) {
+ this.data = data;
+ }
+
+ void icon(int offset, int size) {
+ icons.add(new int[]{offset, size});
+ }
+
+ void group(int offset, int size) {
+ groups.add(new int[]{offset, size});
+ }
+
+ byte[] build() {
+ int iconTypeDir = 16 + 2 * 8;
+ int iconLangDirs = iconTypeDir + 16 + 8 * icons.size();
+ int groupTypeDir = iconLangDirs + DIR_SIZE * icons.size();
+ int groupLangDirs = groupTypeDir + 16 + 8 * groups.size();
+ int dataEntries = groupLangDirs + DIR_SIZE * groups.size();
+ int block = dataEntries + 16 * (icons.size() + groups.size());
+ ByteBuffer rsrc = ByteBuffer.allocate(block + data.length)
+ .order(ByteOrder.LITTLE_ENDIAN);
+ rsrc.position(12).putShort((short) 0).putShort((short) 2);
+ rsrc.putInt(3).putInt(0x80000000 | iconTypeDir);
+ rsrc.putInt(14).putInt(0x80000000 | groupTypeDir);
+ writeType(rsrc, iconTypeDir, iconLangDirs, dataEntries, icons,
block);
+ writeType(rsrc, groupTypeDir, groupLangDirs, dataEntries + 16 *
icons.size(), groups,
+ block);
+ rsrc.position(block);
+ rsrc.put(data);
+ return rsrc.array();
+ }
+
+ private static void writeType(ByteBuffer rsrc, int typeDir, int
langDirs, int dataEntries,
+ List<int[]> resources, int block) {
+ rsrc.position(typeDir + 12).putShort((short) 0).putShort((short)
resources.size());
+ for (int i = 0; i < resources.size(); i++) {
+ rsrc.putInt(i + 1).putInt(0x80000000 | (langDirs + i *
DIR_SIZE));
+ }
+ for (int i = 0; i < resources.size(); i++) {
+ rsrc.position(langDirs + i * DIR_SIZE + 12).putShort((short)
0).putShort((short) 1);
+ rsrc.putInt(1033).putInt(dataEntries + i * 16);
+ rsrc.position(dataEntries + i * 16);
+ rsrc.putInt(SyntheticPE.SECTION_VA + block +
resources.get(i)[0]);
+ rsrc.putInt(resources.get(i)[1]);
+ }
+ }
+ }
+
+ /**
+ * The smallest PE32 file that gets a resource section past
ExecutableParser:
+ * DOS stub, PE signature, COFF header, full optional header, one .rsrc
+ * section header, and the section itself at {@link #SECTION_RAW_PTR}.
+ */
+ private static final class SyntheticPE {
+ static final int SECTION_VA = 0x1000;
+ static final int SECTION_RAW_PTR = 0x200;
+ private static final int PE_OFFSET = 0x40;
+ private static final int OPT_HEADER_SIZE = 224;
+
+ static byte[] build(byte[] section, int rootOffset) {
+ return build(section, rootOffset, section.length);
+ }
+
+ /**
+ * @param declaredSize the section size the header claims, whatever
the section holds
+ */
+ static byte[] build(byte[] section, int rootOffset, int declaredSize) {
+ ByteBuffer pe = ByteBuffer.allocate(SECTION_RAW_PTR +
section.length)
+ .order(ByteOrder.LITTLE_ENDIAN);
+ pe.put((byte) 'M').put((byte) 'Z');
+ pe.putInt(0x3c, PE_OFFSET);
+ pe.position(PE_OFFSET);
+ pe.put("PE\0\0".getBytes(StandardCharsets.US_ASCII));
+ pe.putShort((short) 0x14c).putShort((short) 1); // machine, one
section
+ pe.putInt(0).putInt(0).putInt(0); // timestamp,
symbols
+ pe.putShort((short) OPT_HEADER_SIZE).putShort((short) 0x102);
+ int opt = pe.position();
+ pe.putShort((short) 0x10b); // PE32
+ pe.putInt(opt + 92, 16); //
NumberOfRvaAndSizes
+ pe.putInt(opt + 96 + 2 * 8, SECTION_VA + rootOffset);
+ pe.putInt(opt + 96 + 2 * 8 + 4, declaredSize - rootOffset);
+ int sec = opt + OPT_HEADER_SIZE;
+ pe.position(sec);
+ pe.put(".rsrc\0\0\0".getBytes(StandardCharsets.US_ASCII));
+ pe.putInt(declaredSize).putInt(SECTION_VA);
+ pe.putInt(declaredSize).putInt(SECTION_RAW_PTR);
+ pe.position(SECTION_RAW_PTR);
+ pe.put(section);
+ return pe.array();
+ }
+ }
+}
diff --git
a/tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-code-module/src/test/resources/configs/tika-config-executable-no-icons.json
b/tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-code-module/src/test/resources/configs/tika-config-executable-no-icons.json
new file mode 100644
index 0000000000..cca9978a20
--- /dev/null
+++
b/tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-code-module/src/test/resources/configs/tika-config-executable-no-icons.json
@@ -0,0 +1,10 @@
+{
+ "parsers": [
+ "default-parser",
+ {
+ "executable-parser": {
+ "extractIcons": false
+ }
+ }
+ ]
+}
diff --git
a/tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-code-module/src/test/resources/test-documents/testWindows-icons-app.ico
b/tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-code-module/src/test/resources/test-documents/testWindows-icons-app.ico
new file mode 100644
index 0000000000..ab59ba78bc
Binary files /dev/null and
b/tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-code-module/src/test/resources/test-documents/testWindows-icons-app.ico
differ
diff --git
a/tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-code-module/src/test/resources/test-documents/testWindows-icons-doc.ico
b/tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-code-module/src/test/resources/test-documents/testWindows-icons-doc.ico
new file mode 100644
index 0000000000..5baef96f69
Binary files /dev/null and
b/tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-code-module/src/test/resources/test-documents/testWindows-icons-doc.ico
differ
diff --git
a/tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-code-module/src/test/resources/test-documents/testWindows-x86-32-icons.exe
b/tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-code-module/src/test/resources/test-documents/testWindows-x86-32-icons.exe
new file mode 100755
index 0000000000..816985eb0f
Binary files /dev/null and
b/tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-code-module/src/test/resources/test-documents/testWindows-x86-32-icons.exe
differ
diff --git
a/tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-code-module/src/test/resources/test-documents/testWindows-x86-64-icons-lang.dll
b/tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-code-module/src/test/resources/test-documents/testWindows-x86-64-icons-lang.dll
new file mode 100755
index 0000000000..d40a4d388c
Binary files /dev/null and
b/tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-code-module/src/test/resources/test-documents/testWindows-x86-64-icons-lang.dll
differ
diff --git
a/tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-code-module/src/test/resources/test-documents/testWindows-x86-64-icons.dll
b/tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-code-module/src/test/resources/test-documents/testWindows-x86-64-icons.dll
new file mode 100755
index 0000000000..4220a2fc4f
Binary files /dev/null and
b/tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-code-module/src/test/resources/test-documents/testWindows-x86-64-icons.dll
differ