[
https://issues.apache.org/jira/browse/TIKA-4936?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18119879#comment-18119879
]
ASF GitHub Bot commented on TIKA-4936:
--------------------------------------
dschmidt opened a new pull request, #3263:
URL: https://github.com/apache/tika/pull/3263
https://issues.apache.org/jira/browse/TIKA-4936
`ExecutableParser` so far only read the COFF header of PE files. This PR
makes it continue through the optional header and section table to the resource
section and rebuild every `RT_GROUP_ICON` together with its `RT_ICON` images
into a standalone `.ico` file, which is handed to the
`EmbeddedDocumentExtractor` as `image/vnd.microsoft.icon`.
- The first group in resource order (the icon Windows shows for the file) is
marked `THUMBNAIL`, the others `ATTACHMENT`; `embeddedRelationshipId` carries
`type/name/language` of the resource. Resource names are
`icon_<name-or-id>.ico`, with a language suffix only when a group has several
language variants.
- Single `RT_ICON` entries are deliberately not emitted on their own:
BMP-encoded entries are not valid standalone images and would duplicate the
group data.
- The extractor reads strictly forward (no seeking), bounds-checks every
offset, caps tree depth, entry count and section size, detects cycles and
treats a truncated resource section as "fewer icons" rather than an error.
Failures in the resource section are recorded as an embedded stream exception
and never cost the header metadata.
- New `extractIcons` property on `ExecutableParser` (default `true`) to turn
the feature off.
- `parsePE()` now takes `TikaInputStream` and `ParseContext`; it was only
called internally.
Test fixtures are a MinGW-built PE32 EXE and PE32+ DLL with two icon groups
(numeric id `1` and named `DOCICON`; BMP 16/32/48 px and PNG 256 px images)
plus the source `.ico` files, so the tests check that the rebuilt icons are
byte identical. Further tests cover an EXE without icons, `extractIcons=false`
and truncation at several offsets inside the resource section.
Not included, possible follow-ups: `RT_GROUP_CURSOR`/`RT_CURSOR` as `.cur`,
`RT_MANIFEST` as embedded XML, `VS_VERSIONINFO` as metadata. `CHANGES.txt` not
touched since there is no unreleased section yet.
Module `verify` (tests, checkstyle, forbiddenapis, rat) passes on
`tika-parser-code-module`.
> Extract icons from PE executables (EXE/DLL) as embedded documents
> -----------------------------------------------------------------
>
> Key: TIKA-4936
> URL: https://issues.apache.org/jira/browse/TIKA-4936
> Project: Tika
> Issue Type: New Feature
> Environment:
>
>
>
> Reporter: Dominik Schmidt
> Priority: Major
>
> h3. Background
> {{ExecutableParser}} currently only reads the COFF file header of PE files
> (EXE/DLL) and emits basic metadata (machine type, architecture bits,
> endianness, created date). The resource section ({{.rsrc}}) is not parsed, so
> resources such as the application icon are not accessible via Tika.
> h3. Proposal
> Parse the PE resource directory and emit each icon group as an embedded
> document:
> * Parse optional header, data directories and section table to locate the
> resource directory (RVA → file offset).
> * Walk the resource tree (type → name/ID → language).
> * For each {{RT_GROUP_ICON}} (type 14), reconstruct a standalone {{.ico}}
> file from the {{GRPICONDIR}} and the referenced {{RT_ICON}} (type 3) entries:
> write an {{ICONDIR}} header and replace the 2-byte resource IDs ({{nID}})
> with 4-byte image offsets.
> * Pass each reconstructed file to the {{EmbeddedDocumentExtractor}} with:
> ** {{Content-Type}}: {{image/vnd.microsoft.icon}}
> ** {{resourceName}}: e.g. {{icon_<id-or-name>.ico}}
> ** {{embeddedResourceType}}: {{THUMBNAIL}} for the first icon group (the icon
> shown by Windows Explorer), {{ATTACHMENT}} for all others
> ** the resource language ID
> Single {{RT_ICON}} entries are intentionally not emitted on their own:
> BMP-based entries are not valid standalone images (no {{BITMAPFILEHEADER}},
> double height for the AND mask), and emitting them would duplicate the group
> data.
> h3. Robustness
> The parser must handle malformed or malicious binaries gracefully:
> * bound the resource tree depth (normally 3 levels) and the number of entries
> * validate all offsets and sizes against the file size
> * guard against cycles in the resource directory
> * failures while extracting resources must not break the existing metadata
> extraction
> h3. Out of scope (possible follow-ups)
> * {{RT_GROUP_CURSOR}}/{{RT_CURSOR}} → {{.cur}}
> * {{RT_MANIFEST}} as an embedded XML document
> * {{VS_VERSIONINFO}} (product name, file version, company) as metadata
> h3. Acceptance criteria
> * Icons of 32- and 64-bit EXE and DLL test files are extracted as valid
> {{.ico}} files that are detected as {{image/vnd.microsoft.icon}}.
> * Icon groups containing both PNG- and BMP-encoded entries are supported.
> * Files without a resource section, or without icons, are parsed as before,
> with no embedded documents.
> * Truncated or corrupted resource sections do not throw and still yield the
> existing metadata.
> * Unit tests use small, license-compatible test files.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)