[
https://issues.apache.org/jira/browse/TIKA-4902?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18118739#comment-18118739
]
ASF GitHub Bot commented on TIKA-4902:
--------------------------------------
Copilot commented on code in PR #3227:
URL: https://github.com/apache/tika/pull/3227#discussion_r4091930422
##########
tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-image-module/src/test/java/org/apache/tika/parser/image/ImageMetadataExtractorTest.java:
##########
@@ -193,6 +193,18 @@ public String getDescription(int tagType) {
assertEquals("0.0, 1.0", metadata.get(ImageMetadataExtractor.ICC_NS +
"Red TRC"));
}
+ @Test
+ public void testIccCurveFormattingIgnoresDefaultLocale() {
+ Locale defaultLocale = Locale.getDefault();
+ try {
+ Locale.setDefault(Locale.GERMANY);
Review Comment:
This test changes the JVM-wide default locale while this module enables
JUnit parallel execution (`src/test/resources/junit-platform.properties:17`). A
concurrently running test can observe `Locale.GERMANY` (or restore the wrong
value), causing nondeterministic failures; protect this method with
`@ResourceLock(Resources.LOCALE)` or isolate the test class.
##########
tika-parsers/tika-parsers-standard/tika-parsers-standard-modules/tika-parser-image-module/src/main/java/org/apache/tika/parser/image/ImageMetadataExtractor.java:
##########
@@ -432,6 +438,32 @@ private static boolean isIccCurveOrLut(Directory
directory, int tagType) {
return false;
}
}
+
+ // metadata-extractor 2.21.0 formats curv in the default locale: "0,"
under de_DE
+ static String formatIccCurve(byte[] data) {
+ if (data.length < 12) {
+ return null;
+ }
+ ByteBuffer b = ByteBuffer.wrap(data);
+ long count = Integer.toUnsignedLong(b.getInt(8));
+ if (count > (data.length - 12) / 2) {
+ return null;
+ }
+ StringBuilder sb = new StringBuilder();
+ for (int i = 0; i < count; i++) {
+ if (i > 0) {
+ sb.append(", ");
+ }
+ String v = String.format(Locale.ROOT, "%.7f",
+ (b.getShort(12 + i * 2) & 0xffff) / 65535.0);
Review Comment:
ICC `curv` tags have special encodings for `count == 0` (identity) and
`count == 1` (a single u8Fixed8 value), but this loop treats every sample as a
uint16 normalized by 65535. As a result, identity curves are emitted as an
empty string and one-point curves are numerically wrong (for example, `0x0100`
becomes `0.0039063` instead of `1.0`). Handle those two counts according to the
ICC curveType format before the multi-entry loop, and add cases covering them.
> Do not extract ICC curve/lookup-table tags by default in image parsers
> ----------------------------------------------------------------------
>
> Key: TIKA-4902
> URL: https://issues.apache.org/jira/browse/TIKA-4902
> Project: Tika
> Issue Type: Improvement
> Reporter: Tim Allison
> Priority: Trivial
>
> Tika dumps ~33kb curv/trc numbers into icc:*. I can't imagine there's a huge
> use case for that information. The actual iteration in metadata-extractor is
> expensive...beyond the bloated metadata. So the right fix is not to tell
> people to add a metadata filter.
> Let's avoid iterating that info but leave a parameter for users to turn it
> back on if they need that metadata.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)