[ 
https://issues.apache.org/jira/browse/CAMEL-25303?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Claus Ibsen resolved CAMEL-25303.
---------------------------------
    Resolution: Fixed

> camel-barcode - text outside ISO-8859-1 comes back as '?' (the documented 
> default encoding UTF-8 is not used), and the hints added before the route 
> starts are dropped
> ----------------------------------------------------------------------------------------------------------------------------------------------------------------------
>
>                 Key: CAMEL-25303
>                 URL: https://issues.apache.org/jira/browse/CAMEL-25303
>             Project: Camel
>          Issue Type: Bug
>          Components: camel-barcode
>            Reporter: shashank
>            Assignee: shashank
>            Priority: Major
>             Fix For: 4.23.0
>
>
> h3. 1. Text that ISO-8859-1 cannot represent is lost
> {{BarcodeDataFormat}} gives ZXing no {{EncodeHintType.CHARACTER_SET}}. ZXing 
> then writes the text in ISO-8859-1 (QR code, Aztec; PDF417 refuses such 
> text), replacing every other character by {{?}}. With the default data format:
> * {{日本語のテキスト}} is read back as {{????????}};
> * {{Привет, мир}} as {{??????, ???}} (QR code and Aztec);
> * PDF417 fails with {{WriterException: Non-encodable character detected: П}}.
> The documentation of the data format lists "encoding: UTF-8" among the 
> default values. Text that ISO-8859-1 can represent ({{Grüße aus Köln}}) comes 
> back.
> h3. 2. The hints added before the route starts are dropped (regression in 
> 4.15.0)
> Since CAMEL-22354 the hints are computed in {{doStart()}} by 
> {{optimizeHints()}}, which starts with {{writerHintMap.clear(); 
> readerHintMap.clear();}}. The documented way to configure ZXing is 
> {{addToHintMap(...)}} on the data format instance, which a route does before 
> the route (and so the data format) starts; those hints are cleared and never 
> used. Before 4.15.0 (CAMEL-22354's fix version; so 4.18.x LTS is affected, 
> 4.14.x is not) the hints were computed in the constructors and setters, so 
> hints added afterwards were kept (until a later setter call, CAMEL-7870). 
> This also means the CHARACTER_SET hint cannot be used as a workaround for 1 
> in a route.
> h3. Reproduction
> New {{BarcodeDataFormatCharsetTest}} (marshal and unmarshal through the data 
> format API): Japanese and Cyrillic text in a QR code, Cyrillic in an Aztec 
> code and in a PDF417 code: fail on main; {{testHintsAddedBeforeStart}}: a 
> {{CHARACTER_SET}} and a {{PURE_BARCODE}} hint added before {{start()}} are 
> gone after start ({{expected: <UTF-8> but was: <null>}}); 
> {{testCharacterSetHintIsUsed}}: with a {{CHARACTER_SET=UTF-8}} hint added 
> before start, {{Grüße}} is still written in ISO-8859-1 (byte segment of 5 
> bytes instead of 7). Control: ASCII and ISO-8859-1 text round trip on main 
> and with the fix.
> h3. Proposed fix
> * When no {{CHARACTER_SET}} hint is given and the text has a character that 
> ISO-8859-1 cannot represent, encode it as UTF-8 (ZXing writes an ECI, which 
> the readers use). Text that ISO-8859-1 can represent is encoded exactly as 
> today (no hint, so no ECI: ZXing 3.5.4's QR encoder appends an ECI only when 
> a {{CHARACTER_SET}} hint is present), so existing barcodes do not change. 
> This applies to the formats whose ZXing writers take {{CHARACTER_SET}}: QR 
> code, Aztec, PDF417; the Data Matrix writer uses it only with 
> {{DATA_MATRIX_COMPACT}}, the linear formats never (no change for those).
> * Keep the hints added with {{addToHintMap}} apart and apply them on top of 
> the default hints whenever the hints are computed ({{removeFromHintMap}} 
> removes them from both), so CAMEL-7870's re-optimization still drops the 
> defaults of another format but no longer the user's hints.
> * The documentation's default encoding is corrected (ISO-8859-1, or UTF-8 
> when needed; a CHARACTER_SET hint overrides it; the formats that use it); 
> catalog copy updated.
> With the fix the 7 new tests and the module pass (33 tests).
> Found with a Lean 4 model of the text encoding and of the hint maps: the 
> property "unmarshal(marshal(text)) = text" fails for every text with a 
> character outside ISO-8859-1 (theorem {{main_loses}}) and holds with the fix 
> for every Unicode text ({{fix_roundtrip}}, using the UTF-8 round trip proved 
> for CAMEL-25126), with exactly today's bytes for ISO-8859-1 text 
> ({{fix_same_latin1}}); "a hint added before start is in force after start" 
> fails today ({{main_drops_hint}}) and holds with the fix ({{fix_keeps_hint}}).
> Affected: 1 in every version (4.14.x, 4.18.x, main; the code is the same 
> since the component was added); 2 since 4.15.0 (4.18.x and main; checked on 
> camel-4.14.x and the camel-4.15.0 tag).
> Duplicate check (2026-10-03): JIRA component camel-barcode (CAMEL-15535, 
> CAMEL-7871, CAMEL-7870, CAMEL-7866) and text "barcode" with "UTF-8" / 
> "charset" / "encoding" / "hint": none about the character set or the dropped 
> hints. GitHub pull requests "barcode": none on this (open #22358 only 
> replaces {{getOut()}} in {{readImage}}).
> _Filed with Claude Code on behalf of allthingssecurity._



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to