[ 
https://issues.apache.org/jira/browse/CAMEL-25356?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

shashank reassigned CAMEL-25356:
--------------------------------

    Assignee: shashank

> camel-cm-sms - messages with <, >, ¤ or a form feed are sent as unicode: the 
> GSM 03.38 regex contains HTML entities
> -------------------------------------------------------------------------------------------------------------------
>
>                 Key: CAMEL-25356
>                 URL: https://issues.apache.org/jira/browse/CAMEL-25356
>             Project: Camel
>          Issue Type: Bug
>          Components: camel-cm-sms
>            Reporter: shashank
>            Assignee: shashank
>            Priority: Minor
>
> {{CMConstants.GSM_0338_REGEX}} was copied from a web page together with its 
> HTML entities:
> {code:java}
> "^[A-Za-z0-9 
> \\r\\n@£$Δ_...!\"#$%&amp;'()*+,\\-./:;&lt;=&gt;?¡¿^{}\\\\\\[~\\]|" + 
> "€¥...]*$"
> {code}
> In a Java character class {{&lt;}} is the four characters {{&}}, {{l}}, 
> {{t}}, {{;}}: {{<}} and {{>}} are not in the class. The currency sign {{¤}} 
> (0x24 of the GSM default alphabet) and the form feed of the extension table 
> are missing too. {{CMMessage.setUnicodeAndMultipart}} then sends any message 
> with one of these characters as unicode ({{<DCS>8</DCS>}}): 70 / 67 
> characters per part instead of 160 / 153, and the maximum number of parts is 
> computed for unicode (a 310-character message needs 5 parts instead of 3, and 
> a message of more than 536 characters exceeds the default maximum of 8 
> unicode parts, while it needs 4 GSM parts). GSM 03.38 (3GPP TS 23.038; the 
> mapping https://unicode.org/Public/MAPPINGS/ETSI/GSM0338.TXT, cited in the 
> new test) has these characters, and the {{CMMessage}} javadoc promises 160 / 
> 153 characters per part for a message in GSM 7-bit characters.
> h3. Reproduction
> New {{CMGsm0338Test}}: every character of the GSM basic set and of the 
> extension table checked with {{CMUtils.isGsm0338Encodeable}}: main reports 
> {{U+00A4 U+003C U+003E}} and {{U+000C}} as not GSM; a 310-character text with 
> {{<}} and {{>}} is sent as unicode; Cyrillic, {{ê}} and a backtick are the 
> control (not GSM). Two runs on main.
> h3. Proposed fix
> Write the characters themselves in the class (and drop the duplicated {{$}}). 
> Every message that matched before still matches. Module: 55 tests pass.
> Found with a Lean 4 model of the character class against the GSM 03.38 table: 
> {{main_missing}} computes that exactly {{¤ < >}} and form feed are missing, 
> {{main_sound}} that main accepts nothing outside GSM 03.38, {{fix_complete}} 
> / {{fix_sound}} that the class of the fix is exactly GSM 03.38, 
> {{fix_extends_main}} that no GSM message of main becomes unicode.
> Not in scope: the extension characters count as two septets in GSM 03.38; the 
> part count uses {{String.length()}} (unchanged).
> Affected: 4.14.x, 4.18.x and main (the regex dates from the component, 2016).
> Duplicate check (2026-10-04): JIRA component camel-cm-sms (2 issues), 
> "cm-sms" with unicode: none. GitHub pull requests "cm-sms": dependency bumps 
> only.
> _Filed with Claude Code on behalf of allthingssecurity._



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to