[
https://issues.apache.org/jira/browse/CAMEL-25356?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
shashank reassigned CAMEL-25356:
--------------------------------
Assignee: shashank
> camel-cm-sms - messages with <, >, ¤ or a form feed are sent as unicode: the
> GSM 03.38 regex contains HTML entities
> -------------------------------------------------------------------------------------------------------------------
>
> Key: CAMEL-25356
> URL: https://issues.apache.org/jira/browse/CAMEL-25356
> Project: Camel
> Issue Type: Bug
> Components: camel-cm-sms
> Reporter: shashank
> Assignee: shashank
> Priority: Minor
>
> {{CMConstants.GSM_0338_REGEX}} was copied from a web page together with its
> HTML entities:
> {code:java}
> "^[A-Za-z0-9
> \\r\\n@£$Δ_...!\"#$%&'()*+,\\-./:;<=>?¡¿^{}\\\\\\[~\\]|" +
> "€¥...]*$"
> {code}
> In a Java character class {{<}} is the four characters {{&}}, {{l}},
> {{t}}, {{;}}: {{<}} and {{>}} are not in the class. The currency sign {{¤}}
> (0x24 of the GSM default alphabet) and the form feed of the extension table
> are missing too. {{CMMessage.setUnicodeAndMultipart}} then sends any message
> with one of these characters as unicode ({{<DCS>8</DCS>}}): 70 / 67
> characters per part instead of 160 / 153, and the maximum number of parts is
> computed for unicode (a 310-character message needs 5 parts instead of 3, and
> a message of more than 536 characters exceeds the default maximum of 8
> unicode parts, while it needs 4 GSM parts). GSM 03.38 (3GPP TS 23.038; the
> mapping https://unicode.org/Public/MAPPINGS/ETSI/GSM0338.TXT, cited in the
> new test) has these characters, and the {{CMMessage}} javadoc promises 160 /
> 153 characters per part for a message in GSM 7-bit characters.
> h3. Reproduction
> New {{CMGsm0338Test}}: every character of the GSM basic set and of the
> extension table checked with {{CMUtils.isGsm0338Encodeable}}: main reports
> {{U+00A4 U+003C U+003E}} and {{U+000C}} as not GSM; a 310-character text with
> {{<}} and {{>}} is sent as unicode; Cyrillic, {{ê}} and a backtick are the
> control (not GSM). Two runs on main.
> h3. Proposed fix
> Write the characters themselves in the class (and drop the duplicated {{$}}).
> Every message that matched before still matches. Module: 55 tests pass.
> Found with a Lean 4 model of the character class against the GSM 03.38 table:
> {{main_missing}} computes that exactly {{¤ < >}} and form feed are missing,
> {{main_sound}} that main accepts nothing outside GSM 03.38, {{fix_complete}}
> / {{fix_sound}} that the class of the fix is exactly GSM 03.38,
> {{fix_extends_main}} that no GSM message of main becomes unicode.
> Not in scope: the extension characters count as two septets in GSM 03.38; the
> part count uses {{String.length()}} (unchanged).
> Affected: 4.14.x, 4.18.x and main (the regex dates from the component, 2016).
> Duplicate check (2026-10-04): JIRA component camel-cm-sms (2 issues),
> "cm-sms" with unicode: none. GitHub pull requests "cm-sms": dependency bumps
> only.
> _Filed with Claude Code on behalf of allthingssecurity._
--
This message was sent by Atlassian Jira
(v8.20.10#820010)