https://bz.apache.org/SpamAssassin/show_bug.cgi?id=8429
Bug ID: 8429
Summary: Text-based handlers see raw bytes even when
normalize_charset is on
Product: Spamassassin
Version: 4.0.3
Hardware: PC
OS: Mac OS X
Status: NEW
Severity: normal
Priority: P2
Component: Libraries
Assignee: [email protected]
Reporter: [email protected]
Target Milestone: Undefined
A UTF-16 .js file reached the JavaScript handler NUL-interleaved, because the
JavaScript handler called decode() instead of decode_and_normalize(), as the
HTML handler already did. The same applies to the other text-based handlers
(SVG & ICS).
_normalize() had to be expanded to support an undefined charset, because
handlers can return parts with no charset (e.g. a .js file inside a RAR
archive). decode_and_normalize() now passes an undefined charset through
instead of substituting us-ascii (the us-ascii default is kept for
normalize_charset 0).
Changes:
- Handlers that process text (JavaScript, SVG, ICS) now call
decode_and_normalize() and encode the result back to UTF-8 bytes for rules,
URIs and child parts. Binary handlers (PDF, Image, Archive) still call
decode(). SVG sets utf8_mode when given bytes, as the HTML handler does, so
entities decode to UTF-8.
- _normalize() accepts an undefined charset. For undeclared or declared-UTF-16
data it calls detect_utf16() once. Undeclared data that isn't UTF-16 goes to
the existing fallbacks.
- detect_utf16() now decides both whether data is UTF-16 and its byte order,
from the first 1024 bytes: a BOM, or NULs in at least a quarter of the code
units at one byte position and few at the other. Otherwise it returns undef.
Previously it assumed its input was UTF-16 and reported UTF-16BE even for plain
ASCII.
- A part declared as UTF-16 (charset=UTF-16, UTF-16LE or UTF-16BE) with no BOM
or NUL pattern is decoded using the declared byte order, or as big-endian for
plain UTF-16 (RFC 2781). Previously a declared LE/BE was ignored.
- Mail::SpamAssassin::HTML now passes <script> text and non-base64 data: URI
content to child parts as UTF-8 bytes, not Perl characters.
Behavior change:
- Parts with no declared charset that are UTF-16 (BOM or NUL pattern) are now
decoded. Undeclared UTF-16 with almost no ASCII and no BOM (e.g. CJK) is still
not detected.
- A declared UTF-16 part with no BOM or NUL pattern is decoded using the
declared byte order, or as big-endian. Previously the byte order was guessed
heuristically, or the part fell through to the fallbacks.
- With normalize_charset 0, nothing changes: handlers still see the part's
original bytes.
Tests: t/node_utf16.t (detection, stray NULs, the 1024-byte window, declared
UTF-16 charsets), t/handler_javascript_utf16.t, t/handler_decode_normalize.t.
The patch includes their fixtures: t/data/nice/handler_archive_utf16_bom,
handler_archive_utf16_nobom, handler_javascript_utf16_attach,
handler_javascript_html_utf8, handler_svg_utf8, handler_svg_cp1251,
handler_svg_entity and handler_ics_cp1251.
--
You are receiving this mail because:
You are the assignee for the bug.