https://bz.apache.org/SpamAssassin/show_bug.cgi?id=8429
Kent Oyer <[email protected]> changed: What |Removed |Added ---------------------------------------------------------------------------- Attachment #6099|0 |1 is obsolete| | --- Comment #5 from Kent Oyer <[email protected]> --- Created attachment 6100 --> https://bz.apache.org/SpamAssassin/attachment.cgi?id=6100&action=edit Patch with lenient transcoding Thanks for the upvotes. There is one minor change I need to make. Malicious JavaScript is frequently obfuscated by splitting the code into chunks and concatenating them back together. With UTF-16 source code, sometimes the chunks are split in the middle of a surrogate pair. That leaves a lone surrogate on both ends. If we transcode with FB_CROAK then it fails and falls back to Windows-1251 and the whole thing is corrupted. The solution is to do lenient transcoding in cases where we're fairly certain the data is UTF-16 (BOM or NUL byte pattern). That replaces the lone surrogates with the Unicode replacement character (U+FFFD) but keeps most of the code intact. If we're not sure about the charset, then still use FB_CROAK. It's a small change, plus another test case. I'll go ahead and commit, but if it's problem let me know. --- a/lib/Mail/SpamAssassin/Message/Node.pm +++ b/lib/Mail/SpamAssassin/Message/Node.pm @@ -650,13 +650,20 @@ sub _normalize { # Failing that, a UTF-16 label is trusted: its byte order if it names one, # else big-endian (RFC 2781). Undeclared text that is not UTF-16 is left # to the detection fallbacks below. + # + # With a BOM or the NUL pattern the data is known to be UTF-16, so decode + # leniently: a lone surrogate (legal in a JavaScript string, e.g. a pair + # split across two string literals) becomes U+FFFD instead of failing the + # whole part. A label alone is weaker evidence, so that decode stays strict. my $decoder = detect_utf16( $_[0] ); + my $check = Encode::FB_DEFAULT; if (!defined $decoder && $charset_declared ne '') { $decoder = Encode::find_encoding( $charset_declared =~ /LE\z/i ? 'UTF-16LE' : 'UTF-16BE'); + $check = Encode::FB_CROAK; } if (defined $decoder) { - if (eval { $rv = $decoder->decode($_[0], Encode::FB_CROAK | Encode::LEAVE_SRC); defined $rv }) { + if (eval { $rv = $decoder->decode($_[0], $check | Encode::LEAVE_SRC); defined $rv }) { dbg("message: decoded as charset %s, declared %s", $decoder->name, $charset_declared); utf8::encode($rv) if !$return_decoded; -- You are receiving this mail because: You are the assignee for the bug.
