aizu-m opened a new pull request, #123:
URL: https://github.com/apache/poi-xmlbeans/pull/123

   Noticed while checking how xsd:base64Binary values get validated. A document 
whose base64Binary value carried trailing junk still validated clean:
   
       <t:root>SGVsbG8=!!!!</t:root>   ->   doc.validate() == true
   
   Traced it to `JavaBase64Holder.lex`. It decodes with 
`Base64.getMimeDecoder().decode(v)`, and the JDK MIME decoder silently drops 
every character that is not in the base64 alphabet, not only line separators. 
So `SGVsbG8=!!!!` decodes to `Hello`, `SGV!!!sbG8=` decodes to `Hello`, and 
`!!!!` decodes to an empty array. Each is accepted instead of being reported 
invalid. `lex` is the validator the streaming `Validator` calls for 
base64Binary through `validateLexical`, and `set_text` calls it too, so an 
out-of-alphabet value passes document validation.
   
   The sibling `JavaHexBinaryHolder.lex` does not have this problem: 
`HexBin.decode` returns null on any non-hex character and the value is reported 
invalid. base64 had drifted from that behaviour.
   
   Fix scans the value first and reports it invalid when a character is neither 
in the base64 alphabet nor XML whitespace, then decodes as before. Whitespace 
and line-wrapped values are untouched, so valid input decodes exactly as it did.
   
       before: <t:root>SGVsbG8=!!!!</t:root> validates
       after:  reported invalid; "SGVsbG8=", "SGVs bG8=" and newline-wrapped 
values still validate
   
   Regression test added in `Base64BinaryValidateTest`.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to