Hi,
"I" wrote:
`java.net.IDN` still implements IDNA2003, which is frozen at
Unicode 3.2. This PR moves it to UTS46 (IDNA2008 with the
Unicode compatibility mapping), which is what the big three
browsers and most other languages' IDN libraries use.
Names such as faß.de have been registrable since 2010, and today
`IDN` turns that into fass.de, which is a different domain.
### Changes
Sorry about that Markdown, I wrote a nice readable description for
the pull request and assumed that the bot would produce HTML or
even text/markdown.
I am particularly interested in comments on whether to use
USE_STD3_ASCII_RULES or not. That's a tricky tricky question.
`USE_STD3_ASCII_RULES` works well for hostnames and should
perhaps be used, for better compatibility with IDNA2008 and
registries. However,
`sun.security.util.HostnameChecker.isMatched()` and perhaps
other callers call `IDN.toASCII()` with a domain wildcard, which
leads to friction and general unhappiness. The IDN documentation
suggests that `toAscii()` works on domain names, not wildcards,
so perhaps HostnameChecker is relies on unspecified behaviour
there, I'm not sure, but this code exists. Code outside the JDK
may well make the same assumption.
There's even a unit test that calls
`IDN.toASCII("*.example.com")`. Code outside the JDK may *well*
copy this behaviour.
Because of of a wish for optimal compatibility, the one-argument
methods don't use `USE_STD3_ASCII_RULES`. I'm frankly uncertain
whether it would be better to USE_STD_ASCII_RULES in `IDN.to…()`
and update HostnameChecker.
Note that `SNIHostName` and other callers that want STD3 rules
already pass the flag explicitly.
In general, if anyone wants too much detail about how individual
unicode code points ought to be handled in the DNS or email or
anywhere: just ask ;)
Arnt