`java.net.IDN` still implements IDNA2003, which is frozen at Unicode 3.2. This
PR moves it to UTS46 (IDNA2008 with the Unicode compatibility mapping), which
is what the big three browsers and most other languages' IDN libraries use.
Names such as faß.de have been registrable since 2010, and today `IDN` turns
that into fass.de, which is a different domain.
### Changes
The implementation is ICU4J's UTS46 code, merged into `jdk.internal.icu` with
the smallest reasonable diff. The UTS46 data (`uts46.nrm`) is generated from
Unicode 17, the same version as `Character`, so `IDN` can follow future Unicode
upgrades instead of being stuck at 3.2.
This patch removes the IDNA2003 code and its data file,
`sun/net/idn/uidna.spp`,and does not merge the ICU code that only registries
need.
I also fixed a typo in ICU's `Punycode.isBasicUpperCase()`.
There are three new public constants:
- `CHECK_CONTEXTJ` rejects ZWJ/ZWNJ (U+200D, U+200C) in contexts where RFC5892
doesn't allow them.
- `NONTRANSITIONAL_TO_ASCII` maps ß, ς, ZWJ and ZWNJ to themselves rather than
to ss, σ and nothing. This is what registries and browsers do now.
- `NONTRANSITIONAL_TO_UNICODE` is the same, for `toUnicode()`.
`ALLOW_UNASSIGNED` is kept for source compatibility but ignored, since UTS46
knows the whole Unicode repertoire of the running JVM.
The one-argument `toASCII()` and `toUnicode()` now use
`CHECK_CONTEXTJ|NONTRANSITIONAL_TO_…`. Before, they used no flags.
### Behaviour, before and after
Here are the results from JDK 27 and from this branch.
| Input | Call | JDK 27 | This PR |
|---|---|---|---|
| grå.org | `toASCII(s)` | xn--gr-zia.org | xn--gr-zia.org |
| 例子.中国 | `toASCII(s)` | xn--fsqu00a.xn--fiqs8s | xn--fsqu00a.xn--fiqs8s |
| faß.de | `toASCII(s)` | fass.de | xn--fa-hia.de |
| faß.de | `toASCII(s, 0)` | fass.de | fass.de |
| GRÅ.ORG | `toASCII(s)` | xn--gr-zia.ORG | xn--gr-zia.org |
| a<U+200D>b.example.com | `toASCII(s)` | ab.example.com |
IllegalArgumentException |
| *.example.com | `toASCII(s)` | *.example.com | *.example.com |
| *.example.com | `toASCII(s, USE_STD3_ASCII_RULES)` | IllegalArgumentException
| IllegalArgumentException |
| ab--cd.example.com | `toASCII(s)` | ab--cd.example.com | ab--cd.example.com |
| xn--fa-hia.de | `toUnicode(s)` | xn--fa-hia.de | faß.de |
`toUnicode()` still never throws. It returns any label it can't convert
unchanged, as before.
UTS46 considers labels with `--` in the third and fourth positions as errors.
`IDN` deliberately ignores that error. That rule is meant for public
registries. While you cannot register ab--cd.com, the owner of example.com can
create ab--cd.example.com. One can argue the merits of doing that, but JDK has
accepted it so far, and this PR maintains compatibility when in doubt.
### STD3_ASCII_RULES
`USE_STD3_ASCII_RULES` works well for hostnames and should perhaps be used, for
better compatibility with IDNA2008 and registries. However,
`sun.security.util.HostnameChecker.isMatched()` and perhaps other callers call
`IDN.toASCII()` with a domain wildcard, which leads to friction and general
unhappiness. The IDN documentation suggests that `toAscii()` works on domain
names, not wildcards, so perhaps HostnameChecker is relies on unspecified
behaviour there, I'm not sure, but this code exists. Code outside the JDK may
well make the same assumption.
There's even a unit test that calls `IDN.toASCII("*.example.com")`. Code
outside the JDK may *well* copy this behaviour.
Because of of a wish for optimal compatibility, the one-argument methods don't
use `USE_STD3_ASCII_RULES`. I'm frankly uncertain whether it would be better to
USE_STD_ASCII_RULES in `IDN.to…()` and update HostnameChecker.
Note that `SNIHostName` and other callers that want STD3 rules already pass the
flag explicitly.
### TLS host name matching
`HostnameChecker` now compares nontransitionally. A certificate for
fass.example.com used to match faß.example.com, and now it doesn't. They are
different domains under IDNA2008, so I think that is a fix.
### Spec change
The class javadoc has been extended considerably. An ordinary Java developer
needs some background to choose flags, and the old text described IDNA2003. The
RFC links point to RFC5890/RFC5891 rather than RFC3490. This needs a CSR, which
I will file once the STD3 question above has been settled.
## Unicode versions
I presume that this change needs a note somewhere, so next time someone imports
a new version of the unicode tables, `uts46.nrm` is updated along with the rest
of the family. But I've no idea where.
### Testing
- `test/jdk/java/net/IDN/UTS46.java` is new, and quite comprehensive.
- `sun/net/idn/TestStringPrep.java` no longer tests nameprep, which this PR
removes. `PunycodeTest.java` uses `StringBuilder` instead of `StringBuffer`, as
the merged `Punycode` class does.
- tier1 passes on linux-x64, except the hotspot gtest wrappers, because my
local build happened to be configured without gtest.
---------
- [x] I confirm that I make this contribution in accordance with the [OpenJDK
Interim AI Policy](https://openjdk.org/legal/ai).
I should note that I used an AI tool to rebase the branch and deal with my
lazily incorrect copyright headers, lint problems etc.
-------------
Commit messages:
- 6988055: Update java.net.IDN to UTS#46 and merge parts of ICU.
Changes: https://git.openjdk.org/jdk/pull/33210/files
Webrev: https://webrevs.openjdk.org/?repo=jdk&pr=33210&range=00
Issue: https://bugs.openjdk.org/browse/JDK-6988055
Stats: 4319 lines in 17 files changed: 3774 ins; 379 del; 166 mod
Patch: https://git.openjdk.org/jdk/pull/33210.diff
Fetch: git fetch https://git.openjdk.org/jdk.git pull/33210/head:pull/33210
PR: https://git.openjdk.org/jdk/pull/33210