`java.net.IDN` still implements IDNA2003, which is frozen at Unicode 3.2. This 
PR moves it to UTS46 (IDNA2008 with the Unicode compatibility mapping), which 
is what the big three browsers and most other languages' IDN libraries use.

Names such as faß.de have been registrable since 2010, and today `IDN` turns 
that into fass.de, which is a different domain.

### Changes

The implementation is ICU4J's UTS46 code, merged into `jdk.internal.icu` with 
the smallest reasonable diff. The UTS46 data (`uts46.nrm`) is generated from 
Unicode 17, the same version as `Character`, so `IDN` can follow future Unicode 
upgrades instead of being stuck at 3.2.

This patch removes the IDNA2003 code and its data file, 
`sun/net/idn/uidna.spp`,and does not merge the ICU code that only registries 
need.

I also fixed a typo in ICU's `Punycode.isBasicUpperCase()`.

There are three new public constants:

- `CHECK_CONTEXTJ` rejects ZWJ/ZWNJ (U+200D, U+200C) in contexts where RFC5892 
doesn't allow them.
- `NONTRANSITIONAL_TO_ASCII` maps ß, ς, ZWJ and ZWNJ to themselves rather than 
to ss, σ and nothing. This is what registries and browsers do now.
- `NONTRANSITIONAL_TO_UNICODE` is the same, for `toUnicode()`.

`ALLOW_UNASSIGNED` is kept for source compatibility but ignored, since UTS46 
knows the whole Unicode repertoire of the running JVM.

The one-argument `toASCII()` and `toUnicode()` now use 
`CHECK_CONTEXTJ|NONTRANSITIONAL_TO_…`. Before, they used no flags.

### Behaviour, before and after

Here are the results from JDK 27 and from this branch.

| Input | Call | JDK 27 | This PR |
|---|---|---|---|
| grå.org | `toASCII(s)` | xn--gr-zia.org | xn--gr-zia.org |
| 例子.中国 | `toASCII(s)` | xn--fsqu00a.xn--fiqs8s | xn--fsqu00a.xn--fiqs8s |
| faß.de | `toASCII(s)` | fass.de | xn--fa-hia.de |
| faß.de | `toASCII(s, 0)` | fass.de | fass.de |
| GRÅ.ORG | `toASCII(s)` | xn--gr-zia.ORG | xn--gr-zia.org |
| a<U+200D>b.example.com | `toASCII(s)` | ab.example.com | 
IllegalArgumentException |
| *.example.com | `toASCII(s)` | *.example.com | *.example.com |
| *.example.com | `toASCII(s, USE_STD3_ASCII_RULES)` | IllegalArgumentException 
| IllegalArgumentException |
| ab--cd.example.com | `toASCII(s)` | ab--cd.example.com | ab--cd.example.com |
| xn--fa-hia.de | `toUnicode(s)` | xn--fa-hia.de | faß.de |

`toUnicode()` still never throws. It returns any label it can't convert 
unchanged, as before.

UTS46 considers labels with `--` in the third and fourth positions as errors. 
`IDN` deliberately ignores that error. That rule is meant for public 
registries. While you cannot register ab--cd.com, the owner of example.com can 
create ab--cd.example.com. One can argue the merits of doing that, but JDK has 
accepted it so far, and this PR maintains compatibility when in doubt.

### STD3_ASCII_RULES

`USE_STD3_ASCII_RULES` works well for hostnames and should perhaps be used, for 
better compatibility with IDNA2008 and registries. However, 
`sun.security.util.HostnameChecker.isMatched()` and perhaps other callers call 
`IDN.toASCII()` with a domain wildcard, which leads to friction and general 
unhappiness. The IDN documentation suggests that `toAscii()` works on domain 
names, not wildcards, so perhaps HostnameChecker is relies on unspecified 
behaviour there, I'm not sure, but this code exists. Code outside the JDK may 
well make the same assumption.

There's even a unit test that calls `IDN.toASCII("*.example.com")`. Code 
outside the JDK may *well* copy this behaviour.

Because of of a wish for optimal compatibility, the one-argument methods don't 
use `USE_STD3_ASCII_RULES`. I'm frankly uncertain whether it would be better to 
USE_STD_ASCII_RULES in `IDN.to…()` and update HostnameChecker.

Note that `SNIHostName` and other callers that want STD3 rules already pass the 
flag explicitly.

### TLS host name matching

`HostnameChecker` now compares nontransitionally. A certificate for 
fass.example.com used to match faß.example.com, and now it doesn't. They are 
different domains under IDNA2008, so I think that is a fix.

### Spec change

The class javadoc has been extended considerably. An ordinary Java developer 
needs some background to choose flags, and the old text described IDNA2003. The 
RFC links point to RFC5890/RFC5891 rather than RFC3490. This needs a CSR, which 
I will file once the STD3 question above has been settled.

## Unicode versions

I presume that this change needs a note somewhere, so next time someone imports 
a new version of the unicode tables, `uts46.nrm` is updated along with the rest 
of the family. But I've no idea where.

### Testing

- `test/jdk/java/net/IDN/UTS46.java` is new, and quite comprehensive.
- `sun/net/idn/TestStringPrep.java` no longer tests nameprep, which this PR 
removes. `PunycodeTest.java` uses `StringBuilder` instead of `StringBuffer`, as 
the merged `Punycode` class does.
- tier1 passes on linux-x64, except the hotspot gtest wrappers, because my 
local build happened to be configured without gtest.




---------
- [x] I confirm that I make this contribution in accordance with the [OpenJDK 
Interim AI Policy](https://openjdk.org/legal/ai).

I should note that I used an AI tool to rebase the branch and deal with my 
lazily incorrect copyright headers, lint problems etc.

-------------

Commit messages:
 - 6988055: Update java.net.IDN to UTS#46 and merge parts of ICU.

Changes: https://git.openjdk.org/jdk/pull/33210/files
  Webrev: https://webrevs.openjdk.org/?repo=jdk&pr=33210&range=00
  Issue: https://bugs.openjdk.org/browse/JDK-6988055
  Stats: 4319 lines in 17 files changed: 3774 ins; 379 del; 166 mod
  Patch: https://git.openjdk.org/jdk/pull/33210.diff
  Fetch: git fetch https://git.openjdk.org/jdk.git pull/33210/head:pull/33210

PR: https://git.openjdk.org/jdk/pull/33210

Reply via email to