Hello everyone,

while upgrading ~80 applications from Java 21 to 25 runtime we noticed
that one specifically was crashing with a segfault. Since we had no
idea where this came from, we used Claude Opus 5 to analyze the dump
and write up a report. I validated the report and the proposed fix
worked, so I assume it's correct. Please forgive me for pasting the
output directly, this goes way beyond my knowledge but I think it's
helpful.

We are hitting a reproducible JVM crash in the FFM-based HarfBuzz
layout path added
by JDK-8318364. It looks like a lifetime/thread-safety defect around
the font-table
upcall stub. A 30-line reproducer is attached.

The crash only becomes *visible* when the process uses jemalloc instead of glibc
malloc, for reasons explained under "Why jemalloc matters" below. We
do not believe
jemalloc is the bug.

## Crash
  C  [libjemalloc.so.2+0x1c4c5]
  C  [libfontmanager.so+0x53e8e]  hb_blob_destroy.part.0+0x2e
  C  [libfontmanager.so+0x5f759]  hb_face_t::load_upem() const+0xe9
  C  [libfontmanager.so+0x75b28]  _hb_font_create(hb_face_t*)+0xf8
  C  [libfontmanager.so+0x7f4cf]  hb_font_create+0xf
  C  [libfontmanager.so+0x8347d]  jdk_font_create_hbp+0x1d
  C  [libfontmanager.so+0xd2e0]   jdk_hb_shape+0x100
  v  ~RuntimeStub::nep_invoker_blob
  j  jdk.internal.foreign.abi.DowncallStub...
  j  sun.font.HBShaper.shape(...)                    [email protected]
  j  sun.font.SunLayoutEngine.layout(...)            [email protected]
  j  sun.font.GlyphLayout.layout(...)                [email protected]

  siginfo: si_signo: 11 (SIGSEGV), si_code: 1 (SEGV_MAPERR), si_addr: 0x8

SEGV_MAPERR at a tiny address is a stale/freed pointer being dereferenced, not a
null check we could have avoided.

## Reproducer
Attached HbCrash.java. No dependencies, no font file -- it uses the
logical "Dialog"
font. Run with a single-file launch:

  LD_PRELOAD=/usr/lib64/libjemalloc.so.2 \
    java -Djava.awt.headless=true HbCrash.java

Four threads each render 4000 short strings with
Graphics2D.drawString(). The font
carries TextAttribute.TRACKING, which forces the full GlyphLayout/shaping path
rather than the simple glyph-mapping fast path.

It crashes within a few seconds, 3 runs out of 3.

## What we measured
All local rows are the same machine (x86_64, Fedora 44, jemalloc 5.3.0), same
reproducer, run three times each:

  JDK                          allocator   threads   result
  ---------------------------------------------------------------
  Corretto 25.0.4.1+8          jemalloc    4         crash 3/3
  Red Hat OpenJDK 25.0.4.1+1   jemalloc    4         crash
  Corretto 25.0.4.1+8          jemalloc    1         ok
  Corretto 25.0.4.1+8          glibc       4         ok
  Corretto 21.0.12.1           jemalloc    4         ok
  Corretto 25 + ffm=false      jemalloc    4         ok 3/3

The last row is -Dsun.font.layout.ffm=false, which routes layout back
through the
pre-FFM JNI path in SunLayoutEngine. That is also our current
production workaround
and it has held.

So: needs concurrency, needs the FFM path, absent in 21.

We originally saw this in production on aarch64 (AlmaLinux 10.2, Corretto
25.0.4.1+8, Tomcat request threads rendering PNG labels). Both hs_err files are
attached -- the aarch64 production one and the x86_64 reproducer one. The native
frame sequence is identical across the two architectures.

We measured 21 (path absent) and 25 (path present and default-on). We have not
tested 22/23/24; JBS lists JDK-8318364 with fix version 24, so we
assume 24 is the
first affected release, but have not confirmed that ourselves.

## Why jemalloc matters
We believe jemalloc is only the observability mechanism, not the
cause. glibc malloc
keeps freed chunks mapped, so a use-after-free reads plausible garbage
and usually
goes unnoticed. jemalloc purges freed pages back to the OS -- in our case with
MALLOC_CONF=dirty_decay_ms:30000,muzzy_decay_ms:30000 -- so the same
stale read hits
an unmapped page and faults immediately.

That is consistent with SEGV_MAPERR. We would expect the same defect
to be silently
corrupting memory under glibc rather than being absent there, but we
have not tried
to demonstrate that.

## Suspected mechanism (hypothesis, not measured)
In HBShaper.FaceRef.createFace():

    get_table_data_fn = getBoundUpcallStub(Arena.ofAuto(), ...);
    face = (MemorySegment) create_face_handle.invokeExact(get_table_data_fn);

The upcall stub HarfBuzz calls back into to read font tables is allocated from
Arena.ofAuto(), which frees non-deterministically once the arena becomes
unreachable. The native hb_face_t holds a raw pointer to it. FaceRef
keeps a Java
reference to the segment, and getFace() is synchronized, but the faces
are cached in
a static WeakHashMap<Font2D, FaceRef> shared across all threads, and
hb_face_t's own
lazy state (load_upem, which does reference_table -> upcall ->
hb_blob_destroy) is
reached concurrently.

We have not instrumented this far enough to say which of those is the actual
violation -- that part is inference from reading the code, not something we
measured. What we did establish is that the FFM path is required to
trigger it and
that the JNI path is unaffected.

## Things we ruled out
Listing these because they were our first guesses and cost us time:

- Not fontconfig. libfontconfig is not mapped at all in the production crash
  (we checked the memory map), and it crashes with or without a writable
  fontconfig cache.
- Not duplicate freetype. Only the JDK's own bundled
  lib/libfreetype.so is mapped; the system copy is not loaded.
- Not compact object headers. Present in our production config, but the
  reproducer crashes without it.
- Not OOM or memory pressure. 1.2G RSS against a 1.6G limit, and the
  reproducer crashes with a tiny heap.

## Attachments
  HbCrash.java                    minimal reproducer, no dependencies
  LabelRepro.java                 closer to our real workload (loads a TTF via
                                  Font.createFont); needs NotoMono-Regular.ttf
  hs_err_x86_64_corretto25.log    from HbCrash.java
  hs_err_aarch64_corretto25.log   from production

Happy to run further experiments or test a patch -- the reproducer is fast and
deterministic for us.

-- 

Philipp Trulson
Platform Engineer
mail: [email protected] · web: www.rebuy.de

-- 





rebuy recommerce GmbH* · *Erkelenzdamm 11-13* · *10999 Berlin* · 
*Geschäftsführer: Dr. Philipp Gattner, Marcel ErianSitz und 
Registergericht: Berlin, Amtsgericht Charlottenburg, HRB 109344 B, 
*USt-ID-Nr.:* DE237458635

Reply via email to