Hi Tilman, It's way cleaner than my POC.
I tested your changes locally (the setFont change plus the splitter map). I got the same speedup as my original patch, with no output change. On a 96-page PDF, the time spent drawing text went from 3.1 s to 0.85 s. That's great. Thanks a lot! Olivier On Fri, 17 Jul 2026 at 15:08, Tilman Hausherr <[email protected]> wrote: > Hi, > > I prefer something more simple because I don't want to keep a static map > forever. Like this: > > Map<String,GlyphArraySplitter> map = new WeakHashMap<>(); > > and in applyGsubFeature: > > GlyphArraySplitter glyphArraySplitter = > map.computeIfAbsent(scriptFeature.getName(), (__) -> new > GlyphArraySplitterRegexImpl(allGlyphIdsForSubstitution)); > > is this what you had in mind? > > Tilman > > On 2026/07/17 10:08:44 Olivier Bruchez wrote: > > Hi Tilman, > > > > Thanks for your reply and for opening an issue. > > > > > > > I assume what you did was to check the gsubWorkers map before calling > > > gsubWorkerFactory.getGsubWorker(), correct? > > > > > > > Not exactly. I should have included my patch/"hack" in the original > message > > for more clarity. I'm basically simply adding a global cache for the > > GlyphArraySplitter, inside GsubWorkerForLatin. I left setFont and the > > gsubWorkers map alone. It is built in applyGsubFeature (once per > showText), > > not in the constructor, so reusing the GsubWorker alone would not stop > the > > rebuild and the performance hit. I memoized the GlyphArraySplitter per > font > > and then per script feature (I think this is correct, but I'm not 100% > > sure). > > > > The whole change is one line in applyGsubFeature: > > > > - GlyphArraySplitter glyphArraySplitter = new > > > GlyphArraySplitterRegexImpl(scriptFeature.getAllGlyphIdsForSubstitution()); > > + GlyphArraySplitter glyphArraySplitter = splitterFor(scriptFeature); > > > > > https://github.com/apache/pdfbox/blob/3.0.4/fontbox/src/main/java/org/apache/fontbox/ttf/gsub/GsubWorkerForLatin.java#L91-L92 > > > > Plus a static cache keyed on the font's GsubData and a small helper: > > > > private static final Map<GsubData, Map<String, GlyphArraySplitter>> > > SPLITTER_CACHE = new WeakHashMap<>(); > > > > private GlyphArraySplitter splitterFor(ScriptFeature scriptFeature) > > { > > synchronized (SPLITTER_CACHE) > > { > > Map<String, GlyphArraySplitter> perFont = > > SPLITTER_CACHE.get(gsubData); > > if (perFont == null) > > { > > perFont = new HashMap<>(); > > SPLITTER_CACHE.put(gsubData, perFont); > > } > > GlyphArraySplitter splitter = > perFont.get(scriptFeature.getName()); > > if (splitter == null) > > { > > splitter = new > > > GlyphArraySplitterRegexImpl(scriptFeature.getAllGlyphIdsForSubstitution()); > > perFont.put(scriptFeature.getName(), splitter); > > } > > return splitter; > > } > > } > > > > This is a proof of concept, of course. This was just to convince myself > > that rebuilding GlyphArraySplitter multiple times was indeed the problem > I > > was observing. On larger PDF files (dozens of pages), the speedup with > the > > code above is really huge (3-4x on the total PDF generation times). > > > > Thanks, > > Olivier > > > > --------------------------------------------------------------------- > To unsubscribe, e-mail: [email protected] > For additional commands, e-mail: [email protected] > >
