Hi Eric,

Thank you, Eric! This is incredibly helpful.

I hope our speed improvements for v3 prove otherwise. Here is a quick
rundown of what we have been working hard on:

  - Most of the codebase is now thread-safe (showing about a 2.5x
improvement).
  - We are removing all regex from the code for better performance and
fewer bugs.
  - We have a much cleaner interface and full UTF support.
  - We added several internationalization features and bi-directional
normalization for accurate highlighting.
  - We added a pile of new features to addons, including a PR for a Java
trainer and static embeddings (which can produce 100k/s on many systems).

It is our mission (and my obsession) to make OpenNLP performant and on par
with popular NLP packages in other ecosystems.

So it sounds like GraalVM was a great direction to explore. I have a ticket
open for it now. Outside of containers, is anyone delivering native
binaries for libraries?

Rungs,
Kristian


On Tue, Sep 15, 2026 at 7:28 AM Eric Pugh <[email protected]>
wrote:

> I’ve been really intrigued with GraalVM….     A couple of observations:
>
> 1) We are using it in the forthcoming solr-mcp tool to drastically speed
> up how fast solr-cmp starts when used as a command line tool.  Solr-MCP
> leverages Spring and Spring AI, and as you can imagine, there is a lot of
> wiring dynamically going on in that.  GraalVM changes that by “pre
> compiling” (not sure that is the right word) all of that so you get super
> fast start up.  We ship the GraalVM version in a dockerized version of
> solr-mcp so you don’t need to install anything locally.  And yeah, it adds
> some build complexity on top to get that to work, but then you are done.
>
> 2) Apache Tika actually runs on GraalVM!
> https://docs.quarkiverse.io/quarkus-tika/dev/index.html.  Tika is an
> example of a project that has a million dependencies, so if it can run on
> GraalVM….
>
> 3) OpenNLP is definitely bounded by performance….  When we added it to
> Solr as a UpdateRequestProcessor, any model beyond the most simplistic made
> the indexing pace almost unusuable, which means you need to have a primary
> indexing path that doesn’t touch OpenNLP and then a secondary “enrichment”
> one that comes back async and does the OpenNLP work…
>
>
>
> > On Sep 14, 2026, at 10:18 PM, Kristian Rickert <[email protected]>
> wrote:
> >
> > *Is something that can be achieved via a Maven profile?*
> >
> > Yes, this is typically done using the native-maven-plugin provided by
> > GraalVM.
> >
> > *Is this build done via the project using OpenNLP, or does OpenNLP itself
> > have to be compiled differently?*
> >
> > OpenNLP itself does not need to be compiled differently. The library is
> > still compiled to standard Java bytecode just like normal. Native
> > compilation happens at the final application or distribution level.
> >
> > I'll give you my 2 cents since I've spent the last hour refreshing myself
> > on it.
> >
> > First I want to state my motivation: I want to create a native library
> for
> > apps calling OpenNLP (gRPC is just one example).  I'm convinced what we
> are
> > building can be integrated into many text-heavy apps in Rust, C/C++, or
> > Swift. It could help word processing apps, AI runtime analysis, and RAG
> > pipelines everyone is hyping about.  Basically, between this and gRPC,
> our
> > library can reach any app.
> >
> > The tool that builds the binaries - the GraalVM Native Image - analyzes
> the
> > application entry point and traces all reachable code across all
> > dependencies to produce a single binary executable. This is similar to
> what
> > you would get with gcc, and it comes with its own set of advantages and
> > headaches.
> >
> > I don't understand the magic of this process BUT it sure sounds cool and
> is
> > a low level java thing to learn. Richard's video he references goes over
> > this entire process in detail, and I plan to watch it (it's a great
> > presentation - I'm a fan of his font choice). He explains the compilation
> > process and the performance implications without trying to sell you on
> the
> > technology either way.
> >
> > Everything Richard said is right, and he is rightfully questioning my
> speed
> > claims. Native Image gives you near instant startup time and a very low
> > initial memory footprint. However, a traditional JVM JIT compiler might
> > actually beat it in peak sustained throughput for long running tasks. I
> > have not claimed with certainty that it will be faster for raw
> > computational throughput, but I'm enthusiastic and fueled by hope that it
> > would. We will have to measure the performance and test it extensively.
> >
> > Most modern Java tries to account for this. Both Quarkus and Micronaut
> are
> > built to work with native compilation out of the box.  It's why I use
> > Quarkus in a lot of my projects - a premature optimization I've not
> needed
> > yet but can save your AWS bills by easy double digits.  When you get a
> > multi-million dollar bill from Bezos, this suddenly becomes a more
> > attractive choice :)
> >
> > For OpenNLP, almost all the code should work fine. The biggest friction
> > point will be the model loading. We need to make sure any models loaded
> as
> > classpath resources or anything relying on Java serialization are
> > explicitly registered in the Native Image configuration files or we can
> > refactor the code to load services through an SPI loader.  I'm a big fan
> of
> > SPI loading, so I'd lean toward that.
> >
> > Last but not least, since this is Oracle's baby, they own GraalVM, and
> > there's a community edition.  I've only used the CE, because the Oracle
> one
> > is a commercial product.  You can use it for free in the Oracle cloud -
> but
> > I don't want to be "that guy" in our org to suggest it.
> >
> > Initially, I'd only want to support compiling it and leave the various
> arch
> > binaries up to the user.  IMHO anyone who uses GraalVM right now doesn't
> > look for native binaries, they just build the binaries themselves or
> might
> > use a native docker container instead (I've never seen a java native
> > library in the wild that isn't a container).  We'll make a lot of DevOps
> > folks happy if we can advertise that we are GaalVM compliant - if they
> use
> > it.  To Richard's point, the other project he referenced hasn't seen much
> > action from it.  By doing this, we learn a lot, potentially offer
> something
> > that integrates into a larger ecosystem, and if it fails we remove it
> from
> > the build.
> >
> >
> > On Mon, Sep 14, 2026, 4:29 PM Jeff Zemerick <[email protected]>
> wrote:
> >
> >> I'm not super familiar with GraalVM or what OpenNLP requires to be
> >> compatible with it but I understand the desire to do so.
> >>
> >> A couple questions:
> >>
> >> - Is something that can be achieved via a Maven profile?
> >> - Is this build done via the project using OpenNLP, or does OpenNLP
> >> itself have to be compiled differently?
> >>
> >> Thanks,
> >> Jeff
> >>
> >> On Mon, Sep 14, 2026 at 8:21 AM Richard Zowalla <[email protected]>
> wrote:
> >>>
> >>> Hi,
> >>>
> >>> if the deliverable is "build it yourself", the actual product is the
> >> reachability metadata in our jars and the reflection fixes, and that
> can't
> >> live in the sandbox. It has to go into the main repo, or users still
> can't
> >> build a working image. The sandbox can hold the recipe and the gRPC
> binary
> >> experiment, but the enabling work is main-repo work and needs to be
> scoped
> >> as such.
> >>>
> >>> On the demand argument, note that native image on CE gives startup and
> >> memory footprint (I recommend "Scotty I need Warp Speed" from Gerrit
> >> Grunwald) not peak throughput (no PGO, no G1). For long-running
> validation
> >> pipelines a warm JIT may still win.
> >>>
> >>> Richard
> >>>
> >>> On 2026/09/14 11:53:20 Kristian Rickert wrote:
> >>>> Hi Richard,
> >>>>
> >>>> Thank you for the detailed response. It answered all of my questions
> >> and
> >>>> provides a clear path forward with great historical context we should
> >> heed.
> >>>>
> >>>> I agree with your technical assessment, but see a subtle difference
> >>>> regarding demand: I believe proactive innovation drives demand.  I
> >> call it
> >>>> the Kevin Kostner effect: "if you build it, they will come".
> >>>>
> >>>> But I also hope to drum up collaboration in this thread, as an
> >> opportunity
> >>>> to learn exciting, potential-filled new technologies. Adding native
> >>>> compilation could open new channels for library integration.  Learning
> >> this
> >>>> is a great resume builder and offers experience in profiling and
> >> testing
> >>>> that exceeds our current scope.  And it's low risk but takes a
> >> calculated
> >>>> chance.  This is why it'll remain in the sandbox until the
> >>>> build-then-demand model is proven.
> >>>>
> >>>> But to gain demand - that's when blog posts, demos, and conference
> chit
> >>>> chat come in.  So I promise to focus most of my efforts there once it
> >> is
> >>>> built.
> >>>>
> >>>> Here is why I think it'll work:
> >>>>
> >>>> NLP has been overshadowed by LLM hype, but as the need for fast
> >>>> ground-truth validation grows, performance will be key. People
> unfairly
> >>>> separate LLMs, search, and NLP when they all deal with language and
> >> they
> >>>> need better integration.  Since Java isn't natively fast (though it is
> >>>> still significantly faster than Python), demonstrating GraalVM's speed
> >> and
> >>>> integration capabilities could reignite interest in the ecosystem,
> >> though
> >>>> complex setup remains a risk.
> >>>>
> >>>> Finally I won't dismiss that maintaining native binaries is a headache
> >> and
> >>>> I think we might want to avoid it.  CVE concerns given rapidly
> changing
> >>>> architectures compound this issue. Even if it leaves the sandbox,
> >> keeping
> >>>> this as a build-it-yourself setup may be the best approach. This
> avoids
> >>>> direct binary distribution given the complexity of the setup (ARM/AMD,
> >>>> NVidia vs. Intel, CPU/GPU/NPU, and OS all need to be considered). If
> >> this
> >>>> experiment fails to gain traction, I fully support deprecating and
> >> retiring
> >>>> it quickly.
> >>>>
> >>>> Thanks again for the thorough write-up.
> >>>>
> >>>> Best,
> >>>> Kristian
> >>>>
> >>>>
> >>>>
> >>>> On Mon, Sep 14, 2026, 2:42 AM Richard Zowalla <[email protected]>
> wrote:
> >>>>
> >>>>> Hi,
> >>>>>
> >>>>> I'd frame the scope a bit differently, though. Here are my 2cts from
> >> my
> >>>>> work in the EE ecosystem.
> >>>>>
> >>>>> OpenNLP is a library. Compiling to native code is something the
> >> user's
> >>>>> application does, not something we do IMHO. What OpenNLP can do is
> >> make the
> >>>>> jars native-friendly, so that anyone building a native image (with
> >> the
> >>>>> GraalVM Maven/Gradle plugins, Quarkus, Spring AOT, Apache Geronimo
> >> Arthur,
> >>>>> ...) gets a working build without extra steps. Concretely:
> >>>>>
> >>>>> 1. Ship reachability metadata inside our jars
> >>>>> (META-INF/native-image/org.apache.opennlp/<artifactId>/...), or a
> >> small
> >>>>> GraalVM Feature class.
> >>>>> 2. Fix the parts that break in a native image. We do use reflection,
> >> and
> >>>>> it isn't only in edge cases:
> >>>>> (a) ExtensionLoader / BaseToolFactory / TrainerFactory load classes
> >> by
> >>>>> name. The factory class names even come from the model manifest.
> >>>>> (b) GenericModelReader/Writer and ClassPathModelLoader create
> >> objects
> >>>>> through reflective constructors.
> >>>>> (c) The Snowball stemmers look up methods via
> >> MethodHandles.findVirtual.
> >>>>> (d) The CLI's ArgumentParser uses dynamic proxies.
> >>>>> (e) SimpleClassPathModelFinder reads jdk.internal internals and
> >> scans
> >>>>> jars on the classpath. That can't work in a native image because
> >> there is
> >>>>> no classpath.
> >>>>> (f) opennlp-dl pulls in ONNX Runtime through JNI.
> >>>>> (g) We load a fair number of bundled resources.
> >>>>> 3. Do some experiments outside of the OpenNLP repos, like a native
> >> smoke
> >>>>> test in CI (tokenize, POS tag, NER).
> >>>>>
> >>>>> Custom factories or extensions from users will always need
> >> registering by
> >>>>> the user, and that's fine.
> >>>>>
> >>>>> On binary distributions: the (WIP) gRPC server would be the obvious
> >>>>> candidate for native. It's a standalone, long-running service, it
> >> lives in
> >>>>> the sandbox anyway, and a native container image makes sense there.
> >>>>>
> >>>>> A native CLI could be nice for startup time, but I have no idea how
> >> many
> >>>>> people actually use the CLI, so I'd wait until someone asks for it.
> >>>>>
> >>>>> Before we publish any native binary, we should ask LEGAL. A native
> >> image
> >>>>> contains SubstrateVM and JDK code under GPLv2 with the Classpath
> >> Exception,
> >>>>> and we'd also need one build per OS/arch as part of releases. As far
> >> as I
> >>>>> know, Kafka publishes a native Docker image (KIP-974), so there's at
> >> least
> >>>>> a precedent.
> >>>>>
> >>>>> Some of the expected benefits need benchmarks:
> >>>>>
> >>>>> 1. CE native image has no profile-guided optimization and no G1 GC.
> >> For
> >>>>> long-running pipelines, peak throughput may be lower than on HotSpot.
> >>>>> 2. Startup gets faster, but loading large models still takes time.
> >>>>> 3. GPU / CUDA / OpenVINO / NPU support comes from ONNX Runtime, not
> >> from
> >>>>> GraalVM. We already have that on the JVM.
> >>>>> 4. Rust/C++/Python bindings would mean designing and maintaining a C
> >> API
> >>>>> (@CEntryPoint, isolates, memory ownership). That's a separate
> >> project, not
> >>>>> a side effect of compiling natively.
> >>>>>
> >>>>> In the ASF EE world, we have Apache Geronimo Arthur: it's a thin
> >> Maven
> >>>>> layer over native-image. It downloads GraalVM, generates the
> >> native-image
> >>>>> config, can build images via Jib, and has an extension SPI
> >> ("knights") that
> >>>>> computes reflection/resource config at build time. It's ASF, so the
> >> people
> >>>>> behind it are well known in the EE ecosystem. That said, the last
> >> release
> >>>>> (1.0.9) was in April 2024 and the project has been quiet since. I
> >> wouldn't
> >>>>> make OpenNLP depend on it. Standard metadata in our jars works with
> >> Arthur
> >>>>> as-is. If there's actually interest, an "opennlp-knight" would be a
> >> small
> >>>>> add-on that registers the same things. For building the gRPC binary,
> >> I'd go
> >>>>> with the official GraalVM native-maven-plugin, which is actively
> >> maintained.
> >>>>>
> >>>>> TL;DR: making the library native-friendly is doable, but it's real
> >> work.
> >>>>> Shipping native binaries ourselves has no real gain as long as there
> >> is no
> >>>>> released gRPC server. For the CLI binary and language bindings,
> >> let's wait
> >>>>> for real demand.
> >>>>>
> >>>>> Richard
> >>>>>
> >>>>> On 2026/09/14 00:50:17 Kristian Rickert wrote:
> >>>>>> Devs,
> >>>>>>
> >>>>>> I would like to propose an initiative for a 3.x release (post-
> >> 3.0):
> >>>>>> compiling OpenNLP to native code using GraalVM.
> >>>>>>
> >>>>>> Here's the situation: Java is sandwiched between Python's data
> >> science
> >>>>>> dominance and Rust's performance. To combat this, I propose
> >> leveraging
> >>>>>> GraalVM to compile OpenNLP natively, which offers significant
> >> advantages.
> >>>>>> Having used it in production, I have seen it deliver instant
> >> startup
> >>>>> times,
> >>>>>> lower memory usage, and improved latency.
> >>>>>>
> >>>>>> Given OpenNLP's minimal dependencies and lack of reflection, it is
> >> a
> >>>>> prime
> >>>>>> candidate for this. Initial tests compiling to native code have
> >> yielded
> >>>>> no
> >>>>>> major issues.
> >>>>>>
> >>>>>> Some Pros:
> >>>>>>
> >>>>>>   - Broader Integration: We can package OpenNLP as a Rust crate
> >> or C++
> >>>>>>   library, allowing direct integration into applications, word
> >>>>> processors,
> >>>>>>   and Python (via Cython).
> >>>>>>   - Cross-Language Native Support: OpenNLP could be used natively
> >> in
> >>>>> Rust,
> >>>>>>   C++, Swift, and Python with a much smaller memory footprint.
> >>>>>>   - Performance Gains: By leveraging the pluggable embedding layer
> >>>>> created
> >>>>>>   for the gRPC service, embedding performance could be at least 2x
> >>>>> faster
> >>>>>>   (via GPU or static table creation).
> >>>>>>   - Wide Architecture Support: Native support for Apple Silicon,
> >> Intel
> >>>>>>   NPU, CUDA, OpenVINO, Android, and CPU execution.
> >>>>>>
> >>>>>> Questions for the Team:
> >>>>>>
> >>>>>>   1. Does anyone know of other Apache projects currently using
> >> GraalVM
> >>>>>>   compilation? If so, please reach out directly, I'd love to
> >> connect
> >>>>> with
> >>>>>>   them.
> >>>>>>   2. Do we have any connections with Oracle folks?  They create
> >> it, if I
> >>>>>>   run into issues, having them available to help would be
> >> beneficial.
> >>>>> (Note:
> >>>>>>   We would use the CE edition)
> >>>>>>   3. Are there any constraints / issues this can cause?
> >>>>>>   4. This can be a downstream build, and I'd volunteer to set up
> >> the
> >>>>> CICD
> >>>>>>   for it.  Anyone up for helping?  It can't hurt to understand
> >> Java
> >>>>> native
> >>>>>>   compilations.
> >>>>>>   5. Obviously, I'd set this up in sandbox and it'll be post-gRPC
> >> (I was
> >>>>>>   planning on natively compiling the gRPC server anyway)
> >>>>>>
> >>>>>> If you're interested in helping with this experiment, please let
> >> me know!
> >>>>>>
> >>>>>> Mutant test rungs,
> >>>>>> Kristian
> >>>>>>
> >>>>>
> >>>>
> >>
>
> Disclaimer
>
> The information contained in this communication from the sender is
> confidential. It is intended solely for use by the recipient and others
> authorized to receive it. If you are not the recipient, you are hereby
> notified that any disclosure, copying, distribution or taking action in
> relation of the contents of this information is strictly prohibited and may
> be unlawful.
>
> This email has been scanned for viruses and malware, and may have been
> automatically archived by Mimecast, a leader in email security and cyber
> resilience. Mimecast integrates email defenses with brand protection,
> security awareness training, web security, compliance and other essential
> capabilities. Mimecast helps protect large and small organizations from
> malicious activity, human error and technology failure; and to lead the
> movement toward building a more resilient world. To find out more, visit
> our website.
>

Reply via email to