Hi,

to the delivery question: not really, and that is the point I was trying to 
make earlier.

Native compilation only makes sense if there is something to run at the end: an 
application, a CLI, or a service shipped as a container. If you don't have a 
binary to ship or an app that runs it, there is no need to compile a LIBRARY 
natively. The only thing a library can do is be CAPABLE of being used in a 
NATIVE build.

Eric's two examples show exactly this:

- solr-mcp is an application, and it ships as a native container image.
- Tika does not ship native binaries. The Quarkus extension makes Tika usable 
in a native build, and the Quarkus application that uses Tika is what gets 
compiled.

For OpenNLP this means there is nothing for us to compile natively. What we can 
do is make the jars native-friendly:

- ship reachability metadata in our jars (META-INF/native-image/...)
- fix or replace the reflection-based code paths (factories, model loading, 
classpath scanning)
- document edge cases (like classpath scanning and recommend to use the 
file-based approach instead for native use cases)
- add a native smoke test in CI (tokenize, POS tag, NER)

That work belongs in the main repo, not in the sandbox. Otherwise users still 
can't build a working image.

The one exception is a native shared library (native-image --shared), which is 
what Rust/C/C++/Swift consumers would need. But that isn't "compiling OpenNLP 
natively" either. This would be a new product with its own C API (@CEntryPoint, 
isolates, memory ownership, error handling), one build per OS/arch, and a LEGAL 
review before we distribute any binary. I'd treat that as a separate proposal 
once the project decides to invest effort into the library side.

On Eric's Solr observation: native image won't fix that. It improves startup 
time and memory footprint, not sustained throughput. The v3 work (thread 
safety, removing regex) is what addresses that case far better.

Gruß
Richard

On 2026/09/15 15:18:34 Kristian Rickert wrote:
> Hi Eric,
> 
> Thank you, Eric! This is incredibly helpful.
> 
> I hope our speed improvements for v3 prove otherwise. Here is a quick
> rundown of what we have been working hard on:
> 
>   - Most of the codebase is now thread-safe (showing about a 2.5x
> improvement).
>   - We are removing all regex from the code for better performance and
> fewer bugs.
>   - We have a much cleaner interface and full UTF support.
>   - We added several internationalization features and bi-directional
> normalization for accurate highlighting.
>   - We added a pile of new features to addons, including a PR for a Java
> trainer and static embeddings (which can produce 100k/s on many systems).
> 
> It is our mission (and my obsession) to make OpenNLP performant and on par
> with popular NLP packages in other ecosystems.
> 
> So it sounds like GraalVM was a great direction to explore. I have a ticket
> open for it now. Outside of containers, is anyone delivering native
> binaries for libraries?
> 
> Rungs,
> Kristian
> 
> 
> On Tue, Sep 15, 2026 at 7:28 AM Eric Pugh <[email protected]>
> wrote:
> 
> > I’ve been really intrigued with GraalVM….     A couple of observations:
> >
> > 1) We are using it in the forthcoming solr-mcp tool to drastically speed
> > up how fast solr-cmp starts when used as a command line tool.  Solr-MCP
> > leverages Spring and Spring AI, and as you can imagine, there is a lot of
> > wiring dynamically going on in that.  GraalVM changes that by “pre
> > compiling” (not sure that is the right word) all of that so you get super
> > fast start up.  We ship the GraalVM version in a dockerized version of
> > solr-mcp so you don’t need to install anything locally.  And yeah, it adds
> > some build complexity on top to get that to work, but then you are done.
> >
> > 2) Apache Tika actually runs on GraalVM!
> > https://docs.quarkiverse.io/quarkus-tika/dev/index.html.  Tika is an
> > example of a project that has a million dependencies, so if it can run on
> > GraalVM….
> >
> > 3) OpenNLP is definitely bounded by performance….  When we added it to
> > Solr as a UpdateRequestProcessor, any model beyond the most simplistic made
> > the indexing pace almost unusuable, which means you need to have a primary
> > indexing path that doesn’t touch OpenNLP and then a secondary “enrichment”
> > one that comes back async and does the OpenNLP work…
> >
> >
> >
> > > On Sep 14, 2026, at 10:18 PM, Kristian Rickert <[email protected]>
> > wrote:
> > >
> > > *Is something that can be achieved via a Maven profile?*
> > >
> > > Yes, this is typically done using the native-maven-plugin provided by
> > > GraalVM.
> > >
> > > *Is this build done via the project using OpenNLP, or does OpenNLP itself
> > > have to be compiled differently?*
> > >
> > > OpenNLP itself does not need to be compiled differently. The library is
> > > still compiled to standard Java bytecode just like normal. Native
> > > compilation happens at the final application or distribution level.
> > >
> > > I'll give you my 2 cents since I've spent the last hour refreshing myself
> > > on it.
> > >
> > > First I want to state my motivation: I want to create a native library
> > for
> > > apps calling OpenNLP (gRPC is just one example).  I'm convinced what we
> > are
> > > building can be integrated into many text-heavy apps in Rust, C/C++, or
> > > Swift. It could help word processing apps, AI runtime analysis, and RAG
> > > pipelines everyone is hyping about.  Basically, between this and gRPC,
> > our
> > > library can reach any app.
> > >
> > > The tool that builds the binaries - the GraalVM Native Image - analyzes
> > the
> > > application entry point and traces all reachable code across all
> > > dependencies to produce a single binary executable. This is similar to
> > what
> > > you would get with gcc, and it comes with its own set of advantages and
> > > headaches.
> > >
> > > I don't understand the magic of this process BUT it sure sounds cool and
> > is
> > > a low level java thing to learn. Richard's video he references goes over
> > > this entire process in detail, and I plan to watch it (it's a great
> > > presentation - I'm a fan of his font choice). He explains the compilation
> > > process and the performance implications without trying to sell you on
> > the
> > > technology either way.
> > >
> > > Everything Richard said is right, and he is rightfully questioning my
> > speed
> > > claims. Native Image gives you near instant startup time and a very low
> > > initial memory footprint. However, a traditional JVM JIT compiler might
> > > actually beat it in peak sustained throughput for long running tasks. I
> > > have not claimed with certainty that it will be faster for raw
> > > computational throughput, but I'm enthusiastic and fueled by hope that it
> > > would. We will have to measure the performance and test it extensively.
> > >
> > > Most modern Java tries to account for this. Both Quarkus and Micronaut
> > are
> > > built to work with native compilation out of the box.  It's why I use
> > > Quarkus in a lot of my projects - a premature optimization I've not
> > needed
> > > yet but can save your AWS bills by easy double digits.  When you get a
> > > multi-million dollar bill from Bezos, this suddenly becomes a more
> > > attractive choice :)
> > >
> > > For OpenNLP, almost all the code should work fine. The biggest friction
> > > point will be the model loading. We need to make sure any models loaded
> > as
> > > classpath resources or anything relying on Java serialization are
> > > explicitly registered in the Native Image configuration files or we can
> > > refactor the code to load services through an SPI loader.  I'm a big fan
> > of
> > > SPI loading, so I'd lean toward that.
> > >
> > > Last but not least, since this is Oracle's baby, they own GraalVM, and
> > > there's a community edition.  I've only used the CE, because the Oracle
> > one
> > > is a commercial product.  You can use it for free in the Oracle cloud -
> > but
> > > I don't want to be "that guy" in our org to suggest it.
> > >
> > > Initially, I'd only want to support compiling it and leave the various
> > arch
> > > binaries up to the user.  IMHO anyone who uses GraalVM right now doesn't
> > > look for native binaries, they just build the binaries themselves or
> > might
> > > use a native docker container instead (I've never seen a java native
> > > library in the wild that isn't a container).  We'll make a lot of DevOps
> > > folks happy if we can advertise that we are GaalVM compliant - if they
> > use
> > > it.  To Richard's point, the other project he referenced hasn't seen much
> > > action from it.  By doing this, we learn a lot, potentially offer
> > something
> > > that integrates into a larger ecosystem, and if it fails we remove it
> > from
> > > the build.
> > >
> > >
> > > On Mon, Sep 14, 2026, 4:29 PM Jeff Zemerick <[email protected]>
> > wrote:
> > >
> > >> I'm not super familiar with GraalVM or what OpenNLP requires to be
> > >> compatible with it but I understand the desire to do so.
> > >>
> > >> A couple questions:
> > >>
> > >> - Is something that can be achieved via a Maven profile?
> > >> - Is this build done via the project using OpenNLP, or does OpenNLP
> > >> itself have to be compiled differently?
> > >>
> > >> Thanks,
> > >> Jeff
> > >>
> > >> On Mon, Sep 14, 2026 at 8:21 AM Richard Zowalla <[email protected]>
> > wrote:
> > >>>
> > >>> Hi,
> > >>>
> > >>> if the deliverable is "build it yourself", the actual product is the
> > >> reachability metadata in our jars and the reflection fixes, and that
> > can't
> > >> live in the sandbox. It has to go into the main repo, or users still
> > can't
> > >> build a working image. The sandbox can hold the recipe and the gRPC
> > binary
> > >> experiment, but the enabling work is main-repo work and needs to be
> > scoped
> > >> as such.
> > >>>
> > >>> On the demand argument, note that native image on CE gives startup and
> > >> memory footprint (I recommend "Scotty I need Warp Speed" from Gerrit
> > >> Grunwald) not peak throughput (no PGO, no G1). For long-running
> > validation
> > >> pipelines a warm JIT may still win.
> > >>>
> > >>> Richard
> > >>>
> > >>> On 2026/09/14 11:53:20 Kristian Rickert wrote:
> > >>>> Hi Richard,
> > >>>>
> > >>>> Thank you for the detailed response. It answered all of my questions
> > >> and
> > >>>> provides a clear path forward with great historical context we should
> > >> heed.
> > >>>>
> > >>>> I agree with your technical assessment, but see a subtle difference
> > >>>> regarding demand: I believe proactive innovation drives demand.  I
> > >> call it
> > >>>> the Kevin Kostner effect: "if you build it, they will come".
> > >>>>
> > >>>> But I also hope to drum up collaboration in this thread, as an
> > >> opportunity
> > >>>> to learn exciting, potential-filled new technologies. Adding native
> > >>>> compilation could open new channels for library integration.  Learning
> > >> this
> > >>>> is a great resume builder and offers experience in profiling and
> > >> testing
> > >>>> that exceeds our current scope.  And it's low risk but takes a
> > >> calculated
> > >>>> chance.  This is why it'll remain in the sandbox until the
> > >>>> build-then-demand model is proven.
> > >>>>
> > >>>> But to gain demand - that's when blog posts, demos, and conference
> > chit
> > >>>> chat come in.  So I promise to focus most of my efforts there once it
> > >> is
> > >>>> built.
> > >>>>
> > >>>> Here is why I think it'll work:
> > >>>>
> > >>>> NLP has been overshadowed by LLM hype, but as the need for fast
> > >>>> ground-truth validation grows, performance will be key. People
> > unfairly
> > >>>> separate LLMs, search, and NLP when they all deal with language and
> > >> they
> > >>>> need better integration.  Since Java isn't natively fast (though it is
> > >>>> still significantly faster than Python), demonstrating GraalVM's speed
> > >> and
> > >>>> integration capabilities could reignite interest in the ecosystem,
> > >> though
> > >>>> complex setup remains a risk.
> > >>>>
> > >>>> Finally I won't dismiss that maintaining native binaries is a headache
> > >> and
> > >>>> I think we might want to avoid it.  CVE concerns given rapidly
> > changing
> > >>>> architectures compound this issue. Even if it leaves the sandbox,
> > >> keeping
> > >>>> this as a build-it-yourself setup may be the best approach. This
> > avoids
> > >>>> direct binary distribution given the complexity of the setup (ARM/AMD,
> > >>>> NVidia vs. Intel, CPU/GPU/NPU, and OS all need to be considered). If
> > >> this
> > >>>> experiment fails to gain traction, I fully support deprecating and
> > >> retiring
> > >>>> it quickly.
> > >>>>
> > >>>> Thanks again for the thorough write-up.
> > >>>>
> > >>>> Best,
> > >>>> Kristian
> > >>>>
> > >>>>
> > >>>>
> > >>>> On Mon, Sep 14, 2026, 2:42 AM Richard Zowalla <[email protected]>
> > wrote:
> > >>>>
> > >>>>> Hi,
> > >>>>>
> > >>>>> I'd frame the scope a bit differently, though. Here are my 2cts from
> > >> my
> > >>>>> work in the EE ecosystem.
> > >>>>>
> > >>>>> OpenNLP is a library. Compiling to native code is something the
> > >> user's
> > >>>>> application does, not something we do IMHO. What OpenNLP can do is
> > >> make the
> > >>>>> jars native-friendly, so that anyone building a native image (with
> > >> the
> > >>>>> GraalVM Maven/Gradle plugins, Quarkus, Spring AOT, Apache Geronimo
> > >> Arthur,
> > >>>>> ...) gets a working build without extra steps. Concretely:
> > >>>>>
> > >>>>> 1. Ship reachability metadata inside our jars
> > >>>>> (META-INF/native-image/org.apache.opennlp/<artifactId>/...), or a
> > >> small
> > >>>>> GraalVM Feature class.
> > >>>>> 2. Fix the parts that break in a native image. We do use reflection,
> > >> and
> > >>>>> it isn't only in edge cases:
> > >>>>> (a) ExtensionLoader / BaseToolFactory / TrainerFactory load classes
> > >> by
> > >>>>> name. The factory class names even come from the model manifest.
> > >>>>> (b) GenericModelReader/Writer and ClassPathModelLoader create
> > >> objects
> > >>>>> through reflective constructors.
> > >>>>> (c) The Snowball stemmers look up methods via
> > >> MethodHandles.findVirtual.
> > >>>>> (d) The CLI's ArgumentParser uses dynamic proxies.
> > >>>>> (e) SimpleClassPathModelFinder reads jdk.internal internals and
> > >> scans
> > >>>>> jars on the classpath. That can't work in a native image because
> > >> there is
> > >>>>> no classpath.
> > >>>>> (f) opennlp-dl pulls in ONNX Runtime through JNI.
> > >>>>> (g) We load a fair number of bundled resources.
> > >>>>> 3. Do some experiments outside of the OpenNLP repos, like a native
> > >> smoke
> > >>>>> test in CI (tokenize, POS tag, NER).
> > >>>>>
> > >>>>> Custom factories or extensions from users will always need
> > >> registering by
> > >>>>> the user, and that's fine.
> > >>>>>
> > >>>>> On binary distributions: the (WIP) gRPC server would be the obvious
> > >>>>> candidate for native. It's a standalone, long-running service, it
> > >> lives in
> > >>>>> the sandbox anyway, and a native container image makes sense there.
> > >>>>>
> > >>>>> A native CLI could be nice for startup time, but I have no idea how
> > >> many
> > >>>>> people actually use the CLI, so I'd wait until someone asks for it.
> > >>>>>
> > >>>>> Before we publish any native binary, we should ask LEGAL. A native
> > >> image
> > >>>>> contains SubstrateVM and JDK code under GPLv2 with the Classpath
> > >> Exception,
> > >>>>> and we'd also need one build per OS/arch as part of releases. As far
> > >> as I
> > >>>>> know, Kafka publishes a native Docker image (KIP-974), so there's at
> > >> least
> > >>>>> a precedent.
> > >>>>>
> > >>>>> Some of the expected benefits need benchmarks:
> > >>>>>
> > >>>>> 1. CE native image has no profile-guided optimization and no G1 GC.
> > >> For
> > >>>>> long-running pipelines, peak throughput may be lower than on HotSpot.
> > >>>>> 2. Startup gets faster, but loading large models still takes time.
> > >>>>> 3. GPU / CUDA / OpenVINO / NPU support comes from ONNX Runtime, not
> > >> from
> > >>>>> GraalVM. We already have that on the JVM.
> > >>>>> 4. Rust/C++/Python bindings would mean designing and maintaining a C
> > >> API
> > >>>>> (@CEntryPoint, isolates, memory ownership). That's a separate
> > >> project, not
> > >>>>> a side effect of compiling natively.
> > >>>>>
> > >>>>> In the ASF EE world, we have Apache Geronimo Arthur: it's a thin
> > >> Maven
> > >>>>> layer over native-image. It downloads GraalVM, generates the
> > >> native-image
> > >>>>> config, can build images via Jib, and has an extension SPI
> > >> ("knights") that
> > >>>>> computes reflection/resource config at build time. It's ASF, so the
> > >> people
> > >>>>> behind it are well known in the EE ecosystem. That said, the last
> > >> release
> > >>>>> (1.0.9) was in April 2024 and the project has been quiet since. I
> > >> wouldn't
> > >>>>> make OpenNLP depend on it. Standard metadata in our jars works with
> > >> Arthur
> > >>>>> as-is. If there's actually interest, an "opennlp-knight" would be a
> > >> small
> > >>>>> add-on that registers the same things. For building the gRPC binary,
> > >> I'd go
> > >>>>> with the official GraalVM native-maven-plugin, which is actively
> > >> maintained.
> > >>>>>
> > >>>>> TL;DR: making the library native-friendly is doable, but it's real
> > >> work.
> > >>>>> Shipping native binaries ourselves has no real gain as long as there
> > >> is no
> > >>>>> released gRPC server. For the CLI binary and language bindings,
> > >> let's wait
> > >>>>> for real demand.
> > >>>>>
> > >>>>> Richard
> > >>>>>
> > >>>>> On 2026/09/14 00:50:17 Kristian Rickert wrote:
> > >>>>>> Devs,
> > >>>>>>
> > >>>>>> I would like to propose an initiative for a 3.x release (post-
> > >> 3.0):
> > >>>>>> compiling OpenNLP to native code using GraalVM.
> > >>>>>>
> > >>>>>> Here's the situation: Java is sandwiched between Python's data
> > >> science
> > >>>>>> dominance and Rust's performance. To combat this, I propose
> > >> leveraging
> > >>>>>> GraalVM to compile OpenNLP natively, which offers significant
> > >> advantages.
> > >>>>>> Having used it in production, I have seen it deliver instant
> > >> startup
> > >>>>> times,
> > >>>>>> lower memory usage, and improved latency.
> > >>>>>>
> > >>>>>> Given OpenNLP's minimal dependencies and lack of reflection, it is
> > >> a
> > >>>>> prime
> > >>>>>> candidate for this. Initial tests compiling to native code have
> > >> yielded
> > >>>>> no
> > >>>>>> major issues.
> > >>>>>>
> > >>>>>> Some Pros:
> > >>>>>>
> > >>>>>>   - Broader Integration: We can package OpenNLP as a Rust crate
> > >> or C++
> > >>>>>>   library, allowing direct integration into applications, word
> > >>>>> processors,
> > >>>>>>   and Python (via Cython).
> > >>>>>>   - Cross-Language Native Support: OpenNLP could be used natively
> > >> in
> > >>>>> Rust,
> > >>>>>>   C++, Swift, and Python with a much smaller memory footprint.
> > >>>>>>   - Performance Gains: By leveraging the pluggable embedding layer
> > >>>>> created
> > >>>>>>   for the gRPC service, embedding performance could be at least 2x
> > >>>>> faster
> > >>>>>>   (via GPU or static table creation).
> > >>>>>>   - Wide Architecture Support: Native support for Apple Silicon,
> > >> Intel
> > >>>>>>   NPU, CUDA, OpenVINO, Android, and CPU execution.
> > >>>>>>
> > >>>>>> Questions for the Team:
> > >>>>>>
> > >>>>>>   1. Does anyone know of other Apache projects currently using
> > >> GraalVM
> > >>>>>>   compilation? If so, please reach out directly, I'd love to
> > >> connect
> > >>>>> with
> > >>>>>>   them.
> > >>>>>>   2. Do we have any connections with Oracle folks?  They create
> > >> it, if I
> > >>>>>>   run into issues, having them available to help would be
> > >> beneficial.
> > >>>>> (Note:
> > >>>>>>   We would use the CE edition)
> > >>>>>>   3. Are there any constraints / issues this can cause?
> > >>>>>>   4. This can be a downstream build, and I'd volunteer to set up
> > >> the
> > >>>>> CICD
> > >>>>>>   for it.  Anyone up for helping?  It can't hurt to understand
> > >> Java
> > >>>>> native
> > >>>>>>   compilations.
> > >>>>>>   5. Obviously, I'd set this up in sandbox and it'll be post-gRPC
> > >> (I was
> > >>>>>>   planning on natively compiling the gRPC server anyway)
> > >>>>>>
> > >>>>>> If you're interested in helping with this experiment, please let
> > >> me know!
> > >>>>>>
> > >>>>>> Mutant test rungs,
> > >>>>>> Kristian
> > >>>>>>
> > >>>>>
> > >>>>
> > >>
> >
> > Disclaimer
> >
> > The information contained in this communication from the sender is
> > confidential. It is intended solely for use by the recipient and others
> > authorized to receive it. If you are not the recipient, you are hereby
> > notified that any disclosure, copying, distribution or taking action in
> > relation of the contents of this information is strictly prohibited and may
> > be unlawful.
> >
> > This email has been scanned for viruses and malware, and may have been
> > automatically archived by Mimecast, a leader in email security and cyber
> > resilience. Mimecast integrates email defenses with brand protection,
> > security awareness training, web security, compliance and other essential
> > capabilities. Mimecast helps protect large and small organizations from
> > malicious activity, human error and technology failure; and to lead the
> > movement toward building a more resilient world. To find out more, visit
> > our website.
> >
> 

Reply via email to