Good discussion all.

So it sounds like we need a ticket to close gaps to make OpenNLP more
native-friendly and go from there? Probably heavy on user
documentation.

Eric, was that the most recent work with the ONNX models?

Thanks,
Jeff

On Tue, Sep 15, 2026 at 12:58 PM Richard Zowalla <[email protected]> wrote:
>
> Hi,
>
> to the delivery question: not really, and that is the point I was trying to 
> make earlier.
>
> Native compilation only makes sense if there is something to run at the end: 
> an application, a CLI, or a service shipped as a container. If you don't have 
> a binary to ship or an app that runs it, there is no need to compile a 
> LIBRARY natively. The only thing a library can do is be CAPABLE of being used 
> in a NATIVE build.
>
> Eric's two examples show exactly this:
>
> - solr-mcp is an application, and it ships as a native container image.
> - Tika does not ship native binaries. The Quarkus extension makes Tika usable 
> in a native build, and the Quarkus application that uses Tika is what gets 
> compiled.
>
> For OpenNLP this means there is nothing for us to compile natively. What we 
> can do is make the jars native-friendly:
>
> - ship reachability metadata in our jars (META-INF/native-image/...)
> - fix or replace the reflection-based code paths (factories, model loading, 
> classpath scanning)
> - document edge cases (like classpath scanning and recommend to use the 
> file-based approach instead for native use cases)
> - add a native smoke test in CI (tokenize, POS tag, NER)
>
> That work belongs in the main repo, not in the sandbox. Otherwise users still 
> can't build a working image.
>
> The one exception is a native shared library (native-image --shared), which 
> is what Rust/C/C++/Swift consumers would need. But that isn't "compiling 
> OpenNLP natively" either. This would be a new product with its own C API 
> (@CEntryPoint, isolates, memory ownership, error handling), one build per 
> OS/arch, and a LEGAL review before we distribute any binary. I'd treat that 
> as a separate proposal once the project decides to invest effort into the 
> library side.
>
> On Eric's Solr observation: native image won't fix that. It improves startup 
> time and memory footprint, not sustained throughput. The v3 work (thread 
> safety, removing regex) is what addresses that case far better.
>
> Gruß
> Richard
>
> On 2026/09/15 15:18:34 Kristian Rickert wrote:
> > Hi Eric,
> >
> > Thank you, Eric! This is incredibly helpful.
> >
> > I hope our speed improvements for v3 prove otherwise. Here is a quick
> > rundown of what we have been working hard on:
> >
> >   - Most of the codebase is now thread-safe (showing about a 2.5x
> > improvement).
> >   - We are removing all regex from the code for better performance and
> > fewer bugs.
> >   - We have a much cleaner interface and full UTF support.
> >   - We added several internationalization features and bi-directional
> > normalization for accurate highlighting.
> >   - We added a pile of new features to addons, including a PR for a Java
> > trainer and static embeddings (which can produce 100k/s on many systems).
> >
> > It is our mission (and my obsession) to make OpenNLP performant and on par
> > with popular NLP packages in other ecosystems.
> >
> > So it sounds like GraalVM was a great direction to explore. I have a ticket
> > open for it now. Outside of containers, is anyone delivering native
> > binaries for libraries?
> >
> > Rungs,
> > Kristian
> >
> >
> > On Tue, Sep 15, 2026 at 7:28 AM Eric Pugh <[email protected]>
> > wrote:
> >
> > > I’ve been really intrigued with GraalVM….     A couple of observations:
> > >
> > > 1) We are using it in the forthcoming solr-mcp tool to drastically speed
> > > up how fast solr-cmp starts when used as a command line tool.  Solr-MCP
> > > leverages Spring and Spring AI, and as you can imagine, there is a lot of
> > > wiring dynamically going on in that.  GraalVM changes that by “pre
> > > compiling” (not sure that is the right word) all of that so you get super
> > > fast start up.  We ship the GraalVM version in a dockerized version of
> > > solr-mcp so you don’t need to install anything locally.  And yeah, it adds
> > > some build complexity on top to get that to work, but then you are done.
> > >
> > > 2) Apache Tika actually runs on GraalVM!
> > > https://docs.quarkiverse.io/quarkus-tika/dev/index.html.  Tika is an
> > > example of a project that has a million dependencies, so if it can run on
> > > GraalVM….
> > >
> > > 3) OpenNLP is definitely bounded by performance….  When we added it to
> > > Solr as a UpdateRequestProcessor, any model beyond the most simplistic 
> > > made
> > > the indexing pace almost unusuable, which means you need to have a primary
> > > indexing path that doesn’t touch OpenNLP and then a secondary “enrichment”
> > > one that comes back async and does the OpenNLP work…
> > >
> > >
> > >
> > > > On Sep 14, 2026, at 10:18 PM, Kristian Rickert <[email protected]>
> > > wrote:
> > > >
> > > > *Is something that can be achieved via a Maven profile?*
> > > >
> > > > Yes, this is typically done using the native-maven-plugin provided by
> > > > GraalVM.
> > > >
> > > > *Is this build done via the project using OpenNLP, or does OpenNLP 
> > > > itself
> > > > have to be compiled differently?*
> > > >
> > > > OpenNLP itself does not need to be compiled differently. The library is
> > > > still compiled to standard Java bytecode just like normal. Native
> > > > compilation happens at the final application or distribution level.
> > > >
> > > > I'll give you my 2 cents since I've spent the last hour refreshing 
> > > > myself
> > > > on it.
> > > >
> > > > First I want to state my motivation: I want to create a native library
> > > for
> > > > apps calling OpenNLP (gRPC is just one example).  I'm convinced what we
> > > are
> > > > building can be integrated into many text-heavy apps in Rust, C/C++, or
> > > > Swift. It could help word processing apps, AI runtime analysis, and RAG
> > > > pipelines everyone is hyping about.  Basically, between this and gRPC,
> > > our
> > > > library can reach any app.
> > > >
> > > > The tool that builds the binaries - the GraalVM Native Image - analyzes
> > > the
> > > > application entry point and traces all reachable code across all
> > > > dependencies to produce a single binary executable. This is similar to
> > > what
> > > > you would get with gcc, and it comes with its own set of advantages and
> > > > headaches.
> > > >
> > > > I don't understand the magic of this process BUT it sure sounds cool and
> > > is
> > > > a low level java thing to learn. Richard's video he references goes over
> > > > this entire process in detail, and I plan to watch it (it's a great
> > > > presentation - I'm a fan of his font choice). He explains the 
> > > > compilation
> > > > process and the performance implications without trying to sell you on
> > > the
> > > > technology either way.
> > > >
> > > > Everything Richard said is right, and he is rightfully questioning my
> > > speed
> > > > claims. Native Image gives you near instant startup time and a very low
> > > > initial memory footprint. However, a traditional JVM JIT compiler might
> > > > actually beat it in peak sustained throughput for long running tasks. I
> > > > have not claimed with certainty that it will be faster for raw
> > > > computational throughput, but I'm enthusiastic and fueled by hope that 
> > > > it
> > > > would. We will have to measure the performance and test it extensively.
> > > >
> > > > Most modern Java tries to account for this. Both Quarkus and Micronaut
> > > are
> > > > built to work with native compilation out of the box.  It's why I use
> > > > Quarkus in a lot of my projects - a premature optimization I've not
> > > needed
> > > > yet but can save your AWS bills by easy double digits.  When you get a
> > > > multi-million dollar bill from Bezos, this suddenly becomes a more
> > > > attractive choice :)
> > > >
> > > > For OpenNLP, almost all the code should work fine. The biggest friction
> > > > point will be the model loading. We need to make sure any models loaded
> > > as
> > > > classpath resources or anything relying on Java serialization are
> > > > explicitly registered in the Native Image configuration files or we can
> > > > refactor the code to load services through an SPI loader.  I'm a big fan
> > > of
> > > > SPI loading, so I'd lean toward that.
> > > >
> > > > Last but not least, since this is Oracle's baby, they own GraalVM, and
> > > > there's a community edition.  I've only used the CE, because the Oracle
> > > one
> > > > is a commercial product.  You can use it for free in the Oracle cloud -
> > > but
> > > > I don't want to be "that guy" in our org to suggest it.
> > > >
> > > > Initially, I'd only want to support compiling it and leave the various
> > > arch
> > > > binaries up to the user.  IMHO anyone who uses GraalVM right now doesn't
> > > > look for native binaries, they just build the binaries themselves or
> > > might
> > > > use a native docker container instead (I've never seen a java native
> > > > library in the wild that isn't a container).  We'll make a lot of DevOps
> > > > folks happy if we can advertise that we are GaalVM compliant - if they
> > > use
> > > > it.  To Richard's point, the other project he referenced hasn't seen 
> > > > much
> > > > action from it.  By doing this, we learn a lot, potentially offer
> > > something
> > > > that integrates into a larger ecosystem, and if it fails we remove it
> > > from
> > > > the build.
> > > >
> > > >
> > > > On Mon, Sep 14, 2026, 4:29 PM Jeff Zemerick <[email protected]>
> > > wrote:
> > > >
> > > >> I'm not super familiar with GraalVM or what OpenNLP requires to be
> > > >> compatible with it but I understand the desire to do so.
> > > >>
> > > >> A couple questions:
> > > >>
> > > >> - Is something that can be achieved via a Maven profile?
> > > >> - Is this build done via the project using OpenNLP, or does OpenNLP
> > > >> itself have to be compiled differently?
> > > >>
> > > >> Thanks,
> > > >> Jeff
> > > >>
> > > >> On Mon, Sep 14, 2026 at 8:21 AM Richard Zowalla <[email protected]>
> > > wrote:
> > > >>>
> > > >>> Hi,
> > > >>>
> > > >>> if the deliverable is "build it yourself", the actual product is the
> > > >> reachability metadata in our jars and the reflection fixes, and that
> > > can't
> > > >> live in the sandbox. It has to go into the main repo, or users still
> > > can't
> > > >> build a working image. The sandbox can hold the recipe and the gRPC
> > > binary
> > > >> experiment, but the enabling work is main-repo work and needs to be
> > > scoped
> > > >> as such.
> > > >>>
> > > >>> On the demand argument, note that native image on CE gives startup and
> > > >> memory footprint (I recommend "Scotty I need Warp Speed" from Gerrit
> > > >> Grunwald) not peak throughput (no PGO, no G1). For long-running
> > > validation
> > > >> pipelines a warm JIT may still win.
> > > >>>
> > > >>> Richard
> > > >>>
> > > >>> On 2026/09/14 11:53:20 Kristian Rickert wrote:
> > > >>>> Hi Richard,
> > > >>>>
> > > >>>> Thank you for the detailed response. It answered all of my questions
> > > >> and
> > > >>>> provides a clear path forward with great historical context we should
> > > >> heed.
> > > >>>>
> > > >>>> I agree with your technical assessment, but see a subtle difference
> > > >>>> regarding demand: I believe proactive innovation drives demand.  I
> > > >> call it
> > > >>>> the Kevin Kostner effect: "if you build it, they will come".
> > > >>>>
> > > >>>> But I also hope to drum up collaboration in this thread, as an
> > > >> opportunity
> > > >>>> to learn exciting, potential-filled new technologies. Adding native
> > > >>>> compilation could open new channels for library integration.  
> > > >>>> Learning
> > > >> this
> > > >>>> is a great resume builder and offers experience in profiling and
> > > >> testing
> > > >>>> that exceeds our current scope.  And it's low risk but takes a
> > > >> calculated
> > > >>>> chance.  This is why it'll remain in the sandbox until the
> > > >>>> build-then-demand model is proven.
> > > >>>>
> > > >>>> But to gain demand - that's when blog posts, demos, and conference
> > > chit
> > > >>>> chat come in.  So I promise to focus most of my efforts there once it
> > > >> is
> > > >>>> built.
> > > >>>>
> > > >>>> Here is why I think it'll work:
> > > >>>>
> > > >>>> NLP has been overshadowed by LLM hype, but as the need for fast
> > > >>>> ground-truth validation grows, performance will be key. People
> > > unfairly
> > > >>>> separate LLMs, search, and NLP when they all deal with language and
> > > >> they
> > > >>>> need better integration.  Since Java isn't natively fast (though it 
> > > >>>> is
> > > >>>> still significantly faster than Python), demonstrating GraalVM's 
> > > >>>> speed
> > > >> and
> > > >>>> integration capabilities could reignite interest in the ecosystem,
> > > >> though
> > > >>>> complex setup remains a risk.
> > > >>>>
> > > >>>> Finally I won't dismiss that maintaining native binaries is a 
> > > >>>> headache
> > > >> and
> > > >>>> I think we might want to avoid it.  CVE concerns given rapidly
> > > changing
> > > >>>> architectures compound this issue. Even if it leaves the sandbox,
> > > >> keeping
> > > >>>> this as a build-it-yourself setup may be the best approach. This
> > > avoids
> > > >>>> direct binary distribution given the complexity of the setup 
> > > >>>> (ARM/AMD,
> > > >>>> NVidia vs. Intel, CPU/GPU/NPU, and OS all need to be considered). If
> > > >> this
> > > >>>> experiment fails to gain traction, I fully support deprecating and
> > > >> retiring
> > > >>>> it quickly.
> > > >>>>
> > > >>>> Thanks again for the thorough write-up.
> > > >>>>
> > > >>>> Best,
> > > >>>> Kristian
> > > >>>>
> > > >>>>
> > > >>>>
> > > >>>> On Mon, Sep 14, 2026, 2:42 AM Richard Zowalla <[email protected]>
> > > wrote:
> > > >>>>
> > > >>>>> Hi,
> > > >>>>>
> > > >>>>> I'd frame the scope a bit differently, though. Here are my 2cts from
> > > >> my
> > > >>>>> work in the EE ecosystem.
> > > >>>>>
> > > >>>>> OpenNLP is a library. Compiling to native code is something the
> > > >> user's
> > > >>>>> application does, not something we do IMHO. What OpenNLP can do is
> > > >> make the
> > > >>>>> jars native-friendly, so that anyone building a native image (with
> > > >> the
> > > >>>>> GraalVM Maven/Gradle plugins, Quarkus, Spring AOT, Apache Geronimo
> > > >> Arthur,
> > > >>>>> ...) gets a working build without extra steps. Concretely:
> > > >>>>>
> > > >>>>> 1. Ship reachability metadata inside our jars
> > > >>>>> (META-INF/native-image/org.apache.opennlp/<artifactId>/...), or a
> > > >> small
> > > >>>>> GraalVM Feature class.
> > > >>>>> 2. Fix the parts that break in a native image. We do use reflection,
> > > >> and
> > > >>>>> it isn't only in edge cases:
> > > >>>>> (a) ExtensionLoader / BaseToolFactory / TrainerFactory load classes
> > > >> by
> > > >>>>> name. The factory class names even come from the model manifest.
> > > >>>>> (b) GenericModelReader/Writer and ClassPathModelLoader create
> > > >> objects
> > > >>>>> through reflective constructors.
> > > >>>>> (c) The Snowball stemmers look up methods via
> > > >> MethodHandles.findVirtual.
> > > >>>>> (d) The CLI's ArgumentParser uses dynamic proxies.
> > > >>>>> (e) SimpleClassPathModelFinder reads jdk.internal internals and
> > > >> scans
> > > >>>>> jars on the classpath. That can't work in a native image because
> > > >> there is
> > > >>>>> no classpath.
> > > >>>>> (f) opennlp-dl pulls in ONNX Runtime through JNI.
> > > >>>>> (g) We load a fair number of bundled resources.
> > > >>>>> 3. Do some experiments outside of the OpenNLP repos, like a native
> > > >> smoke
> > > >>>>> test in CI (tokenize, POS tag, NER).
> > > >>>>>
> > > >>>>> Custom factories or extensions from users will always need
> > > >> registering by
> > > >>>>> the user, and that's fine.
> > > >>>>>
> > > >>>>> On binary distributions: the (WIP) gRPC server would be the obvious
> > > >>>>> candidate for native. It's a standalone, long-running service, it
> > > >> lives in
> > > >>>>> the sandbox anyway, and a native container image makes sense there.
> > > >>>>>
> > > >>>>> A native CLI could be nice for startup time, but I have no idea how
> > > >> many
> > > >>>>> people actually use the CLI, so I'd wait until someone asks for it.
> > > >>>>>
> > > >>>>> Before we publish any native binary, we should ask LEGAL. A native
> > > >> image
> > > >>>>> contains SubstrateVM and JDK code under GPLv2 with the Classpath
> > > >> Exception,
> > > >>>>> and we'd also need one build per OS/arch as part of releases. As far
> > > >> as I
> > > >>>>> know, Kafka publishes a native Docker image (KIP-974), so there's at
> > > >> least
> > > >>>>> a precedent.
> > > >>>>>
> > > >>>>> Some of the expected benefits need benchmarks:
> > > >>>>>
> > > >>>>> 1. CE native image has no profile-guided optimization and no G1 GC.
> > > >> For
> > > >>>>> long-running pipelines, peak throughput may be lower than on 
> > > >>>>> HotSpot.
> > > >>>>> 2. Startup gets faster, but loading large models still takes time.
> > > >>>>> 3. GPU / CUDA / OpenVINO / NPU support comes from ONNX Runtime, not
> > > >> from
> > > >>>>> GraalVM. We already have that on the JVM.
> > > >>>>> 4. Rust/C++/Python bindings would mean designing and maintaining a C
> > > >> API
> > > >>>>> (@CEntryPoint, isolates, memory ownership). That's a separate
> > > >> project, not
> > > >>>>> a side effect of compiling natively.
> > > >>>>>
> > > >>>>> In the ASF EE world, we have Apache Geronimo Arthur: it's a thin
> > > >> Maven
> > > >>>>> layer over native-image. It downloads GraalVM, generates the
> > > >> native-image
> > > >>>>> config, can build images via Jib, and has an extension SPI
> > > >> ("knights") that
> > > >>>>> computes reflection/resource config at build time. It's ASF, so the
> > > >> people
> > > >>>>> behind it are well known in the EE ecosystem. That said, the last
> > > >> release
> > > >>>>> (1.0.9) was in April 2024 and the project has been quiet since. I
> > > >> wouldn't
> > > >>>>> make OpenNLP depend on it. Standard metadata in our jars works with
> > > >> Arthur
> > > >>>>> as-is. If there's actually interest, an "opennlp-knight" would be a
> > > >> small
> > > >>>>> add-on that registers the same things. For building the gRPC binary,
> > > >> I'd go
> > > >>>>> with the official GraalVM native-maven-plugin, which is actively
> > > >> maintained.
> > > >>>>>
> > > >>>>> TL;DR: making the library native-friendly is doable, but it's real
> > > >> work.
> > > >>>>> Shipping native binaries ourselves has no real gain as long as there
> > > >> is no
> > > >>>>> released gRPC server. For the CLI binary and language bindings,
> > > >> let's wait
> > > >>>>> for real demand.
> > > >>>>>
> > > >>>>> Richard
> > > >>>>>
> > > >>>>> On 2026/09/14 00:50:17 Kristian Rickert wrote:
> > > >>>>>> Devs,
> > > >>>>>>
> > > >>>>>> I would like to propose an initiative for a 3.x release (post-
> > > >> 3.0):
> > > >>>>>> compiling OpenNLP to native code using GraalVM.
> > > >>>>>>
> > > >>>>>> Here's the situation: Java is sandwiched between Python's data
> > > >> science
> > > >>>>>> dominance and Rust's performance. To combat this, I propose
> > > >> leveraging
> > > >>>>>> GraalVM to compile OpenNLP natively, which offers significant
> > > >> advantages.
> > > >>>>>> Having used it in production, I have seen it deliver instant
> > > >> startup
> > > >>>>> times,
> > > >>>>>> lower memory usage, and improved latency.
> > > >>>>>>
> > > >>>>>> Given OpenNLP's minimal dependencies and lack of reflection, it is
> > > >> a
> > > >>>>> prime
> > > >>>>>> candidate for this. Initial tests compiling to native code have
> > > >> yielded
> > > >>>>> no
> > > >>>>>> major issues.
> > > >>>>>>
> > > >>>>>> Some Pros:
> > > >>>>>>
> > > >>>>>>   - Broader Integration: We can package OpenNLP as a Rust crate
> > > >> or C++
> > > >>>>>>   library, allowing direct integration into applications, word
> > > >>>>> processors,
> > > >>>>>>   and Python (via Cython).
> > > >>>>>>   - Cross-Language Native Support: OpenNLP could be used natively
> > > >> in
> > > >>>>> Rust,
> > > >>>>>>   C++, Swift, and Python with a much smaller memory footprint.
> > > >>>>>>   - Performance Gains: By leveraging the pluggable embedding layer
> > > >>>>> created
> > > >>>>>>   for the gRPC service, embedding performance could be at least 2x
> > > >>>>> faster
> > > >>>>>>   (via GPU or static table creation).
> > > >>>>>>   - Wide Architecture Support: Native support for Apple Silicon,
> > > >> Intel
> > > >>>>>>   NPU, CUDA, OpenVINO, Android, and CPU execution.
> > > >>>>>>
> > > >>>>>> Questions for the Team:
> > > >>>>>>
> > > >>>>>>   1. Does anyone know of other Apache projects currently using
> > > >> GraalVM
> > > >>>>>>   compilation? If so, please reach out directly, I'd love to
> > > >> connect
> > > >>>>> with
> > > >>>>>>   them.
> > > >>>>>>   2. Do we have any connections with Oracle folks?  They create
> > > >> it, if I
> > > >>>>>>   run into issues, having them available to help would be
> > > >> beneficial.
> > > >>>>> (Note:
> > > >>>>>>   We would use the CE edition)
> > > >>>>>>   3. Are there any constraints / issues this can cause?
> > > >>>>>>   4. This can be a downstream build, and I'd volunteer to set up
> > > >> the
> > > >>>>> CICD
> > > >>>>>>   for it.  Anyone up for helping?  It can't hurt to understand
> > > >> Java
> > > >>>>> native
> > > >>>>>>   compilations.
> > > >>>>>>   5. Obviously, I'd set this up in sandbox and it'll be post-gRPC
> > > >> (I was
> > > >>>>>>   planning on natively compiling the gRPC server anyway)
> > > >>>>>>
> > > >>>>>> If you're interested in helping with this experiment, please let
> > > >> me know!
> > > >>>>>>
> > > >>>>>> Mutant test rungs,
> > > >>>>>> Kristian
> > > >>>>>>
> > > >>>>>
> > > >>>>
> > > >>
> > >
> > > Disclaimer
> > >
> > > The information contained in this communication from the sender is
> > > confidential. It is intended solely for use by the recipient and others
> > > authorized to receive it. If you are not the recipient, you are hereby
> > > notified that any disclosure, copying, distribution or taking action in
> > > relation of the contents of this information is strictly prohibited and 
> > > may
> > > be unlawful.
> > >
> > > This email has been scanned for viruses and malware, and may have been
> > > automatically archived by Mimecast, a leader in email security and cyber
> > > resilience. Mimecast integrates email defenses with brand protection,
> > > security awareness training, web security, compliance and other essential
> > > capabilities. Mimecast helps protect large and small organizations from
> > > malicious activity, human error and technology failure; and to lead the
> > > movement toward building a more resilient world. To find out more, visit
> > > our website.
> > >
> >

Reply via email to