Good discussion all. So it sounds like we need a ticket to close gaps to make OpenNLP more native-friendly and go from there? Probably heavy on user documentation.
Eric, was that the most recent work with the ONNX models? Thanks, Jeff On Tue, Sep 15, 2026 at 12:58 PM Richard Zowalla <[email protected]> wrote: > > Hi, > > to the delivery question: not really, and that is the point I was trying to > make earlier. > > Native compilation only makes sense if there is something to run at the end: > an application, a CLI, or a service shipped as a container. If you don't have > a binary to ship or an app that runs it, there is no need to compile a > LIBRARY natively. The only thing a library can do is be CAPABLE of being used > in a NATIVE build. > > Eric's two examples show exactly this: > > - solr-mcp is an application, and it ships as a native container image. > - Tika does not ship native binaries. The Quarkus extension makes Tika usable > in a native build, and the Quarkus application that uses Tika is what gets > compiled. > > For OpenNLP this means there is nothing for us to compile natively. What we > can do is make the jars native-friendly: > > - ship reachability metadata in our jars (META-INF/native-image/...) > - fix or replace the reflection-based code paths (factories, model loading, > classpath scanning) > - document edge cases (like classpath scanning and recommend to use the > file-based approach instead for native use cases) > - add a native smoke test in CI (tokenize, POS tag, NER) > > That work belongs in the main repo, not in the sandbox. Otherwise users still > can't build a working image. > > The one exception is a native shared library (native-image --shared), which > is what Rust/C/C++/Swift consumers would need. But that isn't "compiling > OpenNLP natively" either. This would be a new product with its own C API > (@CEntryPoint, isolates, memory ownership, error handling), one build per > OS/arch, and a LEGAL review before we distribute any binary. I'd treat that > as a separate proposal once the project decides to invest effort into the > library side. > > On Eric's Solr observation: native image won't fix that. It improves startup > time and memory footprint, not sustained throughput. The v3 work (thread > safety, removing regex) is what addresses that case far better. > > Gruß > Richard > > On 2026/09/15 15:18:34 Kristian Rickert wrote: > > Hi Eric, > > > > Thank you, Eric! This is incredibly helpful. > > > > I hope our speed improvements for v3 prove otherwise. Here is a quick > > rundown of what we have been working hard on: > > > > - Most of the codebase is now thread-safe (showing about a 2.5x > > improvement). > > - We are removing all regex from the code for better performance and > > fewer bugs. > > - We have a much cleaner interface and full UTF support. > > - We added several internationalization features and bi-directional > > normalization for accurate highlighting. > > - We added a pile of new features to addons, including a PR for a Java > > trainer and static embeddings (which can produce 100k/s on many systems). > > > > It is our mission (and my obsession) to make OpenNLP performant and on par > > with popular NLP packages in other ecosystems. > > > > So it sounds like GraalVM was a great direction to explore. I have a ticket > > open for it now. Outside of containers, is anyone delivering native > > binaries for libraries? > > > > Rungs, > > Kristian > > > > > > On Tue, Sep 15, 2026 at 7:28 AM Eric Pugh <[email protected]> > > wrote: > > > > > I’ve been really intrigued with GraalVM…. A couple of observations: > > > > > > 1) We are using it in the forthcoming solr-mcp tool to drastically speed > > > up how fast solr-cmp starts when used as a command line tool. Solr-MCP > > > leverages Spring and Spring AI, and as you can imagine, there is a lot of > > > wiring dynamically going on in that. GraalVM changes that by “pre > > > compiling” (not sure that is the right word) all of that so you get super > > > fast start up. We ship the GraalVM version in a dockerized version of > > > solr-mcp so you don’t need to install anything locally. And yeah, it adds > > > some build complexity on top to get that to work, but then you are done. > > > > > > 2) Apache Tika actually runs on GraalVM! > > > https://docs.quarkiverse.io/quarkus-tika/dev/index.html. Tika is an > > > example of a project that has a million dependencies, so if it can run on > > > GraalVM…. > > > > > > 3) OpenNLP is definitely bounded by performance…. When we added it to > > > Solr as a UpdateRequestProcessor, any model beyond the most simplistic > > > made > > > the indexing pace almost unusuable, which means you need to have a primary > > > indexing path that doesn’t touch OpenNLP and then a secondary “enrichment” > > > one that comes back async and does the OpenNLP work… > > > > > > > > > > > > > On Sep 14, 2026, at 10:18 PM, Kristian Rickert <[email protected]> > > > wrote: > > > > > > > > *Is something that can be achieved via a Maven profile?* > > > > > > > > Yes, this is typically done using the native-maven-plugin provided by > > > > GraalVM. > > > > > > > > *Is this build done via the project using OpenNLP, or does OpenNLP > > > > itself > > > > have to be compiled differently?* > > > > > > > > OpenNLP itself does not need to be compiled differently. The library is > > > > still compiled to standard Java bytecode just like normal. Native > > > > compilation happens at the final application or distribution level. > > > > > > > > I'll give you my 2 cents since I've spent the last hour refreshing > > > > myself > > > > on it. > > > > > > > > First I want to state my motivation: I want to create a native library > > > for > > > > apps calling OpenNLP (gRPC is just one example). I'm convinced what we > > > are > > > > building can be integrated into many text-heavy apps in Rust, C/C++, or > > > > Swift. It could help word processing apps, AI runtime analysis, and RAG > > > > pipelines everyone is hyping about. Basically, between this and gRPC, > > > our > > > > library can reach any app. > > > > > > > > The tool that builds the binaries - the GraalVM Native Image - analyzes > > > the > > > > application entry point and traces all reachable code across all > > > > dependencies to produce a single binary executable. This is similar to > > > what > > > > you would get with gcc, and it comes with its own set of advantages and > > > > headaches. > > > > > > > > I don't understand the magic of this process BUT it sure sounds cool and > > > is > > > > a low level java thing to learn. Richard's video he references goes over > > > > this entire process in detail, and I plan to watch it (it's a great > > > > presentation - I'm a fan of his font choice). He explains the > > > > compilation > > > > process and the performance implications without trying to sell you on > > > the > > > > technology either way. > > > > > > > > Everything Richard said is right, and he is rightfully questioning my > > > speed > > > > claims. Native Image gives you near instant startup time and a very low > > > > initial memory footprint. However, a traditional JVM JIT compiler might > > > > actually beat it in peak sustained throughput for long running tasks. I > > > > have not claimed with certainty that it will be faster for raw > > > > computational throughput, but I'm enthusiastic and fueled by hope that > > > > it > > > > would. We will have to measure the performance and test it extensively. > > > > > > > > Most modern Java tries to account for this. Both Quarkus and Micronaut > > > are > > > > built to work with native compilation out of the box. It's why I use > > > > Quarkus in a lot of my projects - a premature optimization I've not > > > needed > > > > yet but can save your AWS bills by easy double digits. When you get a > > > > multi-million dollar bill from Bezos, this suddenly becomes a more > > > > attractive choice :) > > > > > > > > For OpenNLP, almost all the code should work fine. The biggest friction > > > > point will be the model loading. We need to make sure any models loaded > > > as > > > > classpath resources or anything relying on Java serialization are > > > > explicitly registered in the Native Image configuration files or we can > > > > refactor the code to load services through an SPI loader. I'm a big fan > > > of > > > > SPI loading, so I'd lean toward that. > > > > > > > > Last but not least, since this is Oracle's baby, they own GraalVM, and > > > > there's a community edition. I've only used the CE, because the Oracle > > > one > > > > is a commercial product. You can use it for free in the Oracle cloud - > > > but > > > > I don't want to be "that guy" in our org to suggest it. > > > > > > > > Initially, I'd only want to support compiling it and leave the various > > > arch > > > > binaries up to the user. IMHO anyone who uses GraalVM right now doesn't > > > > look for native binaries, they just build the binaries themselves or > > > might > > > > use a native docker container instead (I've never seen a java native > > > > library in the wild that isn't a container). We'll make a lot of DevOps > > > > folks happy if we can advertise that we are GaalVM compliant - if they > > > use > > > > it. To Richard's point, the other project he referenced hasn't seen > > > > much > > > > action from it. By doing this, we learn a lot, potentially offer > > > something > > > > that integrates into a larger ecosystem, and if it fails we remove it > > > from > > > > the build. > > > > > > > > > > > > On Mon, Sep 14, 2026, 4:29 PM Jeff Zemerick <[email protected]> > > > wrote: > > > > > > > >> I'm not super familiar with GraalVM or what OpenNLP requires to be > > > >> compatible with it but I understand the desire to do so. > > > >> > > > >> A couple questions: > > > >> > > > >> - Is something that can be achieved via a Maven profile? > > > >> - Is this build done via the project using OpenNLP, or does OpenNLP > > > >> itself have to be compiled differently? > > > >> > > > >> Thanks, > > > >> Jeff > > > >> > > > >> On Mon, Sep 14, 2026 at 8:21 AM Richard Zowalla <[email protected]> > > > wrote: > > > >>> > > > >>> Hi, > > > >>> > > > >>> if the deliverable is "build it yourself", the actual product is the > > > >> reachability metadata in our jars and the reflection fixes, and that > > > can't > > > >> live in the sandbox. It has to go into the main repo, or users still > > > can't > > > >> build a working image. The sandbox can hold the recipe and the gRPC > > > binary > > > >> experiment, but the enabling work is main-repo work and needs to be > > > scoped > > > >> as such. > > > >>> > > > >>> On the demand argument, note that native image on CE gives startup and > > > >> memory footprint (I recommend "Scotty I need Warp Speed" from Gerrit > > > >> Grunwald) not peak throughput (no PGO, no G1). For long-running > > > validation > > > >> pipelines a warm JIT may still win. > > > >>> > > > >>> Richard > > > >>> > > > >>> On 2026/09/14 11:53:20 Kristian Rickert wrote: > > > >>>> Hi Richard, > > > >>>> > > > >>>> Thank you for the detailed response. It answered all of my questions > > > >> and > > > >>>> provides a clear path forward with great historical context we should > > > >> heed. > > > >>>> > > > >>>> I agree with your technical assessment, but see a subtle difference > > > >>>> regarding demand: I believe proactive innovation drives demand. I > > > >> call it > > > >>>> the Kevin Kostner effect: "if you build it, they will come". > > > >>>> > > > >>>> But I also hope to drum up collaboration in this thread, as an > > > >> opportunity > > > >>>> to learn exciting, potential-filled new technologies. Adding native > > > >>>> compilation could open new channels for library integration. > > > >>>> Learning > > > >> this > > > >>>> is a great resume builder and offers experience in profiling and > > > >> testing > > > >>>> that exceeds our current scope. And it's low risk but takes a > > > >> calculated > > > >>>> chance. This is why it'll remain in the sandbox until the > > > >>>> build-then-demand model is proven. > > > >>>> > > > >>>> But to gain demand - that's when blog posts, demos, and conference > > > chit > > > >>>> chat come in. So I promise to focus most of my efforts there once it > > > >> is > > > >>>> built. > > > >>>> > > > >>>> Here is why I think it'll work: > > > >>>> > > > >>>> NLP has been overshadowed by LLM hype, but as the need for fast > > > >>>> ground-truth validation grows, performance will be key. People > > > unfairly > > > >>>> separate LLMs, search, and NLP when they all deal with language and > > > >> they > > > >>>> need better integration. Since Java isn't natively fast (though it > > > >>>> is > > > >>>> still significantly faster than Python), demonstrating GraalVM's > > > >>>> speed > > > >> and > > > >>>> integration capabilities could reignite interest in the ecosystem, > > > >> though > > > >>>> complex setup remains a risk. > > > >>>> > > > >>>> Finally I won't dismiss that maintaining native binaries is a > > > >>>> headache > > > >> and > > > >>>> I think we might want to avoid it. CVE concerns given rapidly > > > changing > > > >>>> architectures compound this issue. Even if it leaves the sandbox, > > > >> keeping > > > >>>> this as a build-it-yourself setup may be the best approach. This > > > avoids > > > >>>> direct binary distribution given the complexity of the setup > > > >>>> (ARM/AMD, > > > >>>> NVidia vs. Intel, CPU/GPU/NPU, and OS all need to be considered). If > > > >> this > > > >>>> experiment fails to gain traction, I fully support deprecating and > > > >> retiring > > > >>>> it quickly. > > > >>>> > > > >>>> Thanks again for the thorough write-up. > > > >>>> > > > >>>> Best, > > > >>>> Kristian > > > >>>> > > > >>>> > > > >>>> > > > >>>> On Mon, Sep 14, 2026, 2:42 AM Richard Zowalla <[email protected]> > > > wrote: > > > >>>> > > > >>>>> Hi, > > > >>>>> > > > >>>>> I'd frame the scope a bit differently, though. Here are my 2cts from > > > >> my > > > >>>>> work in the EE ecosystem. > > > >>>>> > > > >>>>> OpenNLP is a library. Compiling to native code is something the > > > >> user's > > > >>>>> application does, not something we do IMHO. What OpenNLP can do is > > > >> make the > > > >>>>> jars native-friendly, so that anyone building a native image (with > > > >> the > > > >>>>> GraalVM Maven/Gradle plugins, Quarkus, Spring AOT, Apache Geronimo > > > >> Arthur, > > > >>>>> ...) gets a working build without extra steps. Concretely: > > > >>>>> > > > >>>>> 1. Ship reachability metadata inside our jars > > > >>>>> (META-INF/native-image/org.apache.opennlp/<artifactId>/...), or a > > > >> small > > > >>>>> GraalVM Feature class. > > > >>>>> 2. Fix the parts that break in a native image. We do use reflection, > > > >> and > > > >>>>> it isn't only in edge cases: > > > >>>>> (a) ExtensionLoader / BaseToolFactory / TrainerFactory load classes > > > >> by > > > >>>>> name. The factory class names even come from the model manifest. > > > >>>>> (b) GenericModelReader/Writer and ClassPathModelLoader create > > > >> objects > > > >>>>> through reflective constructors. > > > >>>>> (c) The Snowball stemmers look up methods via > > > >> MethodHandles.findVirtual. > > > >>>>> (d) The CLI's ArgumentParser uses dynamic proxies. > > > >>>>> (e) SimpleClassPathModelFinder reads jdk.internal internals and > > > >> scans > > > >>>>> jars on the classpath. That can't work in a native image because > > > >> there is > > > >>>>> no classpath. > > > >>>>> (f) opennlp-dl pulls in ONNX Runtime through JNI. > > > >>>>> (g) We load a fair number of bundled resources. > > > >>>>> 3. Do some experiments outside of the OpenNLP repos, like a native > > > >> smoke > > > >>>>> test in CI (tokenize, POS tag, NER). > > > >>>>> > > > >>>>> Custom factories or extensions from users will always need > > > >> registering by > > > >>>>> the user, and that's fine. > > > >>>>> > > > >>>>> On binary distributions: the (WIP) gRPC server would be the obvious > > > >>>>> candidate for native. It's a standalone, long-running service, it > > > >> lives in > > > >>>>> the sandbox anyway, and a native container image makes sense there. > > > >>>>> > > > >>>>> A native CLI could be nice for startup time, but I have no idea how > > > >> many > > > >>>>> people actually use the CLI, so I'd wait until someone asks for it. > > > >>>>> > > > >>>>> Before we publish any native binary, we should ask LEGAL. A native > > > >> image > > > >>>>> contains SubstrateVM and JDK code under GPLv2 with the Classpath > > > >> Exception, > > > >>>>> and we'd also need one build per OS/arch as part of releases. As far > > > >> as I > > > >>>>> know, Kafka publishes a native Docker image (KIP-974), so there's at > > > >> least > > > >>>>> a precedent. > > > >>>>> > > > >>>>> Some of the expected benefits need benchmarks: > > > >>>>> > > > >>>>> 1. CE native image has no profile-guided optimization and no G1 GC. > > > >> For > > > >>>>> long-running pipelines, peak throughput may be lower than on > > > >>>>> HotSpot. > > > >>>>> 2. Startup gets faster, but loading large models still takes time. > > > >>>>> 3. GPU / CUDA / OpenVINO / NPU support comes from ONNX Runtime, not > > > >> from > > > >>>>> GraalVM. We already have that on the JVM. > > > >>>>> 4. Rust/C++/Python bindings would mean designing and maintaining a C > > > >> API > > > >>>>> (@CEntryPoint, isolates, memory ownership). That's a separate > > > >> project, not > > > >>>>> a side effect of compiling natively. > > > >>>>> > > > >>>>> In the ASF EE world, we have Apache Geronimo Arthur: it's a thin > > > >> Maven > > > >>>>> layer over native-image. It downloads GraalVM, generates the > > > >> native-image > > > >>>>> config, can build images via Jib, and has an extension SPI > > > >> ("knights") that > > > >>>>> computes reflection/resource config at build time. It's ASF, so the > > > >> people > > > >>>>> behind it are well known in the EE ecosystem. That said, the last > > > >> release > > > >>>>> (1.0.9) was in April 2024 and the project has been quiet since. I > > > >> wouldn't > > > >>>>> make OpenNLP depend on it. Standard metadata in our jars works with > > > >> Arthur > > > >>>>> as-is. If there's actually interest, an "opennlp-knight" would be a > > > >> small > > > >>>>> add-on that registers the same things. For building the gRPC binary, > > > >> I'd go > > > >>>>> with the official GraalVM native-maven-plugin, which is actively > > > >> maintained. > > > >>>>> > > > >>>>> TL;DR: making the library native-friendly is doable, but it's real > > > >> work. > > > >>>>> Shipping native binaries ourselves has no real gain as long as there > > > >> is no > > > >>>>> released gRPC server. For the CLI binary and language bindings, > > > >> let's wait > > > >>>>> for real demand. > > > >>>>> > > > >>>>> Richard > > > >>>>> > > > >>>>> On 2026/09/14 00:50:17 Kristian Rickert wrote: > > > >>>>>> Devs, > > > >>>>>> > > > >>>>>> I would like to propose an initiative for a 3.x release (post- > > > >> 3.0): > > > >>>>>> compiling OpenNLP to native code using GraalVM. > > > >>>>>> > > > >>>>>> Here's the situation: Java is sandwiched between Python's data > > > >> science > > > >>>>>> dominance and Rust's performance. To combat this, I propose > > > >> leveraging > > > >>>>>> GraalVM to compile OpenNLP natively, which offers significant > > > >> advantages. > > > >>>>>> Having used it in production, I have seen it deliver instant > > > >> startup > > > >>>>> times, > > > >>>>>> lower memory usage, and improved latency. > > > >>>>>> > > > >>>>>> Given OpenNLP's minimal dependencies and lack of reflection, it is > > > >> a > > > >>>>> prime > > > >>>>>> candidate for this. Initial tests compiling to native code have > > > >> yielded > > > >>>>> no > > > >>>>>> major issues. > > > >>>>>> > > > >>>>>> Some Pros: > > > >>>>>> > > > >>>>>> - Broader Integration: We can package OpenNLP as a Rust crate > > > >> or C++ > > > >>>>>> library, allowing direct integration into applications, word > > > >>>>> processors, > > > >>>>>> and Python (via Cython). > > > >>>>>> - Cross-Language Native Support: OpenNLP could be used natively > > > >> in > > > >>>>> Rust, > > > >>>>>> C++, Swift, and Python with a much smaller memory footprint. > > > >>>>>> - Performance Gains: By leveraging the pluggable embedding layer > > > >>>>> created > > > >>>>>> for the gRPC service, embedding performance could be at least 2x > > > >>>>> faster > > > >>>>>> (via GPU or static table creation). > > > >>>>>> - Wide Architecture Support: Native support for Apple Silicon, > > > >> Intel > > > >>>>>> NPU, CUDA, OpenVINO, Android, and CPU execution. > > > >>>>>> > > > >>>>>> Questions for the Team: > > > >>>>>> > > > >>>>>> 1. Does anyone know of other Apache projects currently using > > > >> GraalVM > > > >>>>>> compilation? If so, please reach out directly, I'd love to > > > >> connect > > > >>>>> with > > > >>>>>> them. > > > >>>>>> 2. Do we have any connections with Oracle folks? They create > > > >> it, if I > > > >>>>>> run into issues, having them available to help would be > > > >> beneficial. > > > >>>>> (Note: > > > >>>>>> We would use the CE edition) > > > >>>>>> 3. Are there any constraints / issues this can cause? > > > >>>>>> 4. This can be a downstream build, and I'd volunteer to set up > > > >> the > > > >>>>> CICD > > > >>>>>> for it. Anyone up for helping? It can't hurt to understand > > > >> Java > > > >>>>> native > > > >>>>>> compilations. > > > >>>>>> 5. Obviously, I'd set this up in sandbox and it'll be post-gRPC > > > >> (I was > > > >>>>>> planning on natively compiling the gRPC server anyway) > > > >>>>>> > > > >>>>>> If you're interested in helping with this experiment, please let > > > >> me know! > > > >>>>>> > > > >>>>>> Mutant test rungs, > > > >>>>>> Kristian > > > >>>>>> > > > >>>>> > > > >>>> > > > >> > > > > > > Disclaimer > > > > > > The information contained in this communication from the sender is > > > confidential. It is intended solely for use by the recipient and others > > > authorized to receive it. If you are not the recipient, you are hereby > > > notified that any disclosure, copying, distribution or taking action in > > > relation of the contents of this information is strictly prohibited and > > > may > > > be unlawful. > > > > > > This email has been scanned for viruses and malware, and may have been > > > automatically archived by Mimecast, a leader in email security and cyber > > > resilience. Mimecast integrates email defenses with brand protection, > > > security awareness training, web security, compliance and other essential > > > capabilities. Mimecast helps protect large and small organizations from > > > malicious activity, human error and technology failure; and to lead the > > > movement toward building a more resilient world. To find out more, visit > > > our website. > > > > >
