Hi Eric, Thank you, Eric! This is incredibly helpful.
I hope our speed improvements for v3 prove otherwise. Here is a quick rundown of what we have been working hard on: - Most of the codebase is now thread-safe (showing about a 2.5x improvement). - We are removing all regex from the code for better performance and fewer bugs. - We have a much cleaner interface and full UTF support. - We added several internationalization features and bi-directional normalization for accurate highlighting. - We added a pile of new features to addons, including a PR for a Java trainer and static embeddings (which can produce 100k/s on many systems). It is our mission (and my obsession) to make OpenNLP performant and on par with popular NLP packages in other ecosystems. So it sounds like GraalVM was a great direction to explore. I have a ticket open for it now. Outside of containers, is anyone delivering native binaries for libraries? Rungs, Kristian On Tue, Sep 15, 2026 at 7:28 AM Eric Pugh <[email protected]> wrote: > I’ve been really intrigued with GraalVM…. A couple of observations: > > 1) We are using it in the forthcoming solr-mcp tool to drastically speed > up how fast solr-cmp starts when used as a command line tool. Solr-MCP > leverages Spring and Spring AI, and as you can imagine, there is a lot of > wiring dynamically going on in that. GraalVM changes that by “pre > compiling” (not sure that is the right word) all of that so you get super > fast start up. We ship the GraalVM version in a dockerized version of > solr-mcp so you don’t need to install anything locally. And yeah, it adds > some build complexity on top to get that to work, but then you are done. > > 2) Apache Tika actually runs on GraalVM! > https://docs.quarkiverse.io/quarkus-tika/dev/index.html. Tika is an > example of a project that has a million dependencies, so if it can run on > GraalVM…. > > 3) OpenNLP is definitely bounded by performance…. When we added it to > Solr as a UpdateRequestProcessor, any model beyond the most simplistic made > the indexing pace almost unusuable, which means you need to have a primary > indexing path that doesn’t touch OpenNLP and then a secondary “enrichment” > one that comes back async and does the OpenNLP work… > > > > > On Sep 14, 2026, at 10:18 PM, Kristian Rickert <[email protected]> > wrote: > > > > *Is something that can be achieved via a Maven profile?* > > > > Yes, this is typically done using the native-maven-plugin provided by > > GraalVM. > > > > *Is this build done via the project using OpenNLP, or does OpenNLP itself > > have to be compiled differently?* > > > > OpenNLP itself does not need to be compiled differently. The library is > > still compiled to standard Java bytecode just like normal. Native > > compilation happens at the final application or distribution level. > > > > I'll give you my 2 cents since I've spent the last hour refreshing myself > > on it. > > > > First I want to state my motivation: I want to create a native library > for > > apps calling OpenNLP (gRPC is just one example). I'm convinced what we > are > > building can be integrated into many text-heavy apps in Rust, C/C++, or > > Swift. It could help word processing apps, AI runtime analysis, and RAG > > pipelines everyone is hyping about. Basically, between this and gRPC, > our > > library can reach any app. > > > > The tool that builds the binaries - the GraalVM Native Image - analyzes > the > > application entry point and traces all reachable code across all > > dependencies to produce a single binary executable. This is similar to > what > > you would get with gcc, and it comes with its own set of advantages and > > headaches. > > > > I don't understand the magic of this process BUT it sure sounds cool and > is > > a low level java thing to learn. Richard's video he references goes over > > this entire process in detail, and I plan to watch it (it's a great > > presentation - I'm a fan of his font choice). He explains the compilation > > process and the performance implications without trying to sell you on > the > > technology either way. > > > > Everything Richard said is right, and he is rightfully questioning my > speed > > claims. Native Image gives you near instant startup time and a very low > > initial memory footprint. However, a traditional JVM JIT compiler might > > actually beat it in peak sustained throughput for long running tasks. I > > have not claimed with certainty that it will be faster for raw > > computational throughput, but I'm enthusiastic and fueled by hope that it > > would. We will have to measure the performance and test it extensively. > > > > Most modern Java tries to account for this. Both Quarkus and Micronaut > are > > built to work with native compilation out of the box. It's why I use > > Quarkus in a lot of my projects - a premature optimization I've not > needed > > yet but can save your AWS bills by easy double digits. When you get a > > multi-million dollar bill from Bezos, this suddenly becomes a more > > attractive choice :) > > > > For OpenNLP, almost all the code should work fine. The biggest friction > > point will be the model loading. We need to make sure any models loaded > as > > classpath resources or anything relying on Java serialization are > > explicitly registered in the Native Image configuration files or we can > > refactor the code to load services through an SPI loader. I'm a big fan > of > > SPI loading, so I'd lean toward that. > > > > Last but not least, since this is Oracle's baby, they own GraalVM, and > > there's a community edition. I've only used the CE, because the Oracle > one > > is a commercial product. You can use it for free in the Oracle cloud - > but > > I don't want to be "that guy" in our org to suggest it. > > > > Initially, I'd only want to support compiling it and leave the various > arch > > binaries up to the user. IMHO anyone who uses GraalVM right now doesn't > > look for native binaries, they just build the binaries themselves or > might > > use a native docker container instead (I've never seen a java native > > library in the wild that isn't a container). We'll make a lot of DevOps > > folks happy if we can advertise that we are GaalVM compliant - if they > use > > it. To Richard's point, the other project he referenced hasn't seen much > > action from it. By doing this, we learn a lot, potentially offer > something > > that integrates into a larger ecosystem, and if it fails we remove it > from > > the build. > > > > > > On Mon, Sep 14, 2026, 4:29 PM Jeff Zemerick <[email protected]> > wrote: > > > >> I'm not super familiar with GraalVM or what OpenNLP requires to be > >> compatible with it but I understand the desire to do so. > >> > >> A couple questions: > >> > >> - Is something that can be achieved via a Maven profile? > >> - Is this build done via the project using OpenNLP, or does OpenNLP > >> itself have to be compiled differently? > >> > >> Thanks, > >> Jeff > >> > >> On Mon, Sep 14, 2026 at 8:21 AM Richard Zowalla <[email protected]> > wrote: > >>> > >>> Hi, > >>> > >>> if the deliverable is "build it yourself", the actual product is the > >> reachability metadata in our jars and the reflection fixes, and that > can't > >> live in the sandbox. It has to go into the main repo, or users still > can't > >> build a working image. The sandbox can hold the recipe and the gRPC > binary > >> experiment, but the enabling work is main-repo work and needs to be > scoped > >> as such. > >>> > >>> On the demand argument, note that native image on CE gives startup and > >> memory footprint (I recommend "Scotty I need Warp Speed" from Gerrit > >> Grunwald) not peak throughput (no PGO, no G1). For long-running > validation > >> pipelines a warm JIT may still win. > >>> > >>> Richard > >>> > >>> On 2026/09/14 11:53:20 Kristian Rickert wrote: > >>>> Hi Richard, > >>>> > >>>> Thank you for the detailed response. It answered all of my questions > >> and > >>>> provides a clear path forward with great historical context we should > >> heed. > >>>> > >>>> I agree with your technical assessment, but see a subtle difference > >>>> regarding demand: I believe proactive innovation drives demand. I > >> call it > >>>> the Kevin Kostner effect: "if you build it, they will come". > >>>> > >>>> But I also hope to drum up collaboration in this thread, as an > >> opportunity > >>>> to learn exciting, potential-filled new technologies. Adding native > >>>> compilation could open new channels for library integration. Learning > >> this > >>>> is a great resume builder and offers experience in profiling and > >> testing > >>>> that exceeds our current scope. And it's low risk but takes a > >> calculated > >>>> chance. This is why it'll remain in the sandbox until the > >>>> build-then-demand model is proven. > >>>> > >>>> But to gain demand - that's when blog posts, demos, and conference > chit > >>>> chat come in. So I promise to focus most of my efforts there once it > >> is > >>>> built. > >>>> > >>>> Here is why I think it'll work: > >>>> > >>>> NLP has been overshadowed by LLM hype, but as the need for fast > >>>> ground-truth validation grows, performance will be key. People > unfairly > >>>> separate LLMs, search, and NLP when they all deal with language and > >> they > >>>> need better integration. Since Java isn't natively fast (though it is > >>>> still significantly faster than Python), demonstrating GraalVM's speed > >> and > >>>> integration capabilities could reignite interest in the ecosystem, > >> though > >>>> complex setup remains a risk. > >>>> > >>>> Finally I won't dismiss that maintaining native binaries is a headache > >> and > >>>> I think we might want to avoid it. CVE concerns given rapidly > changing > >>>> architectures compound this issue. Even if it leaves the sandbox, > >> keeping > >>>> this as a build-it-yourself setup may be the best approach. This > avoids > >>>> direct binary distribution given the complexity of the setup (ARM/AMD, > >>>> NVidia vs. Intel, CPU/GPU/NPU, and OS all need to be considered). If > >> this > >>>> experiment fails to gain traction, I fully support deprecating and > >> retiring > >>>> it quickly. > >>>> > >>>> Thanks again for the thorough write-up. > >>>> > >>>> Best, > >>>> Kristian > >>>> > >>>> > >>>> > >>>> On Mon, Sep 14, 2026, 2:42 AM Richard Zowalla <[email protected]> > wrote: > >>>> > >>>>> Hi, > >>>>> > >>>>> I'd frame the scope a bit differently, though. Here are my 2cts from > >> my > >>>>> work in the EE ecosystem. > >>>>> > >>>>> OpenNLP is a library. Compiling to native code is something the > >> user's > >>>>> application does, not something we do IMHO. What OpenNLP can do is > >> make the > >>>>> jars native-friendly, so that anyone building a native image (with > >> the > >>>>> GraalVM Maven/Gradle plugins, Quarkus, Spring AOT, Apache Geronimo > >> Arthur, > >>>>> ...) gets a working build without extra steps. Concretely: > >>>>> > >>>>> 1. Ship reachability metadata inside our jars > >>>>> (META-INF/native-image/org.apache.opennlp/<artifactId>/...), or a > >> small > >>>>> GraalVM Feature class. > >>>>> 2. Fix the parts that break in a native image. We do use reflection, > >> and > >>>>> it isn't only in edge cases: > >>>>> (a) ExtensionLoader / BaseToolFactory / TrainerFactory load classes > >> by > >>>>> name. The factory class names even come from the model manifest. > >>>>> (b) GenericModelReader/Writer and ClassPathModelLoader create > >> objects > >>>>> through reflective constructors. > >>>>> (c) The Snowball stemmers look up methods via > >> MethodHandles.findVirtual. > >>>>> (d) The CLI's ArgumentParser uses dynamic proxies. > >>>>> (e) SimpleClassPathModelFinder reads jdk.internal internals and > >> scans > >>>>> jars on the classpath. That can't work in a native image because > >> there is > >>>>> no classpath. > >>>>> (f) opennlp-dl pulls in ONNX Runtime through JNI. > >>>>> (g) We load a fair number of bundled resources. > >>>>> 3. Do some experiments outside of the OpenNLP repos, like a native > >> smoke > >>>>> test in CI (tokenize, POS tag, NER). > >>>>> > >>>>> Custom factories or extensions from users will always need > >> registering by > >>>>> the user, and that's fine. > >>>>> > >>>>> On binary distributions: the (WIP) gRPC server would be the obvious > >>>>> candidate for native. It's a standalone, long-running service, it > >> lives in > >>>>> the sandbox anyway, and a native container image makes sense there. > >>>>> > >>>>> A native CLI could be nice for startup time, but I have no idea how > >> many > >>>>> people actually use the CLI, so I'd wait until someone asks for it. > >>>>> > >>>>> Before we publish any native binary, we should ask LEGAL. A native > >> image > >>>>> contains SubstrateVM and JDK code under GPLv2 with the Classpath > >> Exception, > >>>>> and we'd also need one build per OS/arch as part of releases. As far > >> as I > >>>>> know, Kafka publishes a native Docker image (KIP-974), so there's at > >> least > >>>>> a precedent. > >>>>> > >>>>> Some of the expected benefits need benchmarks: > >>>>> > >>>>> 1. CE native image has no profile-guided optimization and no G1 GC. > >> For > >>>>> long-running pipelines, peak throughput may be lower than on HotSpot. > >>>>> 2. Startup gets faster, but loading large models still takes time. > >>>>> 3. GPU / CUDA / OpenVINO / NPU support comes from ONNX Runtime, not > >> from > >>>>> GraalVM. We already have that on the JVM. > >>>>> 4. Rust/C++/Python bindings would mean designing and maintaining a C > >> API > >>>>> (@CEntryPoint, isolates, memory ownership). That's a separate > >> project, not > >>>>> a side effect of compiling natively. > >>>>> > >>>>> In the ASF EE world, we have Apache Geronimo Arthur: it's a thin > >> Maven > >>>>> layer over native-image. It downloads GraalVM, generates the > >> native-image > >>>>> config, can build images via Jib, and has an extension SPI > >> ("knights") that > >>>>> computes reflection/resource config at build time. It's ASF, so the > >> people > >>>>> behind it are well known in the EE ecosystem. That said, the last > >> release > >>>>> (1.0.9) was in April 2024 and the project has been quiet since. I > >> wouldn't > >>>>> make OpenNLP depend on it. Standard metadata in our jars works with > >> Arthur > >>>>> as-is. If there's actually interest, an "opennlp-knight" would be a > >> small > >>>>> add-on that registers the same things. For building the gRPC binary, > >> I'd go > >>>>> with the official GraalVM native-maven-plugin, which is actively > >> maintained. > >>>>> > >>>>> TL;DR: making the library native-friendly is doable, but it's real > >> work. > >>>>> Shipping native binaries ourselves has no real gain as long as there > >> is no > >>>>> released gRPC server. For the CLI binary and language bindings, > >> let's wait > >>>>> for real demand. > >>>>> > >>>>> Richard > >>>>> > >>>>> On 2026/09/14 00:50:17 Kristian Rickert wrote: > >>>>>> Devs, > >>>>>> > >>>>>> I would like to propose an initiative for a 3.x release (post- > >> 3.0): > >>>>>> compiling OpenNLP to native code using GraalVM. > >>>>>> > >>>>>> Here's the situation: Java is sandwiched between Python's data > >> science > >>>>>> dominance and Rust's performance. To combat this, I propose > >> leveraging > >>>>>> GraalVM to compile OpenNLP natively, which offers significant > >> advantages. > >>>>>> Having used it in production, I have seen it deliver instant > >> startup > >>>>> times, > >>>>>> lower memory usage, and improved latency. > >>>>>> > >>>>>> Given OpenNLP's minimal dependencies and lack of reflection, it is > >> a > >>>>> prime > >>>>>> candidate for this. Initial tests compiling to native code have > >> yielded > >>>>> no > >>>>>> major issues. > >>>>>> > >>>>>> Some Pros: > >>>>>> > >>>>>> - Broader Integration: We can package OpenNLP as a Rust crate > >> or C++ > >>>>>> library, allowing direct integration into applications, word > >>>>> processors, > >>>>>> and Python (via Cython). > >>>>>> - Cross-Language Native Support: OpenNLP could be used natively > >> in > >>>>> Rust, > >>>>>> C++, Swift, and Python with a much smaller memory footprint. > >>>>>> - Performance Gains: By leveraging the pluggable embedding layer > >>>>> created > >>>>>> for the gRPC service, embedding performance could be at least 2x > >>>>> faster > >>>>>> (via GPU or static table creation). > >>>>>> - Wide Architecture Support: Native support for Apple Silicon, > >> Intel > >>>>>> NPU, CUDA, OpenVINO, Android, and CPU execution. > >>>>>> > >>>>>> Questions for the Team: > >>>>>> > >>>>>> 1. Does anyone know of other Apache projects currently using > >> GraalVM > >>>>>> compilation? If so, please reach out directly, I'd love to > >> connect > >>>>> with > >>>>>> them. > >>>>>> 2. Do we have any connections with Oracle folks? They create > >> it, if I > >>>>>> run into issues, having them available to help would be > >> beneficial. > >>>>> (Note: > >>>>>> We would use the CE edition) > >>>>>> 3. Are there any constraints / issues this can cause? > >>>>>> 4. This can be a downstream build, and I'd volunteer to set up > >> the > >>>>> CICD > >>>>>> for it. Anyone up for helping? It can't hurt to understand > >> Java > >>>>> native > >>>>>> compilations. > >>>>>> 5. Obviously, I'd set this up in sandbox and it'll be post-gRPC > >> (I was > >>>>>> planning on natively compiling the gRPC server anyway) > >>>>>> > >>>>>> If you're interested in helping with this experiment, please let > >> me know! > >>>>>> > >>>>>> Mutant test rungs, > >>>>>> Kristian > >>>>>> > >>>>> > >>>> > >> > > Disclaimer > > The information contained in this communication from the sender is > confidential. It is intended solely for use by the recipient and others > authorized to receive it. If you are not the recipient, you are hereby > notified that any disclosure, copying, distribution or taking action in > relation of the contents of this information is strictly prohibited and may > be unlawful. > > This email has been scanned for viruses and malware, and may have been > automatically archived by Mimecast, a leader in email security and cyber > resilience. Mimecast integrates email defenses with brand protection, > security awareness training, web security, compliance and other essential > capabilities. Mimecast helps protect large and small organizations from > malicious activity, human error and technology failure; and to lead the > movement toward building a more resilient world. To find out more, visit > our website. >
