*Is something that can be achieved via a Maven profile?* Yes, this is typically done using the native-maven-plugin provided by GraalVM.
*Is this build done via the project using OpenNLP, or does OpenNLP itself have to be compiled differently?* OpenNLP itself does not need to be compiled differently. The library is still compiled to standard Java bytecode just like normal. Native compilation happens at the final application or distribution level. I'll give you my 2 cents since I've spent the last hour refreshing myself on it. First I want to state my motivation: I want to create a native library for apps calling OpenNLP (gRPC is just one example). I'm convinced what we are building can be integrated into many text-heavy apps in Rust, C/C++, or Swift. It could help word processing apps, AI runtime analysis, and RAG pipelines everyone is hyping about. Basically, between this and gRPC, our library can reach any app. The tool that builds the binaries - the GraalVM Native Image - analyzes the application entry point and traces all reachable code across all dependencies to produce a single binary executable. This is similar to what you would get with gcc, and it comes with its own set of advantages and headaches. I don't understand the magic of this process BUT it sure sounds cool and is a low level java thing to learn. Richard's video he references goes over this entire process in detail, and I plan to watch it (it's a great presentation - I'm a fan of his font choice). He explains the compilation process and the performance implications without trying to sell you on the technology either way. Everything Richard said is right, and he is rightfully questioning my speed claims. Native Image gives you near instant startup time and a very low initial memory footprint. However, a traditional JVM JIT compiler might actually beat it in peak sustained throughput for long running tasks. I have not claimed with certainty that it will be faster for raw computational throughput, but I'm enthusiastic and fueled by hope that it would. We will have to measure the performance and test it extensively. Most modern Java tries to account for this. Both Quarkus and Micronaut are built to work with native compilation out of the box. It's why I use Quarkus in a lot of my projects - a premature optimization I've not needed yet but can save your AWS bills by easy double digits. When you get a multi-million dollar bill from Bezos, this suddenly becomes a more attractive choice :) For OpenNLP, almost all the code should work fine. The biggest friction point will be the model loading. We need to make sure any models loaded as classpath resources or anything relying on Java serialization are explicitly registered in the Native Image configuration files or we can refactor the code to load services through an SPI loader. I'm a big fan of SPI loading, so I'd lean toward that. Last but not least, since this is Oracle's baby, they own GraalVM, and there's a community edition. I've only used the CE, because the Oracle one is a commercial product. You can use it for free in the Oracle cloud - but I don't want to be "that guy" in our org to suggest it. Initially, I'd only want to support compiling it and leave the various arch binaries up to the user. IMHO anyone who uses GraalVM right now doesn't look for native binaries, they just build the binaries themselves or might use a native docker container instead (I've never seen a java native library in the wild that isn't a container). We'll make a lot of DevOps folks happy if we can advertise that we are GaalVM compliant - if they use it. To Richard's point, the other project he referenced hasn't seen much action from it. By doing this, we learn a lot, potentially offer something that integrates into a larger ecosystem, and if it fails we remove it from the build. On Mon, Sep 14, 2026, 4:29 PM Jeff Zemerick <[email protected]> wrote: > I'm not super familiar with GraalVM or what OpenNLP requires to be > compatible with it but I understand the desire to do so. > > A couple questions: > > - Is something that can be achieved via a Maven profile? > - Is this build done via the project using OpenNLP, or does OpenNLP > itself have to be compiled differently? > > Thanks, > Jeff > > On Mon, Sep 14, 2026 at 8:21 AM Richard Zowalla <[email protected]> wrote: > > > > Hi, > > > > if the deliverable is "build it yourself", the actual product is the > reachability metadata in our jars and the reflection fixes, and that can't > live in the sandbox. It has to go into the main repo, or users still can't > build a working image. The sandbox can hold the recipe and the gRPC binary > experiment, but the enabling work is main-repo work and needs to be scoped > as such. > > > > On the demand argument, note that native image on CE gives startup and > memory footprint (I recommend "Scotty I need Warp Speed" from Gerrit > Grunwald) not peak throughput (no PGO, no G1). For long-running validation > pipelines a warm JIT may still win. > > > > Richard > > > > On 2026/09/14 11:53:20 Kristian Rickert wrote: > > > Hi Richard, > > > > > > Thank you for the detailed response. It answered all of my questions > and > > > provides a clear path forward with great historical context we should > heed. > > > > > > I agree with your technical assessment, but see a subtle difference > > > regarding demand: I believe proactive innovation drives demand. I > call it > > > the Kevin Kostner effect: "if you build it, they will come". > > > > > > But I also hope to drum up collaboration in this thread, as an > opportunity > > > to learn exciting, potential-filled new technologies. Adding native > > > compilation could open new channels for library integration. Learning > this > > > is a great resume builder and offers experience in profiling and > testing > > > that exceeds our current scope. And it's low risk but takes a > calculated > > > chance. This is why it'll remain in the sandbox until the > > > build-then-demand model is proven. > > > > > > But to gain demand - that's when blog posts, demos, and conference chit > > > chat come in. So I promise to focus most of my efforts there once it > is > > > built. > > > > > > Here is why I think it'll work: > > > > > > NLP has been overshadowed by LLM hype, but as the need for fast > > > ground-truth validation grows, performance will be key. People unfairly > > > separate LLMs, search, and NLP when they all deal with language and > they > > > need better integration. Since Java isn't natively fast (though it is > > > still significantly faster than Python), demonstrating GraalVM's speed > and > > > integration capabilities could reignite interest in the ecosystem, > though > > > complex setup remains a risk. > > > > > > Finally I won't dismiss that maintaining native binaries is a headache > and > > > I think we might want to avoid it. CVE concerns given rapidly changing > > > architectures compound this issue. Even if it leaves the sandbox, > keeping > > > this as a build-it-yourself setup may be the best approach. This avoids > > > direct binary distribution given the complexity of the setup (ARM/AMD, > > > NVidia vs. Intel, CPU/GPU/NPU, and OS all need to be considered). If > this > > > experiment fails to gain traction, I fully support deprecating and > retiring > > > it quickly. > > > > > > Thanks again for the thorough write-up. > > > > > > Best, > > > Kristian > > > > > > > > > > > > On Mon, Sep 14, 2026, 2:42 AM Richard Zowalla <[email protected]> wrote: > > > > > > > Hi, > > > > > > > > I'd frame the scope a bit differently, though. Here are my 2cts from > my > > > > work in the EE ecosystem. > > > > > > > > OpenNLP is a library. Compiling to native code is something the > user's > > > > application does, not something we do IMHO. What OpenNLP can do is > make the > > > > jars native-friendly, so that anyone building a native image (with > the > > > > GraalVM Maven/Gradle plugins, Quarkus, Spring AOT, Apache Geronimo > Arthur, > > > > ...) gets a working build without extra steps. Concretely: > > > > > > > > 1. Ship reachability metadata inside our jars > > > > (META-INF/native-image/org.apache.opennlp/<artifactId>/...), or a > small > > > > GraalVM Feature class. > > > > 2. Fix the parts that break in a native image. We do use reflection, > and > > > > it isn't only in edge cases: > > > > (a) ExtensionLoader / BaseToolFactory / TrainerFactory load classes > by > > > > name. The factory class names even come from the model manifest. > > > > (b) GenericModelReader/Writer and ClassPathModelLoader create > objects > > > > through reflective constructors. > > > > (c) The Snowball stemmers look up methods via > MethodHandles.findVirtual. > > > > (d) The CLI's ArgumentParser uses dynamic proxies. > > > > (e) SimpleClassPathModelFinder reads jdk.internal internals and > scans > > > > jars on the classpath. That can't work in a native image because > there is > > > > no classpath. > > > > (f) opennlp-dl pulls in ONNX Runtime through JNI. > > > > (g) We load a fair number of bundled resources. > > > > 3. Do some experiments outside of the OpenNLP repos, like a native > smoke > > > > test in CI (tokenize, POS tag, NER). > > > > > > > > Custom factories or extensions from users will always need > registering by > > > > the user, and that's fine. > > > > > > > > On binary distributions: the (WIP) gRPC server would be the obvious > > > > candidate for native. It's a standalone, long-running service, it > lives in > > > > the sandbox anyway, and a native container image makes sense there. > > > > > > > > A native CLI could be nice for startup time, but I have no idea how > many > > > > people actually use the CLI, so I'd wait until someone asks for it. > > > > > > > > Before we publish any native binary, we should ask LEGAL. A native > image > > > > contains SubstrateVM and JDK code under GPLv2 with the Classpath > Exception, > > > > and we'd also need one build per OS/arch as part of releases. As far > as I > > > > know, Kafka publishes a native Docker image (KIP-974), so there's at > least > > > > a precedent. > > > > > > > > Some of the expected benefits need benchmarks: > > > > > > > > 1. CE native image has no profile-guided optimization and no G1 GC. > For > > > > long-running pipelines, peak throughput may be lower than on HotSpot. > > > > 2. Startup gets faster, but loading large models still takes time. > > > > 3. GPU / CUDA / OpenVINO / NPU support comes from ONNX Runtime, not > from > > > > GraalVM. We already have that on the JVM. > > > > 4. Rust/C++/Python bindings would mean designing and maintaining a C > API > > > > (@CEntryPoint, isolates, memory ownership). That's a separate > project, not > > > > a side effect of compiling natively. > > > > > > > > In the ASF EE world, we have Apache Geronimo Arthur: it's a thin > Maven > > > > layer over native-image. It downloads GraalVM, generates the > native-image > > > > config, can build images via Jib, and has an extension SPI > ("knights") that > > > > computes reflection/resource config at build time. It's ASF, so the > people > > > > behind it are well known in the EE ecosystem. That said, the last > release > > > > (1.0.9) was in April 2024 and the project has been quiet since. I > wouldn't > > > > make OpenNLP depend on it. Standard metadata in our jars works with > Arthur > > > > as-is. If there's actually interest, an "opennlp-knight" would be a > small > > > > add-on that registers the same things. For building the gRPC binary, > I'd go > > > > with the official GraalVM native-maven-plugin, which is actively > maintained. > > > > > > > > TL;DR: making the library native-friendly is doable, but it's real > work. > > > > Shipping native binaries ourselves has no real gain as long as there > is no > > > > released gRPC server. For the CLI binary and language bindings, > let's wait > > > > for real demand. > > > > > > > > Richard > > > > > > > > On 2026/09/14 00:50:17 Kristian Rickert wrote: > > > > > Devs, > > > > > > > > > > I would like to propose an initiative for a 3.x release (post- > 3.0): > > > > > compiling OpenNLP to native code using GraalVM. > > > > > > > > > > Here's the situation: Java is sandwiched between Python's data > science > > > > > dominance and Rust's performance. To combat this, I propose > leveraging > > > > > GraalVM to compile OpenNLP natively, which offers significant > advantages. > > > > > Having used it in production, I have seen it deliver instant > startup > > > > times, > > > > > lower memory usage, and improved latency. > > > > > > > > > > Given OpenNLP's minimal dependencies and lack of reflection, it is > a > > > > prime > > > > > candidate for this. Initial tests compiling to native code have > yielded > > > > no > > > > > major issues. > > > > > > > > > > Some Pros: > > > > > > > > > > - Broader Integration: We can package OpenNLP as a Rust crate > or C++ > > > > > library, allowing direct integration into applications, word > > > > processors, > > > > > and Python (via Cython). > > > > > - Cross-Language Native Support: OpenNLP could be used natively > in > > > > Rust, > > > > > C++, Swift, and Python with a much smaller memory footprint. > > > > > - Performance Gains: By leveraging the pluggable embedding layer > > > > created > > > > > for the gRPC service, embedding performance could be at least 2x > > > > faster > > > > > (via GPU or static table creation). > > > > > - Wide Architecture Support: Native support for Apple Silicon, > Intel > > > > > NPU, CUDA, OpenVINO, Android, and CPU execution. > > > > > > > > > > Questions for the Team: > > > > > > > > > > 1. Does anyone know of other Apache projects currently using > GraalVM > > > > > compilation? If so, please reach out directly, I'd love to > connect > > > > with > > > > > them. > > > > > 2. Do we have any connections with Oracle folks? They create > it, if I > > > > > run into issues, having them available to help would be > beneficial. > > > > (Note: > > > > > We would use the CE edition) > > > > > 3. Are there any constraints / issues this can cause? > > > > > 4. This can be a downstream build, and I'd volunteer to set up > the > > > > CICD > > > > > for it. Anyone up for helping? It can't hurt to understand > Java > > > > native > > > > > compilations. > > > > > 5. Obviously, I'd set this up in sandbox and it'll be post-gRPC > (I was > > > > > planning on natively compiling the gRPC server anyway) > > > > > > > > > > If you're interested in helping with this experiment, please let > me know! > > > > > > > > > > Mutant test rungs, > > > > > Kristian > > > > > > > > > > > > >
