I'm not super familiar with GraalVM or what OpenNLP requires to be compatible with it but I understand the desire to do so.
A couple questions: - Is something that can be achieved via a Maven profile? - Is this build done via the project using OpenNLP, or does OpenNLP itself have to be compiled differently? Thanks, Jeff On Mon, Sep 14, 2026 at 8:21 AM Richard Zowalla <[email protected]> wrote: > > Hi, > > if the deliverable is "build it yourself", the actual product is the > reachability metadata in our jars and the reflection fixes, and that can't > live in the sandbox. It has to go into the main repo, or users still can't > build a working image. The sandbox can hold the recipe and the gRPC binary > experiment, but the enabling work is main-repo work and needs to be scoped as > such. > > On the demand argument, note that native image on CE gives startup and memory > footprint (I recommend "Scotty I need Warp Speed" from Gerrit Grunwald) not > peak throughput (no PGO, no G1). For long-running validation pipelines a warm > JIT may still win. > > Richard > > On 2026/09/14 11:53:20 Kristian Rickert wrote: > > Hi Richard, > > > > Thank you for the detailed response. It answered all of my questions and > > provides a clear path forward with great historical context we should heed. > > > > I agree with your technical assessment, but see a subtle difference > > regarding demand: I believe proactive innovation drives demand. I call it > > the Kevin Kostner effect: "if you build it, they will come". > > > > But I also hope to drum up collaboration in this thread, as an opportunity > > to learn exciting, potential-filled new technologies. Adding native > > compilation could open new channels for library integration. Learning this > > is a great resume builder and offers experience in profiling and testing > > that exceeds our current scope. And it's low risk but takes a calculated > > chance. This is why it'll remain in the sandbox until the > > build-then-demand model is proven. > > > > But to gain demand - that's when blog posts, demos, and conference chit > > chat come in. So I promise to focus most of my efforts there once it is > > built. > > > > Here is why I think it'll work: > > > > NLP has been overshadowed by LLM hype, but as the need for fast > > ground-truth validation grows, performance will be key. People unfairly > > separate LLMs, search, and NLP when they all deal with language and they > > need better integration. Since Java isn't natively fast (though it is > > still significantly faster than Python), demonstrating GraalVM's speed and > > integration capabilities could reignite interest in the ecosystem, though > > complex setup remains a risk. > > > > Finally I won't dismiss that maintaining native binaries is a headache and > > I think we might want to avoid it. CVE concerns given rapidly changing > > architectures compound this issue. Even if it leaves the sandbox, keeping > > this as a build-it-yourself setup may be the best approach. This avoids > > direct binary distribution given the complexity of the setup (ARM/AMD, > > NVidia vs. Intel, CPU/GPU/NPU, and OS all need to be considered). If this > > experiment fails to gain traction, I fully support deprecating and retiring > > it quickly. > > > > Thanks again for the thorough write-up. > > > > Best, > > Kristian > > > > > > > > On Mon, Sep 14, 2026, 2:42 AM Richard Zowalla <[email protected]> wrote: > > > > > Hi, > > > > > > I'd frame the scope a bit differently, though. Here are my 2cts from my > > > work in the EE ecosystem. > > > > > > OpenNLP is a library. Compiling to native code is something the user's > > > application does, not something we do IMHO. What OpenNLP can do is make > > > the > > > jars native-friendly, so that anyone building a native image (with the > > > GraalVM Maven/Gradle plugins, Quarkus, Spring AOT, Apache Geronimo Arthur, > > > ...) gets a working build without extra steps. Concretely: > > > > > > 1. Ship reachability metadata inside our jars > > > (META-INF/native-image/org.apache.opennlp/<artifactId>/...), or a small > > > GraalVM Feature class. > > > 2. Fix the parts that break in a native image. We do use reflection, and > > > it isn't only in edge cases: > > > (a) ExtensionLoader / BaseToolFactory / TrainerFactory load classes by > > > name. The factory class names even come from the model manifest. > > > (b) GenericModelReader/Writer and ClassPathModelLoader create objects > > > through reflective constructors. > > > (c) The Snowball stemmers look up methods via MethodHandles.findVirtual. > > > (d) The CLI's ArgumentParser uses dynamic proxies. > > > (e) SimpleClassPathModelFinder reads jdk.internal internals and scans > > > jars on the classpath. That can't work in a native image because there is > > > no classpath. > > > (f) opennlp-dl pulls in ONNX Runtime through JNI. > > > (g) We load a fair number of bundled resources. > > > 3. Do some experiments outside of the OpenNLP repos, like a native smoke > > > test in CI (tokenize, POS tag, NER). > > > > > > Custom factories or extensions from users will always need registering by > > > the user, and that's fine. > > > > > > On binary distributions: the (WIP) gRPC server would be the obvious > > > candidate for native. It's a standalone, long-running service, it lives in > > > the sandbox anyway, and a native container image makes sense there. > > > > > > A native CLI could be nice for startup time, but I have no idea how many > > > people actually use the CLI, so I'd wait until someone asks for it. > > > > > > Before we publish any native binary, we should ask LEGAL. A native image > > > contains SubstrateVM and JDK code under GPLv2 with the Classpath > > > Exception, > > > and we'd also need one build per OS/arch as part of releases. As far as I > > > know, Kafka publishes a native Docker image (KIP-974), so there's at least > > > a precedent. > > > > > > Some of the expected benefits need benchmarks: > > > > > > 1. CE native image has no profile-guided optimization and no G1 GC. For > > > long-running pipelines, peak throughput may be lower than on HotSpot. > > > 2. Startup gets faster, but loading large models still takes time. > > > 3. GPU / CUDA / OpenVINO / NPU support comes from ONNX Runtime, not from > > > GraalVM. We already have that on the JVM. > > > 4. Rust/C++/Python bindings would mean designing and maintaining a C API > > > (@CEntryPoint, isolates, memory ownership). That's a separate project, not > > > a side effect of compiling natively. > > > > > > In the ASF EE world, we have Apache Geronimo Arthur: it's a thin Maven > > > layer over native-image. It downloads GraalVM, generates the native-image > > > config, can build images via Jib, and has an extension SPI ("knights") > > > that > > > computes reflection/resource config at build time. It's ASF, so the people > > > behind it are well known in the EE ecosystem. That said, the last release > > > (1.0.9) was in April 2024 and the project has been quiet since. I wouldn't > > > make OpenNLP depend on it. Standard metadata in our jars works with Arthur > > > as-is. If there's actually interest, an "opennlp-knight" would be a small > > > add-on that registers the same things. For building the gRPC binary, I'd > > > go > > > with the official GraalVM native-maven-plugin, which is actively > > > maintained. > > > > > > TL;DR: making the library native-friendly is doable, but it's real work. > > > Shipping native binaries ourselves has no real gain as long as there is no > > > released gRPC server. For the CLI binary and language bindings, let's wait > > > for real demand. > > > > > > Richard > > > > > > On 2026/09/14 00:50:17 Kristian Rickert wrote: > > > > Devs, > > > > > > > > I would like to propose an initiative for a 3.x release (post- 3.0): > > > > compiling OpenNLP to native code using GraalVM. > > > > > > > > Here's the situation: Java is sandwiched between Python's data science > > > > dominance and Rust's performance. To combat this, I propose leveraging > > > > GraalVM to compile OpenNLP natively, which offers significant > > > > advantages. > > > > Having used it in production, I have seen it deliver instant startup > > > times, > > > > lower memory usage, and improved latency. > > > > > > > > Given OpenNLP's minimal dependencies and lack of reflection, it is a > > > prime > > > > candidate for this. Initial tests compiling to native code have yielded > > > no > > > > major issues. > > > > > > > > Some Pros: > > > > > > > > - Broader Integration: We can package OpenNLP as a Rust crate or C++ > > > > library, allowing direct integration into applications, word > > > processors, > > > > and Python (via Cython). > > > > - Cross-Language Native Support: OpenNLP could be used natively in > > > Rust, > > > > C++, Swift, and Python with a much smaller memory footprint. > > > > - Performance Gains: By leveraging the pluggable embedding layer > > > created > > > > for the gRPC service, embedding performance could be at least 2x > > > faster > > > > (via GPU or static table creation). > > > > - Wide Architecture Support: Native support for Apple Silicon, Intel > > > > NPU, CUDA, OpenVINO, Android, and CPU execution. > > > > > > > > Questions for the Team: > > > > > > > > 1. Does anyone know of other Apache projects currently using GraalVM > > > > compilation? If so, please reach out directly, I'd love to connect > > > with > > > > them. > > > > 2. Do we have any connections with Oracle folks? They create it, if > > > > I > > > > run into issues, having them available to help would be beneficial. > > > (Note: > > > > We would use the CE edition) > > > > 3. Are there any constraints / issues this can cause? > > > > 4. This can be a downstream build, and I'd volunteer to set up the > > > CICD > > > > for it. Anyone up for helping? It can't hurt to understand Java > > > native > > > > compilations. > > > > 5. Obviously, I'd set this up in sandbox and it'll be post-gRPC (I > > > > was > > > > planning on natively compiling the gRPC server anyway) > > > > > > > > If you're interested in helping with this experiment, please let me > > > > know! > > > > > > > > Mutant test rungs, > > > > Kristian > > > > > > > > >
