Hi,
I'd frame the scope a bit differently, though. Here are my 2cts from my work in
the EE ecosystem.
OpenNLP is a library. Compiling to native code is something the user's
application does, not something we do IMHO. What OpenNLP can do is make the
jars native-friendly, so that anyone building a native image (with the GraalVM
Maven/Gradle plugins, Quarkus, Spring AOT, Apache Geronimo Arthur, ...) gets a
working build without extra steps. Concretely:
1. Ship reachability metadata inside our jars
(META-INF/native-image/org.apache.opennlp/<artifactId>/...), or a small GraalVM
Feature class.
2. Fix the parts that break in a native image. We do use reflection, and it
isn't only in edge cases:
(a) ExtensionLoader / BaseToolFactory / TrainerFactory load classes by name.
The factory class names even come from the model manifest.
(b) GenericModelReader/Writer and ClassPathModelLoader create objects through
reflective constructors.
(c) The Snowball stemmers look up methods via MethodHandles.findVirtual.
(d) The CLI's ArgumentParser uses dynamic proxies.
(e) SimpleClassPathModelFinder reads jdk.internal internals and scans jars on
the classpath. That can't work in a native image because there is no classpath.
(f) opennlp-dl pulls in ONNX Runtime through JNI.
(g) We load a fair number of bundled resources.
3. Do some experiments outside of the OpenNLP repos, like a native smoke test
in CI (tokenize, POS tag, NER).
Custom factories or extensions from users will always need registering by the
user, and that's fine.
On binary distributions: the (WIP) gRPC server would be the obvious candidate
for native. It's a standalone, long-running service, it lives in the sandbox
anyway, and a native container image makes sense there.
A native CLI could be nice for startup time, but I have no idea how many people
actually use the CLI, so I'd wait until someone asks for it.
Before we publish any native binary, we should ask LEGAL. A native image
contains SubstrateVM and JDK code under GPLv2 with the Classpath Exception, and
we'd also need one build per OS/arch as part of releases. As far as I know,
Kafka publishes a native Docker image (KIP-974), so there's at least a
precedent.
Some of the expected benefits need benchmarks:
1. CE native image has no profile-guided optimization and no G1 GC. For
long-running pipelines, peak throughput may be lower than on HotSpot.
2. Startup gets faster, but loading large models still takes time.
3. GPU / CUDA / OpenVINO / NPU support comes from ONNX Runtime, not from
GraalVM. We already have that on the JVM.
4. Rust/C++/Python bindings would mean designing and maintaining a C API
(@CEntryPoint, isolates, memory ownership). That's a separate project, not a
side effect of compiling natively.
In the ASF EE world, we have Apache Geronimo Arthur: it's a thin Maven layer
over native-image. It downloads GraalVM, generates the native-image config, can
build images via Jib, and has an extension SPI ("knights") that computes
reflection/resource config at build time. It's ASF, so the people behind it are
well known in the EE ecosystem. That said, the last release (1.0.9) was in
April 2024 and the project has been quiet since. I wouldn't make OpenNLP depend
on it. Standard metadata in our jars works with Arthur as-is. If there's
actually interest, an "opennlp-knight" would be a small add-on that registers
the same things. For building the gRPC binary, I'd go with the official GraalVM
native-maven-plugin, which is actively maintained.
TL;DR: making the library native-friendly is doable, but it's real work.
Shipping native binaries ourselves has no real gain as long as there is no
released gRPC server. For the CLI binary and language bindings, let's wait for
real demand.
Richard
On 2026/09/14 00:50:17 Kristian Rickert wrote:
> Devs,
>
> I would like to propose an initiative for a 3.x release (post- 3.0):
> compiling OpenNLP to native code using GraalVM.
>
> Here's the situation: Java is sandwiched between Python's data science
> dominance and Rust's performance. To combat this, I propose leveraging
> GraalVM to compile OpenNLP natively, which offers significant advantages.
> Having used it in production, I have seen it deliver instant startup times,
> lower memory usage, and improved latency.
>
> Given OpenNLP's minimal dependencies and lack of reflection, it is a prime
> candidate for this. Initial tests compiling to native code have yielded no
> major issues.
>
> Some Pros:
>
> - Broader Integration: We can package OpenNLP as a Rust crate or C++
> library, allowing direct integration into applications, word processors,
> and Python (via Cython).
> - Cross-Language Native Support: OpenNLP could be used natively in Rust,
> C++, Swift, and Python with a much smaller memory footprint.
> - Performance Gains: By leveraging the pluggable embedding layer created
> for the gRPC service, embedding performance could be at least 2x faster
> (via GPU or static table creation).
> - Wide Architecture Support: Native support for Apple Silicon, Intel
> NPU, CUDA, OpenVINO, Android, and CPU execution.
>
> Questions for the Team:
>
> 1. Does anyone know of other Apache projects currently using GraalVM
> compilation? If so, please reach out directly, I'd love to connect with
> them.
> 2. Do we have any connections with Oracle folks? They create it, if I
> run into issues, having them available to help would be beneficial. (Note:
> We would use the CE edition)
> 3. Are there any constraints / issues this can cause?
> 4. This can be a downstream build, and I'd volunteer to set up the CICD
> for it. Anyone up for helping? It can't hurt to understand Java native
> compilations.
> 5. Obviously, I'd set this up in sandbox and it'll be post-gRPC (I was
> planning on natively compiling the gRPC server anyway)
>
> If you're interested in helping with this experiment, please let me know!
>
> Mutant test rungs,
> Kristian
>