Hi,

I'd frame the scope a bit differently, though. Here are my 2cts from my work in 
the EE ecosystem.

OpenNLP is a library. Compiling to native code is something the user's 
application does, not something we do IMHO. What OpenNLP can do is make the 
jars native-friendly, so that anyone building a native image (with the GraalVM 
Maven/Gradle plugins, Quarkus, Spring AOT, Apache Geronimo Arthur, ...) gets a 
working build without extra steps. Concretely:

1. Ship reachability metadata inside our jars 
(META-INF/native-image/org.apache.opennlp/<artifactId>/...), or a small GraalVM 
Feature class.
2. Fix the parts that break in a native image. We do use reflection, and it 
isn't only in edge cases:
 (a) ExtensionLoader / BaseToolFactory / TrainerFactory load classes by name. 
The factory class names even come from the model manifest.
 (b) GenericModelReader/Writer and ClassPathModelLoader create objects through 
reflective constructors.
 (c) The Snowball stemmers look up methods via MethodHandles.findVirtual.
 (d) The CLI's ArgumentParser uses dynamic proxies.
 (e) SimpleClassPathModelFinder reads jdk.internal internals and scans jars on 
the classpath. That can't work in a native image because there is no classpath.
 (f) opennlp-dl pulls in ONNX Runtime through JNI.
 (g) We load a fair number of bundled resources.
3. Do some experiments outside of the OpenNLP repos, like a native smoke test 
in CI (tokenize, POS tag, NER).

Custom factories or extensions from users will always need registering by the 
user, and that's fine.

On binary distributions: the (WIP) gRPC server would be the obvious candidate 
for native. It's a standalone, long-running service, it lives in the sandbox 
anyway, and a native container image makes sense there.

A native CLI could be nice for startup time, but I have no idea how many people 
actually use the CLI, so I'd wait until someone asks for it.

Before we publish any native binary, we should ask LEGAL. A native image 
contains SubstrateVM and JDK code under GPLv2 with the Classpath Exception, and 
we'd also need one build per OS/arch as part of releases. As far as I know, 
Kafka publishes a native Docker image (KIP-974), so there's at least a 
precedent.

Some of the expected benefits need benchmarks:

1. CE native image has no profile-guided optimization and no G1 GC. For 
long-running pipelines, peak throughput may be lower than on HotSpot.
2. Startup gets faster, but loading large models still takes time.
3. GPU / CUDA / OpenVINO / NPU support comes from ONNX Runtime, not from 
GraalVM. We already have that on the JVM.
4. Rust/C++/Python bindings would mean designing and maintaining a C API 
(@CEntryPoint, isolates, memory ownership). That's a separate project, not a 
side effect of compiling natively.

In the ASF EE world, we have Apache Geronimo Arthur: it's a thin Maven layer 
over native-image. It downloads GraalVM, generates the native-image config, can 
build images via Jib, and has an extension SPI ("knights") that computes 
reflection/resource config at build time. It's ASF, so the people behind it are 
well known in the EE ecosystem. That said, the last release (1.0.9) was in 
April 2024 and the project has been quiet since. I wouldn't make OpenNLP depend 
on it. Standard metadata in our jars works with Arthur as-is. If there's 
actually interest, an "opennlp-knight" would be a small add-on that registers 
the same things. For building the gRPC binary, I'd go with the official GraalVM 
native-maven-plugin, which is actively maintained.

TL;DR: making the library native-friendly is doable, but it's real work. 
Shipping native binaries ourselves has no real gain as long as there is no 
released gRPC server. For the CLI binary and language bindings, let's wait for 
real demand. 

Richard

On 2026/09/14 00:50:17 Kristian Rickert wrote:
> Devs,
> 
> I would like to propose an initiative for a 3.x release (post- 3.0):
> compiling OpenNLP to native code using GraalVM.
> 
> Here's the situation: Java is sandwiched between Python's data science
> dominance and Rust's performance. To combat this, I propose leveraging
> GraalVM to compile OpenNLP natively, which offers significant advantages.
> Having used it in production, I have seen it deliver instant startup times,
> lower memory usage, and improved latency.
> 
> Given OpenNLP's minimal dependencies and lack of reflection, it is a prime
> candidate for this. Initial tests compiling to native code have yielded no
> major issues.
> 
> Some Pros:
> 
>    - Broader Integration: We can package OpenNLP as a Rust crate or C++
>    library, allowing direct integration into applications, word processors,
>    and Python (via Cython).
>    - Cross-Language Native Support: OpenNLP could be used natively in Rust,
>    C++, Swift, and Python with a much smaller memory footprint.
>    - Performance Gains: By leveraging the pluggable embedding layer created
>    for the gRPC service, embedding performance could be at least 2x faster
>    (via GPU or static table creation).
>    - Wide Architecture Support: Native support for Apple Silicon, Intel
>    NPU, CUDA, OpenVINO, Android, and CPU execution.
> 
> Questions for the Team:
> 
>    1. Does anyone know of other Apache projects currently using GraalVM
>    compilation? If so, please reach out directly, I'd love to connect with
>    them.
>    2. Do we have any connections with Oracle folks?  They create it, if I
>    run into issues, having them available to help would be beneficial. (Note:
>    We would use the CE edition)
>    3. Are there any constraints / issues this can cause?
>    4. This can be a downstream build, and I'd volunteer to set up the CICD
>    for it.  Anyone up for helping?  It can't hurt to understand Java native
>    compilations.
>    5. Obviously, I'd set this up in sandbox and it'll be post-gRPC (I was
>    planning on natively compiling the gRPC server anyway)
> 
> If you're interested in helping with this experiment, please let me know!
> 
> Mutant test rungs,
> Kristian
> 

Reply via email to