Hi,

if the deliverable is "build it yourself", the actual product is the 
reachability metadata in our jars and the reflection fixes, and that can't live 
in the sandbox. It has to go into the main repo, or users still can't build a 
working image. The sandbox can hold the recipe and the gRPC binary experiment, 
but the enabling work is main-repo work and needs to be scoped as such.

On the demand argument, note that native image on CE gives startup and memory 
footprint (I recommend "Scotty I need Warp Speed" from Gerrit Grunwald) not 
peak throughput (no PGO, no G1). For long-running validation pipelines a warm 
JIT may still win.

Richard

On 2026/09/14 11:53:20 Kristian Rickert wrote:
> Hi Richard,
> 
> Thank you for the detailed response. It answered all of my questions and
> provides a clear path forward with great historical context we should heed.
> 
> I agree with your technical assessment, but see a subtle difference
> regarding demand: I believe proactive innovation drives demand.  I call it
> the Kevin Kostner effect: "if you build it, they will come".
> 
> But I also hope to drum up collaboration in this thread, as an opportunity
> to learn exciting, potential-filled new technologies. Adding native
> compilation could open new channels for library integration.  Learning this
> is a great resume builder and offers experience in profiling and testing
> that exceeds our current scope.  And it's low risk but takes a calculated
> chance.  This is why it'll remain in the sandbox until the
> build-then-demand model is proven.
> 
> But to gain demand - that's when blog posts, demos, and conference chit
> chat come in.  So I promise to focus most of my efforts there once it is
> built.
> 
> Here is why I think it'll work:
> 
> NLP has been overshadowed by LLM hype, but as the need for fast
> ground-truth validation grows, performance will be key. People unfairly
> separate LLMs, search, and NLP when they all deal with language and they
> need better integration.  Since Java isn't natively fast (though it is
> still significantly faster than Python), demonstrating GraalVM's speed and
> integration capabilities could reignite interest in the ecosystem, though
> complex setup remains a risk.
> 
> Finally I won't dismiss that maintaining native binaries is a headache and
> I think we might want to avoid it.  CVE concerns given rapidly changing
> architectures compound this issue. Even if it leaves the sandbox, keeping
> this as a build-it-yourself setup may be the best approach. This avoids
> direct binary distribution given the complexity of the setup (ARM/AMD,
> NVidia vs. Intel, CPU/GPU/NPU, and OS all need to be considered). If this
> experiment fails to gain traction, I fully support deprecating and retiring
> it quickly.
> 
> Thanks again for the thorough write-up.
> 
> Best,
> Kristian
> 
> 
> 
> On Mon, Sep 14, 2026, 2:42 AM Richard Zowalla <[email protected]> wrote:
> 
> > Hi,
> >
> > I'd frame the scope a bit differently, though. Here are my 2cts from my
> > work in the EE ecosystem.
> >
> > OpenNLP is a library. Compiling to native code is something the user's
> > application does, not something we do IMHO. What OpenNLP can do is make the
> > jars native-friendly, so that anyone building a native image (with the
> > GraalVM Maven/Gradle plugins, Quarkus, Spring AOT, Apache Geronimo Arthur,
> > ...) gets a working build without extra steps. Concretely:
> >
> > 1. Ship reachability metadata inside our jars
> > (META-INF/native-image/org.apache.opennlp/<artifactId>/...), or a small
> > GraalVM Feature class.
> > 2. Fix the parts that break in a native image. We do use reflection, and
> > it isn't only in edge cases:
> >  (a) ExtensionLoader / BaseToolFactory / TrainerFactory load classes by
> > name. The factory class names even come from the model manifest.
> >  (b) GenericModelReader/Writer and ClassPathModelLoader create objects
> > through reflective constructors.
> >  (c) The Snowball stemmers look up methods via MethodHandles.findVirtual.
> >  (d) The CLI's ArgumentParser uses dynamic proxies.
> >  (e) SimpleClassPathModelFinder reads jdk.internal internals and scans
> > jars on the classpath. That can't work in a native image because there is
> > no classpath.
> >  (f) opennlp-dl pulls in ONNX Runtime through JNI.
> >  (g) We load a fair number of bundled resources.
> > 3. Do some experiments outside of the OpenNLP repos, like a native smoke
> > test in CI (tokenize, POS tag, NER).
> >
> > Custom factories or extensions from users will always need registering by
> > the user, and that's fine.
> >
> > On binary distributions: the (WIP) gRPC server would be the obvious
> > candidate for native. It's a standalone, long-running service, it lives in
> > the sandbox anyway, and a native container image makes sense there.
> >
> > A native CLI could be nice for startup time, but I have no idea how many
> > people actually use the CLI, so I'd wait until someone asks for it.
> >
> > Before we publish any native binary, we should ask LEGAL. A native image
> > contains SubstrateVM and JDK code under GPLv2 with the Classpath Exception,
> > and we'd also need one build per OS/arch as part of releases. As far as I
> > know, Kafka publishes a native Docker image (KIP-974), so there's at least
> > a precedent.
> >
> > Some of the expected benefits need benchmarks:
> >
> > 1. CE native image has no profile-guided optimization and no G1 GC. For
> > long-running pipelines, peak throughput may be lower than on HotSpot.
> > 2. Startup gets faster, but loading large models still takes time.
> > 3. GPU / CUDA / OpenVINO / NPU support comes from ONNX Runtime, not from
> > GraalVM. We already have that on the JVM.
> > 4. Rust/C++/Python bindings would mean designing and maintaining a C API
> > (@CEntryPoint, isolates, memory ownership). That's a separate project, not
> > a side effect of compiling natively.
> >
> > In the ASF EE world, we have Apache Geronimo Arthur: it's a thin Maven
> > layer over native-image. It downloads GraalVM, generates the native-image
> > config, can build images via Jib, and has an extension SPI ("knights") that
> > computes reflection/resource config at build time. It's ASF, so the people
> > behind it are well known in the EE ecosystem. That said, the last release
> > (1.0.9) was in April 2024 and the project has been quiet since. I wouldn't
> > make OpenNLP depend on it. Standard metadata in our jars works with Arthur
> > as-is. If there's actually interest, an "opennlp-knight" would be a small
> > add-on that registers the same things. For building the gRPC binary, I'd go
> > with the official GraalVM native-maven-plugin, which is actively maintained.
> >
> > TL;DR: making the library native-friendly is doable, but it's real work.
> > Shipping native binaries ourselves has no real gain as long as there is no
> > released gRPC server. For the CLI binary and language bindings, let's wait
> > for real demand.
> >
> > Richard
> >
> > On 2026/09/14 00:50:17 Kristian Rickert wrote:
> > > Devs,
> > >
> > > I would like to propose an initiative for a 3.x release (post- 3.0):
> > > compiling OpenNLP to native code using GraalVM.
> > >
> > > Here's the situation: Java is sandwiched between Python's data science
> > > dominance and Rust's performance. To combat this, I propose leveraging
> > > GraalVM to compile OpenNLP natively, which offers significant advantages.
> > > Having used it in production, I have seen it deliver instant startup
> > times,
> > > lower memory usage, and improved latency.
> > >
> > > Given OpenNLP's minimal dependencies and lack of reflection, it is a
> > prime
> > > candidate for this. Initial tests compiling to native code have yielded
> > no
> > > major issues.
> > >
> > > Some Pros:
> > >
> > >    - Broader Integration: We can package OpenNLP as a Rust crate or C++
> > >    library, allowing direct integration into applications, word
> > processors,
> > >    and Python (via Cython).
> > >    - Cross-Language Native Support: OpenNLP could be used natively in
> > Rust,
> > >    C++, Swift, and Python with a much smaller memory footprint.
> > >    - Performance Gains: By leveraging the pluggable embedding layer
> > created
> > >    for the gRPC service, embedding performance could be at least 2x
> > faster
> > >    (via GPU or static table creation).
> > >    - Wide Architecture Support: Native support for Apple Silicon, Intel
> > >    NPU, CUDA, OpenVINO, Android, and CPU execution.
> > >
> > > Questions for the Team:
> > >
> > >    1. Does anyone know of other Apache projects currently using GraalVM
> > >    compilation? If so, please reach out directly, I'd love to connect
> > with
> > >    them.
> > >    2. Do we have any connections with Oracle folks?  They create it, if I
> > >    run into issues, having them available to help would be beneficial.
> > (Note:
> > >    We would use the CE edition)
> > >    3. Are there any constraints / issues this can cause?
> > >    4. This can be a downstream build, and I'd volunteer to set up the
> > CICD
> > >    for it.  Anyone up for helping?  It can't hurt to understand Java
> > native
> > >    compilations.
> > >    5. Obviously, I'd set this up in sandbox and it'll be post-gRPC (I was
> > >    planning on natively compiling the gRPC server anyway)
> > >
> > > If you're interested in helping with this experiment, please let me know!
> > >
> > > Mutant test rungs,
> > > Kristian
> > >
> >
> 

Reply via email to