Devs, I would like to propose an initiative for a 3.x release (post- 3.0): compiling OpenNLP to native code using GraalVM.
Here's the situation: Java is sandwiched between Python's data science dominance and Rust's performance. To combat this, I propose leveraging GraalVM to compile OpenNLP natively, which offers significant advantages. Having used it in production, I have seen it deliver instant startup times, lower memory usage, and improved latency. Given OpenNLP's minimal dependencies and lack of reflection, it is a prime candidate for this. Initial tests compiling to native code have yielded no major issues. Some Pros: - Broader Integration: We can package OpenNLP as a Rust crate or C++ library, allowing direct integration into applications, word processors, and Python (via Cython). - Cross-Language Native Support: OpenNLP could be used natively in Rust, C++, Swift, and Python with a much smaller memory footprint. - Performance Gains: By leveraging the pluggable embedding layer created for the gRPC service, embedding performance could be at least 2x faster (via GPU or static table creation). - Wide Architecture Support: Native support for Apple Silicon, Intel NPU, CUDA, OpenVINO, Android, and CPU execution. Questions for the Team: 1. Does anyone know of other Apache projects currently using GraalVM compilation? If so, please reach out directly, I'd love to connect with them. 2. Do we have any connections with Oracle folks? They create it, if I run into issues, having them available to help would be beneficial. (Note: We would use the CE edition) 3. Are there any constraints / issues this can cause? 4. This can be a downstream build, and I'd volunteer to set up the CICD for it. Anyone up for helping? It can't hurt to understand Java native compilations. 5. Obviously, I'd set this up in sandbox and it'll be post-gRPC (I was planning on natively compiling the gRPC server anyway) If you're interested in helping with this experiment, please let me know! Mutant test rungs, Kristian
