Hi team, Thank you for weighing in. I was feeling overwhelmed thinking about it and am relieved to hear we will focus on getting version 3.0 released.
There is no rush when I post these issues. All new software releases come with bugs, but the release is still in great shape. Before posting any issues, I will take the time to validate them with solid examples. Version 3.0-SNAPSHOT is in a good place. Assuming everyone agrees, my work will focus on chipping away at fixing bugs, anti-regex, creating backports of bugs, and grpc at a steady pace. This will give me a solid, challenging queue focused on stability and quality, while helping me become more familiar with other parts of the codebase. I'll continue this effort until the queue is fully managable. I plan to limit any new PRs to 0–2 in the near future so I can focus on reviewing PRs from others. I will also continue working on gRPC; so feel free to check the sandbox or ask me directly if you're interested in the current progress / would like to help. I'll regularly post updates there.. Initial tests for the anti-regex work are showing a conservative 2x+ speedup in many areas. I am truly excited about how much work has been done and am looking forward to the release. I have learned a lot from everyone here. Best regards, Kristian P.S. I would prefer not to post the very first issue, so please feel free to open one! On Sun, Sep 20, 2026, 11:52 AM Richard Zowalla <[email protected]> wrote: > Ok, I think I am fine if we postpone things to 3.1.x where they need more > time. Makes sense to re-schedule the related Jiras to 3.1.x then, so the > 3.0.0 scope stays predictable and the regex work can go step by step > instead of as a rush. > > +1 on staying on the regex path rather than opening 16 tickets at once. > Whack-a-mole is a real risk, and the validation work is worth more than the > ticket count. S > ame for the correctness backports to 2.x - those are the ones that matter, > and they can flow at their own pace once all is validated. > > Side note on the ticket migration: we now have GitHub issues enabled, and > Jira is going read-only for the public but stays writable for PMC and > committers. > So there's no need to block on migration, we can leave the existing issues > where they are for now (at least for 3.0.0), and once 3.x is out we can > migrate the remaining open ones to GH. > > Re: the queue: agreed. Anything green with no outstanding comments is fair > game for review now. > > If we all agree, that we don’t need the whole regex removal stuff in 3.x., > we are fine. > > Gruß > R > > > Am 20.09.2026 um 14:12 schrieb Kristian Rickert <[email protected]>: > > > > The regex sweep uncovered several bugs. A thorough regex pass would fix > > them and most fixes should be backported to 2.x - especially the > > correctness ones. This doesn't include the embedding issues recently > > filed, which would also need backporting. > > > > It doesn't change the merge queue, it just means a bigger queue to > process > > everything. I won't post anything until I complete the validation work. > > I'll stick to the current queue, but I'd like the team's thoughts on how > to > > triage this. > > > > Posting 16 bugs might cause a whack-a-mole effect and will take too much > > time away from the first sweep. I'll post to our issues after ticket > > migration happens. > > > > I think staying methodically on the regex path is the way to go. > > > > There are a few tickets that are mergeable; anything green should OK to > > review if there's no outstanding comment. Drafts are not ready for > review > > (by definition), but comments are always welcome. > > > > > > > > On Fri, Sep 11, 2026 at 8:08 AM Kristian Rickert <[email protected]> > wrote: > > > >> Martin, I made them an epic with linked issues. Do we need them to be > >> sub-issues? > >> > >> > >> > https://issues.apache.org/jira/projects/OPENNLP/issues/OPENNLP-1926?filter=allopenissues > >> > >> But glad we're both fans of epics. > >> > >> On Fri, Sep 11, 2026 at 3:53 AM Martin Wiesner <[email protected]> > >> wrote: > >> > >>> I’m also +1 for this direction. We should aim for this to be achieved > for > >>> the 3.0.0 (GA) release in November (see roadmap in Jira). > >>> Think a separate epic with sub-issues might be helpful. > >>> > >>> Thanks, > >>> Martin > >>> > >>> > >>>> Am 09.09.2026 um 14:55 schrieb Jeff Zemerick <[email protected]>: > >>>> > >>>> I'm in favor of removing regex usage where possible for the reasons > >>>> you gave, and because regex usage is often a source of CVEs. > >>>> > >>>> Thanks, > >>>> Jeff > >>>> > >>>> On Mon, Sep 7, 2026 at 2:39 PM Kristian Rickert <[email protected]> > >>> wrote: > >>>>> > >>>>> Hey everyone, > >>>>> > >>>>> While reviewing the code I had noticed a few places where we are > still > >>>>> using regex unnecessarily. These improvements aren't a major rush, > >>> and I > >>>>> have one epic tracking all of the necessary changes. These > >>> improvements > >>>>> are all heavily tested. Only the first one is stacked because it > >>> improves > >>>>> StringUtil, which the others depend on. That means no major rush > >>> merging > >>>>> them and none are blockers for a 3.0 release (but would be a great > >>> story to > >>>>> tell > >>>>> > >>>>> So OPENNLP-1928 is the only one stacked; the rest will hang off of > >>> main and > >>>>> merge cleanly. > >>>>> > >>>>> After the first ticket is done, I'll create a speed test to show any > >>>>> performance/memory improvement from the fix. As we've seen from the > >>> other > >>>>> tickets where we reduced RegEx, using cursors and string pointers > over > >>>>> regex gives us far less memory churn and better speed. I also have > an > >>>>> easier time understanding it as I never really "get" a lot of regex. > >>>>> > >>>>> Feel free to chime in, make changes, test, or review... > >>>>> > >>>>> Here's the list: > >>>>> Ticket Scope PR > >>>>> OPENNLP-1928 <https://issues.apache.org/jira/browse/OPENNLP-1928> > >>> part 1: > >>>>> trivial batch, plus the shared StringUtil helpers isAsciiWhitespace, > >>>>> splitOnAsciiWhitespace, containsAsciiUpperCase, containsAsciiDigit > >>>>> apache/opennlp#1275 > >>>>> OPENNLP-1930 <https://issues.apache.org/jira/browse/OPENNLP-1930> > >>> part 2: > >>>>> Arvores Deitadas markup parsing apache/opennlp#1276 > >>>>> OPENNLP-1931 <https://issues.apache.org/jira/browse/OPENNLP-1931> > >>> part 3: > >>>>> JSON vocabulary and id2label scrape in opennlp-dl apache/opennlp#1277 > >>>>> OPENNLP-1932 <https://issues.apache.org/jira/browse/OPENNLP-1932> > >>> part 4: > >>>>> wildcard matching in the model resolver apache/opennlp#1278 > >>>>> OPENNLP-1933 <https://issues.apache.org/jira/browse/OPENNLP-1933> > >>> part 5: > >>>>> per-call String regex splits and replacements apache/opennlp#1279 > >>>>> OPENNLP-1929 <https://issues.apache.org/jira/browse/OPENNLP-1929> > bug: > >>>>> BasicContextGenerator splits on its separator as a regular expression > >>>>> apache/opennlp#1280 > >>>>> OPENNLP-1934 <https://issues.apache.org/jira/browse/OPENNLP-1934> > >>> part 6: > >>>>> tokenizer alphanumeric pattern evaluated as a character set > >>>>> apache/opennlp#1281 > >>>>> OPENNLP-1935 <https://issues.apache.org/jira/browse/OPENNLP-1935> > >>> part 7: > >>>>> checkstyle guard and the exempt list in checkstyle-suppressions.xml > >>>>> apache/opennlp#1282 > >>> > >>> > >
