Hi team,

Thank you for weighing in. I was feeling overwhelmed thinking about it and
am relieved to hear we will focus on getting version 3.0 released.

There is no rush when I post these issues. All new software releases come
with bugs, but the release is still in great shape. Before posting any
issues, I will take the time to validate them with solid examples.

Version 3.0-SNAPSHOT is in a good place. Assuming everyone agrees, my work
will focus on chipping away at fixing bugs, anti-regex, creating backports
of bugs, and grpc at a steady pace. This will give me a solid, challenging
queue focused on stability and quality, while helping me become more
familiar with other parts of the codebase.

I'll continue this effort until the queue is fully managable. I plan to
limit any new PRs to 0–2 in the near future so I can focus on reviewing PRs
from others. I will also continue working on gRPC; so feel free to check
the sandbox or ask me directly if you're interested in the current progress
/ would like to help.  I'll regularly post updates there..

Initial tests for the anti-regex work are showing a conservative 2x+
speedup in many areas. I am truly excited about how much work has been done
and am looking forward to the release. I have learned a lot from everyone
here.

Best regards,
Kristian

P.S. I would prefer not to post the very first issue, so please feel free
to open one!


On Sun, Sep 20, 2026, 11:52 AM Richard Zowalla <[email protected]> wrote:

> Ok, I think I am fine if we postpone things to 3.1.x where they need more
> time. Makes sense to re-schedule the related Jiras to 3.1.x then, so the
> 3.0.0 scope stays predictable and the regex work can go step by step
> instead of as a rush.
>
> +1 on staying on the regex path rather than opening 16 tickets at once.
> Whack-a-mole is a real risk, and the validation work is worth more than the
> ticket count. S
> ame for the correctness backports to 2.x - those are the ones that matter,
> and they can flow at their own pace once all is validated.
>
> Side note on the ticket migration: we now have GitHub issues enabled, and
> Jira is going read-only for the public but stays writable for PMC and
> committers.
> So there's no need to block on migration, we can leave the existing issues
> where they are for now (at least for 3.0.0), and once 3.x is out we can
> migrate the remaining open ones to GH.
>
> Re: the queue: agreed. Anything green with no outstanding comments is fair
> game for review now.
>
> If we all agree, that we don’t need the whole regex removal stuff in 3.x.,
> we are fine.
>
> Gruß
> R
>
> > Am 20.09.2026 um 14:12 schrieb Kristian Rickert <[email protected]>:
> >
> > The regex sweep uncovered several bugs.   A thorough regex pass would fix
> > them and most fixes should be backported to 2.x - especially the
> > correctness ones.  This doesn't include the embedding issues recently
> > filed, which would also need backporting.
> >
> > It doesn't change the merge queue, it just means a bigger queue to
> process
> > everything.  I won't post anything until I complete the validation work.
> > I'll stick to the current queue, but I'd like the team's thoughts on how
> to
> > triage this.
> >
> > Posting 16 bugs might cause a whack-a-mole effect and will take too much
> > time away from the first sweep.  I'll post to our issues after ticket
> > migration happens.
> >
> > I think staying methodically on the regex path is the way to go.
> >
> > There are a few tickets that are mergeable; anything green should OK to
> > review if there's no outstanding comment.  Drafts are not ready for
> review
> > (by definition), but comments are always welcome.
> >
> >
> >
> > On Fri, Sep 11, 2026 at 8:08 AM Kristian Rickert <[email protected]>
> wrote:
> >
> >> Martin, I made them an epic with linked issues. Do we need them to be
> >> sub-issues?
> >>
> >>
> >>
> https://issues.apache.org/jira/projects/OPENNLP/issues/OPENNLP-1926?filter=allopenissues
> >>
> >> But glad we're both fans of epics.
> >>
> >> On Fri, Sep 11, 2026 at 3:53 AM Martin Wiesner <[email protected]>
> >> wrote:
> >>
> >>> I’m also +1 for this direction. We should aim for this to be achieved
> for
> >>> the 3.0.0 (GA) release in November (see roadmap in Jira).
> >>> Think a separate epic with sub-issues might be helpful.
> >>>
> >>> Thanks,
> >>> Martin
> >>>
> >>>
> >>>> Am 09.09.2026 um 14:55 schrieb Jeff Zemerick <[email protected]>:
> >>>>
> >>>> I'm in favor of removing regex usage where possible for the reasons
> >>>> you gave, and because regex usage is often a source of CVEs.
> >>>>
> >>>> Thanks,
> >>>> Jeff
> >>>>
> >>>> On Mon, Sep 7, 2026 at 2:39 PM Kristian Rickert <[email protected]>
> >>> wrote:
> >>>>>
> >>>>> Hey everyone,
> >>>>>
> >>>>> While reviewing the code I had noticed a few places where we are
> still
> >>>>> using regex unnecessarily.  These improvements aren't a major rush,
> >>> and I
> >>>>> have one epic tracking all of the necessary changes.  These
> >>> improvements
> >>>>> are all heavily tested. Only the first one is stacked because it
> >>> improves
> >>>>> StringUtil, which the others depend on. That means no major rush
> >>> merging
> >>>>> them and none are blockers for a 3.0 release (but would be a great
> >>> story to
> >>>>> tell
> >>>>>
> >>>>> So OPENNLP-1928 is the only one stacked; the rest will hang off of
> >>> main and
> >>>>> merge cleanly.
> >>>>>
> >>>>> After the first ticket is done, I'll create a speed test to show any
> >>>>> performance/memory improvement from the fix.  As we've seen from the
> >>> other
> >>>>> tickets where we reduced RegEx, using cursors and string pointers
> over
> >>>>> regex gives us far less memory churn and better speed.  I also have
> an
> >>>>> easier time understanding it as I never really "get" a lot of regex.
> >>>>>
> >>>>> Feel free to chime in, make changes, test, or review...
> >>>>>
> >>>>> Here's the list:
> >>>>> Ticket Scope PR
> >>>>> OPENNLP-1928 <https://issues.apache.org/jira/browse/OPENNLP-1928>
> >>> part 1:
> >>>>> trivial batch, plus the shared StringUtil helpers isAsciiWhitespace,
> >>>>> splitOnAsciiWhitespace, containsAsciiUpperCase, containsAsciiDigit
> >>>>> apache/opennlp#1275
> >>>>> OPENNLP-1930 <https://issues.apache.org/jira/browse/OPENNLP-1930>
> >>> part 2:
> >>>>> Arvores Deitadas markup parsing apache/opennlp#1276
> >>>>> OPENNLP-1931 <https://issues.apache.org/jira/browse/OPENNLP-1931>
> >>> part 3:
> >>>>> JSON vocabulary and id2label scrape in opennlp-dl apache/opennlp#1277
> >>>>> OPENNLP-1932 <https://issues.apache.org/jira/browse/OPENNLP-1932>
> >>> part 4:
> >>>>> wildcard matching in the model resolver apache/opennlp#1278
> >>>>> OPENNLP-1933 <https://issues.apache.org/jira/browse/OPENNLP-1933>
> >>> part 5:
> >>>>> per-call String regex splits and replacements apache/opennlp#1279
> >>>>> OPENNLP-1929 <https://issues.apache.org/jira/browse/OPENNLP-1929>
> bug:
> >>>>> BasicContextGenerator splits on its separator as a regular expression
> >>>>> apache/opennlp#1280
> >>>>> OPENNLP-1934 <https://issues.apache.org/jira/browse/OPENNLP-1934>
> >>> part 6:
> >>>>> tokenizer alphanumeric pattern evaluated as a character set
> >>>>> apache/opennlp#1281
> >>>>> OPENNLP-1935 <https://issues.apache.org/jira/browse/OPENNLP-1935>
> >>> part 7:
> >>>>> checkstyle guard and the exempt list in checkstyle-suppressions.xml
> >>>>> apache/opennlp#1282
> >>>
> >>>
>
>

Reply via email to