Ok, I think I am fine if we postpone things to 3.1.x where they need more time. 
Makes sense to re-schedule the related Jiras to 3.1.x then, so the 3.0.0 scope 
stays predictable and the regex work can go step by step instead of as a rush.

+1 on staying on the regex path rather than opening 16 tickets at once. 
Whack-a-mole is a real risk, and the validation work is worth more than the 
ticket count. S
ame for the correctness backports to 2.x - those are the ones that matter, and 
they can flow at their own pace once all is validated.

Side note on the ticket migration: we now have GitHub issues enabled, and Jira 
is going read-only for the public but stays writable for PMC and committers. 
So there's no need to block on migration, we can leave the existing issues 
where they are for now (at least for 3.0.0), and once 3.x is out we can migrate 
the remaining open ones to GH.

Re: the queue: agreed. Anything green with no outstanding comments is fair game 
for review now. 

If we all agree, that we don’t need the whole regex removal stuff in 3.x., we 
are fine. 

Gruß
R

> Am 20.09.2026 um 14:12 schrieb Kristian Rickert <[email protected]>:
> 
> The regex sweep uncovered several bugs.   A thorough regex pass would fix
> them and most fixes should be backported to 2.x - especially the
> correctness ones.  This doesn't include the embedding issues recently
> filed, which would also need backporting.
> 
> It doesn't change the merge queue, it just means a bigger queue to process
> everything.  I won't post anything until I complete the validation work.
> I'll stick to the current queue, but I'd like the team's thoughts on how to
> triage this.
> 
> Posting 16 bugs might cause a whack-a-mole effect and will take too much
> time away from the first sweep.  I'll post to our issues after ticket
> migration happens.
> 
> I think staying methodically on the regex path is the way to go.
> 
> There are a few tickets that are mergeable; anything green should OK to
> review if there's no outstanding comment.  Drafts are not ready for review
> (by definition), but comments are always welcome.
> 
> 
> 
> On Fri, Sep 11, 2026 at 8:08 AM Kristian Rickert <[email protected]> wrote:
> 
>> Martin, I made them an epic with linked issues. Do we need them to be
>> sub-issues?
>> 
>> 
>> https://issues.apache.org/jira/projects/OPENNLP/issues/OPENNLP-1926?filter=allopenissues
>> 
>> But glad we're both fans of epics.
>> 
>> On Fri, Sep 11, 2026 at 3:53 AM Martin Wiesner <[email protected]>
>> wrote:
>> 
>>> I’m also +1 for this direction. We should aim for this to be achieved for
>>> the 3.0.0 (GA) release in November (see roadmap in Jira).
>>> Think a separate epic with sub-issues might be helpful.
>>> 
>>> Thanks,
>>> Martin
>>> 
>>> 
>>>> Am 09.09.2026 um 14:55 schrieb Jeff Zemerick <[email protected]>:
>>>> 
>>>> I'm in favor of removing regex usage where possible for the reasons
>>>> you gave, and because regex usage is often a source of CVEs.
>>>> 
>>>> Thanks,
>>>> Jeff
>>>> 
>>>> On Mon, Sep 7, 2026 at 2:39 PM Kristian Rickert <[email protected]>
>>> wrote:
>>>>> 
>>>>> Hey everyone,
>>>>> 
>>>>> While reviewing the code I had noticed a few places where we are still
>>>>> using regex unnecessarily.  These improvements aren't a major rush,
>>> and I
>>>>> have one epic tracking all of the necessary changes.  These
>>> improvements
>>>>> are all heavily tested. Only the first one is stacked because it
>>> improves
>>>>> StringUtil, which the others depend on. That means no major rush
>>> merging
>>>>> them and none are blockers for a 3.0 release (but would be a great
>>> story to
>>>>> tell
>>>>> 
>>>>> So OPENNLP-1928 is the only one stacked; the rest will hang off of
>>> main and
>>>>> merge cleanly.
>>>>> 
>>>>> After the first ticket is done, I'll create a speed test to show any
>>>>> performance/memory improvement from the fix.  As we've seen from the
>>> other
>>>>> tickets where we reduced RegEx, using cursors and string pointers over
>>>>> regex gives us far less memory churn and better speed.  I also have an
>>>>> easier time understanding it as I never really "get" a lot of regex.
>>>>> 
>>>>> Feel free to chime in, make changes, test, or review...
>>>>> 
>>>>> Here's the list:
>>>>> Ticket Scope PR
>>>>> OPENNLP-1928 <https://issues.apache.org/jira/browse/OPENNLP-1928>
>>> part 1:
>>>>> trivial batch, plus the shared StringUtil helpers isAsciiWhitespace,
>>>>> splitOnAsciiWhitespace, containsAsciiUpperCase, containsAsciiDigit
>>>>> apache/opennlp#1275
>>>>> OPENNLP-1930 <https://issues.apache.org/jira/browse/OPENNLP-1930>
>>> part 2:
>>>>> Arvores Deitadas markup parsing apache/opennlp#1276
>>>>> OPENNLP-1931 <https://issues.apache.org/jira/browse/OPENNLP-1931>
>>> part 3:
>>>>> JSON vocabulary and id2label scrape in opennlp-dl apache/opennlp#1277
>>>>> OPENNLP-1932 <https://issues.apache.org/jira/browse/OPENNLP-1932>
>>> part 4:
>>>>> wildcard matching in the model resolver apache/opennlp#1278
>>>>> OPENNLP-1933 <https://issues.apache.org/jira/browse/OPENNLP-1933>
>>> part 5:
>>>>> per-call String regex splits and replacements apache/opennlp#1279
>>>>> OPENNLP-1929 <https://issues.apache.org/jira/browse/OPENNLP-1929> bug:
>>>>> BasicContextGenerator splits on its separator as a regular expression
>>>>> apache/opennlp#1280
>>>>> OPENNLP-1934 <https://issues.apache.org/jira/browse/OPENNLP-1934>
>>> part 6:
>>>>> tokenizer alphanumeric pattern evaluated as a character set
>>>>> apache/opennlp#1281
>>>>> OPENNLP-1935 <https://issues.apache.org/jira/browse/OPENNLP-1935>
>>> part 7:
>>>>> checkstyle guard and the exempt list in checkstyle-suppressions.xml
>>>>> apache/opennlp#1282
>>> 
>>> 

Reply via email to