The regex sweep uncovered several bugs.   A thorough regex pass would fix
them and most fixes should be backported to 2.x - especially the
correctness ones.  This doesn't include the embedding issues recently
filed, which would also need backporting.

It doesn't change the merge queue, it just means a bigger queue to process
everything.  I won't post anything until I complete the validation work.
I'll stick to the current queue, but I'd like the team's thoughts on how to
triage this.

Posting 16 bugs might cause a whack-a-mole effect and will take too much
time away from the first sweep.  I'll post to our issues after ticket
migration happens.

I think staying methodically on the regex path is the way to go.

There are a few tickets that are mergeable; anything green should OK to
review if there's no outstanding comment.  Drafts are not ready for review
(by definition), but comments are always welcome.



On Fri, Sep 11, 2026 at 8:08 AM Kristian Rickert <[email protected]> wrote:

> Martin, I made them an epic with linked issues. Do we need them to be
> sub-issues?
>
>
> https://issues.apache.org/jira/projects/OPENNLP/issues/OPENNLP-1926?filter=allopenissues
>
> But glad we're both fans of epics.
>
> On Fri, Sep 11, 2026 at 3:53 AM Martin Wiesner <[email protected]>
> wrote:
>
>> I’m also +1 for this direction. We should aim for this to be achieved for
>> the 3.0.0 (GA) release in November (see roadmap in Jira).
>> Think a separate epic with sub-issues might be helpful.
>>
>> Thanks,
>> Martin
>>
>>
>> > Am 09.09.2026 um 14:55 schrieb Jeff Zemerick <[email protected]>:
>> >
>> > I'm in favor of removing regex usage where possible for the reasons
>> > you gave, and because regex usage is often a source of CVEs.
>> >
>> > Thanks,
>> > Jeff
>> >
>> > On Mon, Sep 7, 2026 at 2:39 PM Kristian Rickert <[email protected]>
>> wrote:
>> >>
>> >> Hey everyone,
>> >>
>> >> While reviewing the code I had noticed a few places where we are still
>> >> using regex unnecessarily.  These improvements aren't a major rush,
>> and I
>> >> have one epic tracking all of the necessary changes.  These
>> improvements
>> >> are all heavily tested. Only the first one is stacked because it
>> improves
>> >> StringUtil, which the others depend on. That means no major rush
>> merging
>> >> them and none are blockers for a 3.0 release (but would be a great
>> story to
>> >> tell
>> >>
>> >> So OPENNLP-1928 is the only one stacked; the rest will hang off of
>> main and
>> >> merge cleanly.
>> >>
>> >> After the first ticket is done, I'll create a speed test to show any
>> >> performance/memory improvement from the fix.  As we've seen from the
>> other
>> >> tickets where we reduced RegEx, using cursors and string pointers over
>> >> regex gives us far less memory churn and better speed.  I also have an
>> >> easier time understanding it as I never really "get" a lot of regex.
>> >>
>> >> Feel free to chime in, make changes, test, or review...
>> >>
>> >> Here's the list:
>> >> Ticket Scope PR
>> >> OPENNLP-1928 <https://issues.apache.org/jira/browse/OPENNLP-1928>
>> part 1:
>> >> trivial batch, plus the shared StringUtil helpers isAsciiWhitespace,
>> >> splitOnAsciiWhitespace, containsAsciiUpperCase, containsAsciiDigit
>> >> apache/opennlp#1275
>> >> OPENNLP-1930 <https://issues.apache.org/jira/browse/OPENNLP-1930>
>> part 2:
>> >> Arvores Deitadas markup parsing apache/opennlp#1276
>> >> OPENNLP-1931 <https://issues.apache.org/jira/browse/OPENNLP-1931>
>> part 3:
>> >> JSON vocabulary and id2label scrape in opennlp-dl apache/opennlp#1277
>> >> OPENNLP-1932 <https://issues.apache.org/jira/browse/OPENNLP-1932>
>> part 4:
>> >> wildcard matching in the model resolver apache/opennlp#1278
>> >> OPENNLP-1933 <https://issues.apache.org/jira/browse/OPENNLP-1933>
>> part 5:
>> >> per-call String regex splits and replacements apache/opennlp#1279
>> >> OPENNLP-1929 <https://issues.apache.org/jira/browse/OPENNLP-1929> bug:
>> >> BasicContextGenerator splits on its separator as a regular expression
>> >> apache/opennlp#1280
>> >> OPENNLP-1934 <https://issues.apache.org/jira/browse/OPENNLP-1934>
>> part 6:
>> >> tokenizer alphanumeric pattern evaluated as a character set
>> >> apache/opennlp#1281
>> >> OPENNLP-1935 <https://issues.apache.org/jira/browse/OPENNLP-1935>
>> part 7:
>> >> checkstyle guard and the exempt list in checkstyle-suppressions.xml
>> >> apache/opennlp#1282
>>
>>

Reply via email to