The regex sweep uncovered several bugs. A thorough regex pass would fix them and most fixes should be backported to 2.x - especially the correctness ones. This doesn't include the embedding issues recently filed, which would also need backporting.
It doesn't change the merge queue, it just means a bigger queue to process everything. I won't post anything until I complete the validation work. I'll stick to the current queue, but I'd like the team's thoughts on how to triage this. Posting 16 bugs might cause a whack-a-mole effect and will take too much time away from the first sweep. I'll post to our issues after ticket migration happens. I think staying methodically on the regex path is the way to go. There are a few tickets that are mergeable; anything green should OK to review if there's no outstanding comment. Drafts are not ready for review (by definition), but comments are always welcome. On Fri, Sep 11, 2026 at 8:08 AM Kristian Rickert <[email protected]> wrote: > Martin, I made them an epic with linked issues. Do we need them to be > sub-issues? > > > https://issues.apache.org/jira/projects/OPENNLP/issues/OPENNLP-1926?filter=allopenissues > > But glad we're both fans of epics. > > On Fri, Sep 11, 2026 at 3:53 AM Martin Wiesner <[email protected]> > wrote: > >> I’m also +1 for this direction. We should aim for this to be achieved for >> the 3.0.0 (GA) release in November (see roadmap in Jira). >> Think a separate epic with sub-issues might be helpful. >> >> Thanks, >> Martin >> >> >> > Am 09.09.2026 um 14:55 schrieb Jeff Zemerick <[email protected]>: >> > >> > I'm in favor of removing regex usage where possible for the reasons >> > you gave, and because regex usage is often a source of CVEs. >> > >> > Thanks, >> > Jeff >> > >> > On Mon, Sep 7, 2026 at 2:39 PM Kristian Rickert <[email protected]> >> wrote: >> >> >> >> Hey everyone, >> >> >> >> While reviewing the code I had noticed a few places where we are still >> >> using regex unnecessarily. These improvements aren't a major rush, >> and I >> >> have one epic tracking all of the necessary changes. These >> improvements >> >> are all heavily tested. Only the first one is stacked because it >> improves >> >> StringUtil, which the others depend on. That means no major rush >> merging >> >> them and none are blockers for a 3.0 release (but would be a great >> story to >> >> tell >> >> >> >> So OPENNLP-1928 is the only one stacked; the rest will hang off of >> main and >> >> merge cleanly. >> >> >> >> After the first ticket is done, I'll create a speed test to show any >> >> performance/memory improvement from the fix. As we've seen from the >> other >> >> tickets where we reduced RegEx, using cursors and string pointers over >> >> regex gives us far less memory churn and better speed. I also have an >> >> easier time understanding it as I never really "get" a lot of regex. >> >> >> >> Feel free to chime in, make changes, test, or review... >> >> >> >> Here's the list: >> >> Ticket Scope PR >> >> OPENNLP-1928 <https://issues.apache.org/jira/browse/OPENNLP-1928> >> part 1: >> >> trivial batch, plus the shared StringUtil helpers isAsciiWhitespace, >> >> splitOnAsciiWhitespace, containsAsciiUpperCase, containsAsciiDigit >> >> apache/opennlp#1275 >> >> OPENNLP-1930 <https://issues.apache.org/jira/browse/OPENNLP-1930> >> part 2: >> >> Arvores Deitadas markup parsing apache/opennlp#1276 >> >> OPENNLP-1931 <https://issues.apache.org/jira/browse/OPENNLP-1931> >> part 3: >> >> JSON vocabulary and id2label scrape in opennlp-dl apache/opennlp#1277 >> >> OPENNLP-1932 <https://issues.apache.org/jira/browse/OPENNLP-1932> >> part 4: >> >> wildcard matching in the model resolver apache/opennlp#1278 >> >> OPENNLP-1933 <https://issues.apache.org/jira/browse/OPENNLP-1933> >> part 5: >> >> per-call String regex splits and replacements apache/opennlp#1279 >> >> OPENNLP-1929 <https://issues.apache.org/jira/browse/OPENNLP-1929> bug: >> >> BasicContextGenerator splits on its separator as a regular expression >> >> apache/opennlp#1280 >> >> OPENNLP-1934 <https://issues.apache.org/jira/browse/OPENNLP-1934> >> part 6: >> >> tokenizer alphanumeric pattern evaluated as a character set >> >> apache/opennlp#1281 >> >> OPENNLP-1935 <https://issues.apache.org/jira/browse/OPENNLP-1935> >> part 7: >> >> checkstyle guard and the exempt list in checkstyle-suppressions.xml >> >> apache/opennlp#1282 >> >>
