Just bumping this thread to see if there is any feedback. If there's no further feedback after a few days, I second Richard's proposal to include those guidelines in the contribution guidelines and PR template, and to hold a vote to adopt the AI Policy adopted by Lucene.
Thanks, Jeff On Thu, Sep 17, 2026 at 6:07 AM Jeff Zemerick <[email protected]> wrote: > > Good points. I agree with putting all of these points in the > contributor guidelines, and I especially like the expectations for > committers in the last section. Ultimately, any policy affects the > committers the most. > > It's reassuring to see that OpenNLP is not the only ASF (and I'm sure > non-ASF) project adapting to the new ways of working. Lucene has > recently adopted an AI Policy [1] over using AI-generated responses > for communication. I would favor the same or a similar policy for > OpenNLP, too. While I don't think OpenNLP is facing this issue to the > degree Lucene is, I think it would be good to go ahead and get a > policy in place becase it's probably only a matter of time until > OpenNLP is. > > Thanks, > Jeff > > [1] https://github.com/apache/lucene/blob/main/AI_POLICY.md > > On Thu, Sep 17, 2026 at 5:53 AM Richard Zowalla <[email protected]> wrote: > > > > Hi all, > > > > Over the last few weeks the number and size of open pull requests has grown > > far beyond what we can review, and I think we all share part of the blame. > > > > So I'd like to agree on a few practices for everyone, contributors and > > committers alike. > > > > 1.) Why this matters > > > > Every change we merge has to be read and understood by a human committer, > > who then takes responsibility for it. That's how Apache works, and a green > > build or a machine-generated review doesn't replace it. I'm including > > myself here: I've posted LLM-generated reviews on some PRs when I was short > > on time. They can help, but they aren't a review, and we shouldn't pretend > > they are. > > > > This has also become a problem for the community itself. Some people have > > told me privately that they can't handle the review load anymore, and some > > have stopped reviewing completely. That's the worst outcome for a volunteer > > project. Reviewers are our scarcest resource, and if we burn them out, > > nothing gets merged, no matter how good the code is. I don't want anyone to > > feel they have to step back because the queue has become impossible to keep > > up with. > > > > Right now there are about 24 open PRs [1] adding about 246k lines. Much of > > that code is duplicated across PRs, and nobody can tell which part of a > > diff belongs to which change; description or comments are outdated and a > > human is totally lost. Some examples, only to show the pattern: > > > > - Stacked PRs that all target main repeat their parents' diffs: #1213, > > #1214 and #1215 each contain all of #1152 (+30k to +40k lines, 138–165 > > commits each) [2], and #1276–#1282 all contain #1275 [3]. > > - The same code is in several PRs at once (the TextEmbedder SPI is in #1290 > > and #1152) [4], so review comments on one copy never reach the others. > > - Descriptions no longer match the diff, for example PRs that list > > dependencies which were merged long ago, or file counts and scope that have > > since changed [5]. Several of us have left comments saying we simply don't > > understand what a PR is about. > > - Some PRs have 100+ commits, including merge commits from main. > > - New scope gets added to PRs while they are under review, and new drafts > > keep opening before the existing ones are done. > > > > None of this says the work is bad. Much of it fixes real bugs and adds > > useful features. But a PR that a volunteer can't review in one sitting > > doesn't get a real review. > > It either sits there (forever) or gets merged on trust, and neither is > > acceptable > > > > 2.) Proposal > > > > A. One PR, one JIRA, one unit a human can review. > > > > Squash to a single commit, or to a few commits that each make sense on > > their own. Rebase on main instead of merging main in. The title and > > description should say what the PR does *now*: rewrite them when the scope > > changes rather than adding "review round" sections, since the history is > > already in the comments. > > > > B. No stacking. > > > > A PR targets main and contains only its own change. If it depends on > > another open PR, it waits on the contributor's fork until that parent is > > merged. A PR should never show the diff of another open PR. > > > > C. Optional modules go to opennlp-addons. > > > > New optional modules with their own data, dependencies or follow-up plans > > belong in opennlp-addons [6]. The current embeddings, vector index, > > gazetteer and wordnet drafts are examples [7]. Core only gets the small API > > contracts those modules really need, each in its own focused PR. > > > > D. Finish before starting something new. > > > > Limit how many PRs one person has ready for review at a time (3–4, say). > > Don't add new scope to a PR that is under review; put new ideas in JIRA > > first. > > The speed at which we can review sets how fast code gets in, not the speed > > at which code gets written. Five finished PRs are worth more to us than 24 > > open ones. > > > > E. Committers, too. > > > > We should review in time and reply clearly: approve, request specific > > changes, or say it doesn't fit and close it. We shouldn't post generated > > reviews as if they were our own. If we use tools, we check and sign off on > > every point we post. And it's fine to say "this is too big for me to > > review", which is useful feedback, not a failure. > > > > > > If we agree, we should add this to the contribution guidelines and the PR > > template. > > For the PRs already open, I suggest we agree on a merge order together, and > > close or park stacked and duplicated ones until their turn comes. > > > > And to those who stepped back: thank you for everything you've reviewed so > > far. I hope this makes it possible for you to come back. > > > > Thoughts? > > > > Gruß > > Richard > > > > --- > > > > ## References > > > > **[1] Open pull requests** > > - https://github.com/apache/opennlp/pulls > > > > **[2] Static embeddings stack** > > - https://github.com/apache/opennlp/pull/1152 > > ([OPENNLP-1877](https://issues.apache.org/jira/browse/OPENNLP-1877)) > > - https://github.com/apache/opennlp/pull/1213 > > ([OPENNLP-1895](https://issues.apache.org/jira/browse/OPENNLP-1895)) > > - https://github.com/apache/opennlp/pull/1214 > > ([OPENNLP-1910](https://issues.apache.org/jira/browse/OPENNLP-1910)) > > - https://github.com/apache/opennlp/pull/1215 > > ([OPENNLP-1911](https://issues.apache.org/jira/browse/OPENNLP-1911)) > > > > **[3] Regex removal** > > - https://github.com/apache/opennlp/pull/1275 > > ([OPENNLP-1928](https://issues.apache.org/jira/browse/OPENNLP-1928)) > > - https://github.com/apache/opennlp/pull/1276 > > ([OPENNLP-1930](https://issues.apache.org/jira/browse/OPENNLP-1930)) > > - https://github.com/apache/opennlp/pull/1277 > > ([OPENNLP-1931](https://issues.apache.org/jira/browse/OPENNLP-1931)) > > - https://github.com/apache/opennlp/pull/1278 > > ([OPENNLP-1932](https://issues.apache.org/jira/browse/OPENNLP-1932)) > > - https://github.com/apache/opennlp/pull/1279 > > ([OPENNLP-1933](https://issues.apache.org/jira/browse/OPENNLP-1933)) > > - https://github.com/apache/opennlp/pull/1280 > > ([OPENNLP-1929](https://issues.apache.org/jira/browse/OPENNLP-1929)) > > - https://github.com/apache/opennlp/pull/1281 > > ([OPENNLP-1934](https://issues.apache.org/jira/browse/OPENNLP-1934)) > > - https://github.com/apache/opennlp/pull/1282 > > ([OPENNLP-1935](https://issues.apache.org/jira/browse/OPENNLP-1935)) > > > > **[4] TextEmbedder SPI and shared deep-learning encoder tests** > > - https://github.com/apache/opennlp/pull/1290 > > ([OPENNLP-1937](https://issues.apache.org/jira/browse/OPENNLP-1937)) > > - https://github.com/apache/opennlp/pull/1288 > > ([OPENNLP-1942](https://issues.apache.org/jira/browse/OPENNLP-1942)) > > > > **[5] Dependency parser stack** > > - https://github.com/apache/opennlp/pull/1236 > > ([OPENNLP-547](https://issues.apache.org/jira/browse/OPENNLP-547)) > > - https://github.com/apache/opennlp/pull/1237 > > ([OPENNLP-1919](https://issues.apache.org/jira/browse/OPENNLP-1919)) > > - https://github.com/apache/opennlp/pull/1238 > > ([OPENNLP-1920](https://issues.apache.org/jira/browse/OPENNLP-1920)) > > > > **[6] opennlp-addons, with an example add-on PR** > > - https://github.com/apache/opennlp-addons > > - https://github.com/apache/opennlp-addons/pull/184 > > ([OPENNLP-1885](https://issues.apache.org/jira/browse/OPENNLP-1885)) > > > > **[7] Drafts that belong in opennlp-addons** > > - https://github.com/apache/opennlp/pull/1154 > > ([OPENNLP-1879](https://issues.apache.org/jira/browse/OPENNLP-1879)) > > - https://github.com/apache/opennlp/pull/1155 > > ([OPENNLP-1880](https://issues.apache.org/jira/browse/OPENNLP-1880)) > > - https://github.com/apache/opennlp/pull/1167 > > ([OPENNLP-1887](https://issues.apache.org/jira/browse/OPENNLP-1887)) > > - plus the embeddings stack in [2]
