Just bumping this thread to see if there is any feedback. If there's
no further feedback after a few days, I second Richard's proposal to
include those guidelines in the contribution guidelines and PR
template, and to hold a vote to adopt the AI Policy adopted by Lucene.

Thanks,
Jeff

On Thu, Sep 17, 2026 at 6:07 AM Jeff Zemerick <[email protected]> wrote:
>
> Good points. I agree with putting all of these points in the
> contributor guidelines, and I especially like the expectations for
> committers in the last section. Ultimately, any policy affects the
> committers the most.
>
> It's reassuring to see that OpenNLP is not the only ASF (and I'm sure
> non-ASF) project adapting to the new ways of working. Lucene has
> recently adopted an AI Policy [1] over using AI-generated responses
> for communication. I would favor the same or a similar policy for
> OpenNLP, too. While I don't think OpenNLP is facing this issue to the
> degree Lucene is, I think it would be good to go ahead and get a
> policy in place becase it's probably only a matter of time until
> OpenNLP is.
>
> Thanks,
> Jeff
>
> [1] https://github.com/apache/lucene/blob/main/AI_POLICY.md
>
> On Thu, Sep 17, 2026 at 5:53 AM Richard Zowalla <[email protected]> wrote:
> >
> > Hi all,
> >
> > Over the last few weeks the number and size of open pull requests has grown 
> > far beyond what we can review, and I think we all share part of the blame.
> >
> > So I'd like to agree on a few practices for everyone, contributors and 
> > committers alike.
> >
> > 1.) Why this matters
> >
> > Every change we merge has to be read and understood by a human committer, 
> > who then takes responsibility for it. That's how Apache works, and a green 
> > build or a machine-generated review doesn't replace it. I'm including 
> > myself here: I've posted LLM-generated reviews on some PRs when I was short 
> > on time. They can help, but they aren't a review, and we shouldn't pretend 
> > they are.
> >
> > This has also become a problem for the community itself. Some people have 
> > told me privately that they can't handle the review load anymore, and some 
> > have stopped reviewing completely. That's the worst outcome for a volunteer 
> > project. Reviewers are our scarcest resource, and if we burn them out, 
> > nothing gets merged, no matter how good the code is. I don't want anyone to 
> > feel they have to step back because the queue has become impossible to keep 
> > up with.
> >
> > Right now there are about 24 open PRs [1] adding about 246k lines. Much of 
> > that code is duplicated across PRs, and nobody can tell which part of a 
> > diff belongs to which change; description or comments are outdated and a 
> > human is totally lost. Some examples, only to show the pattern:
> >
> > - Stacked PRs that all target main repeat their parents' diffs: #1213, 
> > #1214 and #1215 each contain all of #1152 (+30k to +40k lines, 138–165 
> > commits each) [2], and #1276–#1282 all contain #1275 [3].
> > - The same code is in several PRs at once (the TextEmbedder SPI is in #1290 
> > and #1152) [4], so review comments on one copy never reach the others.
> > - Descriptions no longer match the diff, for example PRs that list 
> > dependencies which were merged long ago, or file counts and scope that have 
> > since changed [5]. Several of us have left comments saying we simply don't 
> > understand what a PR is about.
> > - Some PRs have 100+ commits, including merge commits from main.
> > - New scope gets added to PRs while they are under review, and new drafts 
> > keep opening before the existing ones are done.
> >
> > None of this says the work is bad. Much of it fixes real bugs and adds 
> > useful features. But a PR that a volunteer can't review in one sitting 
> > doesn't get a real review.
> > It either sits there (forever) or gets merged on trust, and neither is 
> > acceptable
> >
> > 2.) Proposal
> >
> > A. One PR, one JIRA, one unit a human can review.
> >
> > Squash to a single commit, or to a few commits that each make sense on 
> > their own. Rebase on main instead of merging main in. The title and 
> > description should say what the PR does *now*: rewrite them when the scope 
> > changes rather than adding "review round" sections, since the history is 
> > already in the comments.
> >
> > B. No stacking.
> >
> > A PR targets main and contains only its own change. If it depends on 
> > another open PR, it waits on the contributor's fork until that parent is 
> > merged. A PR should never show the diff of another open PR.
> >
> > C. Optional modules go to opennlp-addons.
> >
> > New optional modules with their own data, dependencies or follow-up plans 
> > belong in opennlp-addons [6]. The current embeddings, vector index, 
> > gazetteer and wordnet drafts are examples [7]. Core only gets the small API 
> > contracts those modules really need, each in its own focused PR.
> >
> > D. Finish before starting something new.
> >
> > Limit how many PRs one person has ready for review at a time (3–4, say). 
> > Don't add new scope to a PR that is under review; put new ideas in JIRA 
> > first.
> > The speed at which we can review sets how fast code gets in, not the speed 
> > at which code gets written. Five finished PRs are worth more to us than 24 
> > open ones.
> >
> > E. Committers, too.
> >
> > We should review in time and reply clearly: approve, request specific 
> > changes, or say it doesn't fit and close it. We shouldn't post generated 
> > reviews as if they were our own. If we use tools, we check and sign off on 
> > every point we post. And it's fine to say "this is too big for me to 
> > review", which is useful feedback, not a failure.
> >
> >
> > If we agree, we should add this to the contribution guidelines and the PR 
> > template.
> > For the PRs already open, I suggest we agree on a merge order together, and 
> > close or park stacked and duplicated ones until their turn comes.
> >
> > And to those who stepped back: thank you for everything you've reviewed so 
> > far. I hope this makes it possible for you to come back.
> >
> > Thoughts?
> >
> > Gruß
> > Richard
> >
> > ---
> >
> > ## References
> >
> > **[1] Open pull requests**
> > - https://github.com/apache/opennlp/pulls
> >
> > **[2] Static embeddings stack**
> > - https://github.com/apache/opennlp/pull/1152 
> > ([OPENNLP-1877](https://issues.apache.org/jira/browse/OPENNLP-1877))
> > - https://github.com/apache/opennlp/pull/1213 
> > ([OPENNLP-1895](https://issues.apache.org/jira/browse/OPENNLP-1895))
> > - https://github.com/apache/opennlp/pull/1214 
> > ([OPENNLP-1910](https://issues.apache.org/jira/browse/OPENNLP-1910))
> > - https://github.com/apache/opennlp/pull/1215 
> > ([OPENNLP-1911](https://issues.apache.org/jira/browse/OPENNLP-1911))
> >
> > **[3] Regex removal**
> > - https://github.com/apache/opennlp/pull/1275 
> > ([OPENNLP-1928](https://issues.apache.org/jira/browse/OPENNLP-1928))
> > - https://github.com/apache/opennlp/pull/1276 
> > ([OPENNLP-1930](https://issues.apache.org/jira/browse/OPENNLP-1930))
> > - https://github.com/apache/opennlp/pull/1277 
> > ([OPENNLP-1931](https://issues.apache.org/jira/browse/OPENNLP-1931))
> > - https://github.com/apache/opennlp/pull/1278 
> > ([OPENNLP-1932](https://issues.apache.org/jira/browse/OPENNLP-1932))
> > - https://github.com/apache/opennlp/pull/1279 
> > ([OPENNLP-1933](https://issues.apache.org/jira/browse/OPENNLP-1933))
> > - https://github.com/apache/opennlp/pull/1280 
> > ([OPENNLP-1929](https://issues.apache.org/jira/browse/OPENNLP-1929))
> > - https://github.com/apache/opennlp/pull/1281 
> > ([OPENNLP-1934](https://issues.apache.org/jira/browse/OPENNLP-1934))
> > - https://github.com/apache/opennlp/pull/1282 
> > ([OPENNLP-1935](https://issues.apache.org/jira/browse/OPENNLP-1935))
> >
> > **[4] TextEmbedder SPI and shared deep-learning encoder tests**
> > - https://github.com/apache/opennlp/pull/1290 
> > ([OPENNLP-1937](https://issues.apache.org/jira/browse/OPENNLP-1937))
> > - https://github.com/apache/opennlp/pull/1288 
> > ([OPENNLP-1942](https://issues.apache.org/jira/browse/OPENNLP-1942))
> >
> > **[5] Dependency parser stack**
> > - https://github.com/apache/opennlp/pull/1236 
> > ([OPENNLP-547](https://issues.apache.org/jira/browse/OPENNLP-547))
> > - https://github.com/apache/opennlp/pull/1237 
> > ([OPENNLP-1919](https://issues.apache.org/jira/browse/OPENNLP-1919))
> > - https://github.com/apache/opennlp/pull/1238 
> > ([OPENNLP-1920](https://issues.apache.org/jira/browse/OPENNLP-1920))
> >
> > **[6] opennlp-addons, with an example add-on PR**
> > - https://github.com/apache/opennlp-addons
> > - https://github.com/apache/opennlp-addons/pull/184 
> > ([OPENNLP-1885](https://issues.apache.org/jira/browse/OPENNLP-1885))
> >
> > **[7] Drafts that belong in opennlp-addons**
> > - https://github.com/apache/opennlp/pull/1154 
> > ([OPENNLP-1879](https://issues.apache.org/jira/browse/OPENNLP-1879))
> > - https://github.com/apache/opennlp/pull/1155 
> > ([OPENNLP-1880](https://issues.apache.org/jira/browse/OPENNLP-1880))
> > - https://github.com/apache/opennlp/pull/1167 
> > ([OPENNLP-1887](https://issues.apache.org/jira/browse/OPENNLP-1887))
> > - plus the embeddings stack in [2]

Reply via email to