Sounds good. Agree on most but the stacks... - *Hand written guideline tutorial - *I'd volunteer to write up a good sorta "tutorial" that goes over "dos" and "do nots" of how you should push a PR. - *I disagree with not having stacks for complex features, but try to minimize the noise and rebase* -Stacks are normal in open source projects for larger features. I think anything on a downstream stack should remain in draft. Github even has a stack feature now. That being said, I'm fine with not doing a stack because it's not hard to have it otherwise. - *Drafts are OK if they start *green but are later converted to draft. Otherwise the work should be on a collaboration branch without being in draft. * Please let's view drafts as not reviewable* - as the button to turn green says "ready for review?" will get users regularly confused about the deviation on github. - *Limit PRs open - *Unless the team agrees on the number of PRs (i.e. the anti-regex push), we should try to limit the amount. I promised that after the current stack is done, I'll limit my PRs to 0-2 with a primary focus on bug fixing, backporting, and features to be done collaboratively/ discussed on. - *Discussions on github - *I hope we can bring more discussions to github - as design and collaboration have a lot of features you don't get in email (i.e. mermaid arch diagrams).
Thanks for bringing it to the thread. On Thu, Sep 17, 2026 at 9:07 AM Jeff Zemerick <[email protected]> wrote: > Good points. I agree with putting all of these points in the > contributor guidelines, and I especially like the expectations for > committers in the last section. Ultimately, any policy affects the > committers the most. > > It's reassuring to see that OpenNLP is not the only ASF (and I'm sure > non-ASF) project adapting to the new ways of working. Lucene has > recently adopted an AI Policy [1] over using AI-generated responses > for communication. I would favor the same or a similar policy for > OpenNLP, too. While I don't think OpenNLP is facing this issue to the > degree Lucene is, I think it would be good to go ahead and get a > policy in place becase it's probably only a matter of time until > OpenNLP is. > > Thanks, > Jeff > > [1] https://github.com/apache/lucene/blob/main/AI_POLICY.md > > On Thu, Sep 17, 2026 at 5:53 AM Richard Zowalla <[email protected]> wrote: > > > > Hi all, > > > > Over the last few weeks the number and size of open pull requests has > grown far beyond what we can review, and I think we all share part of the > blame. > > > > So I'd like to agree on a few practices for everyone, contributors and > committers alike. > > > > 1.) Why this matters > > > > Every change we merge has to be read and understood by a human > committer, who then takes responsibility for it. That's how Apache works, > and a green build or a machine-generated review doesn't replace it. I'm > including myself here: I've posted LLM-generated reviews on some PRs when I > was short on time. They can help, but they aren't a review, and we > shouldn't pretend they are. > > > > This has also become a problem for the community itself. Some people > have told me privately that they can't handle the review load anymore, and > some have stopped reviewing completely. That's the worst outcome for a > volunteer project. Reviewers are our scarcest resource, and if we burn them > out, nothing gets merged, no matter how good the code is. I don't want > anyone to feel they have to step back because the queue has become > impossible to keep up with. > > > > Right now there are about 24 open PRs [1] adding about 246k lines. Much > of that code is duplicated across PRs, and nobody can tell which part of a > diff belongs to which change; description or comments are outdated and a > human is totally lost. Some examples, only to show the pattern: > > > > - Stacked PRs that all target main repeat their parents' diffs: #1213, > #1214 and #1215 each contain all of #1152 (+30k to +40k lines, 138–165 > commits each) [2], and #1276–#1282 all contain #1275 [3]. > > - The same code is in several PRs at once (the TextEmbedder SPI is in > #1290 and #1152) [4], so review comments on one copy never reach the others. > > - Descriptions no longer match the diff, for example PRs that list > dependencies which were merged long ago, or file counts and scope that have > since changed [5]. Several of us have left comments saying we simply don't > understand what a PR is about. > > - Some PRs have 100+ commits, including merge commits from main. > > - New scope gets added to PRs while they are under review, and new > drafts keep opening before the existing ones are done. > > > > None of this says the work is bad. Much of it fixes real bugs and adds > useful features. But a PR that a volunteer can't review in one sitting > doesn't get a real review. > > It either sits there (forever) or gets merged on trust, and neither is > acceptable > > > > 2.) Proposal > > > > A. One PR, one JIRA, one unit a human can review. > > > > Squash to a single commit, or to a few commits that each make sense on > their own. Rebase on main instead of merging main in. The title and > description should say what the PR does *now*: rewrite them when the scope > changes rather than adding "review round" sections, since the history is > already in the comments. > > > > B. No stacking. > > > > A PR targets main and contains only its own change. If it depends on > another open PR, it waits on the contributor's fork until that parent is > merged. A PR should never show the diff of another open PR. > > > > C. Optional modules go to opennlp-addons. > > > > New optional modules with their own data, dependencies or follow-up > plans belong in opennlp-addons [6]. The current embeddings, vector index, > gazetteer and wordnet drafts are examples [7]. Core only gets the small API > contracts those modules really need, each in its own focused PR. > > > > D. Finish before starting something new. > > > > Limit how many PRs one person has ready for review at a time (3–4, say). > Don't add new scope to a PR that is under review; put new ideas in JIRA > first. > > The speed at which we can review sets how fast code gets in, not the > speed at which code gets written. Five finished PRs are worth more to us > than 24 open ones. > > > > E. Committers, too. > > > > We should review in time and reply clearly: approve, request specific > changes, or say it doesn't fit and close it. We shouldn't post generated > reviews as if they were our own. If we use tools, we check and sign off on > every point we post. And it's fine to say "this is too big for me to > review", which is useful feedback, not a failure. > > > > > > If we agree, we should add this to the contribution guidelines and the > PR template. > > For the PRs already open, I suggest we agree on a merge order together, > and close or park stacked and duplicated ones until their turn comes. > > > > And to those who stepped back: thank you for everything you've reviewed > so far. I hope this makes it possible for you to come back. > > > > Thoughts? > > > > Gruß > > Richard > > > > --- > > > > ## References > > > > **[1] Open pull requests** > > - https://github.com/apache/opennlp/pulls > > > > **[2] Static embeddings stack** > > - https://github.com/apache/opennlp/pull/1152 ([OPENNLP-1877]( > https://issues.apache.org/jira/browse/OPENNLP-1877)) > > - https://github.com/apache/opennlp/pull/1213 ([OPENNLP-1895]( > https://issues.apache.org/jira/browse/OPENNLP-1895)) > > - https://github.com/apache/opennlp/pull/1214 ([OPENNLP-1910]( > https://issues.apache.org/jira/browse/OPENNLP-1910)) > > - https://github.com/apache/opennlp/pull/1215 ([OPENNLP-1911]( > https://issues.apache.org/jira/browse/OPENNLP-1911)) > > > > **[3] Regex removal** > > - https://github.com/apache/opennlp/pull/1275 ([OPENNLP-1928]( > https://issues.apache.org/jira/browse/OPENNLP-1928)) > > - https://github.com/apache/opennlp/pull/1276 ([OPENNLP-1930]( > https://issues.apache.org/jira/browse/OPENNLP-1930)) > > - https://github.com/apache/opennlp/pull/1277 ([OPENNLP-1931]( > https://issues.apache.org/jira/browse/OPENNLP-1931)) > > - https://github.com/apache/opennlp/pull/1278 ([OPENNLP-1932]( > https://issues.apache.org/jira/browse/OPENNLP-1932)) > > - https://github.com/apache/opennlp/pull/1279 ([OPENNLP-1933]( > https://issues.apache.org/jira/browse/OPENNLP-1933)) > > - https://github.com/apache/opennlp/pull/1280 ([OPENNLP-1929]( > https://issues.apache.org/jira/browse/OPENNLP-1929)) > > - https://github.com/apache/opennlp/pull/1281 ([OPENNLP-1934]( > https://issues.apache.org/jira/browse/OPENNLP-1934)) > > - https://github.com/apache/opennlp/pull/1282 ([OPENNLP-1935]( > https://issues.apache.org/jira/browse/OPENNLP-1935)) > > > > **[4] TextEmbedder SPI and shared deep-learning encoder tests** > > - https://github.com/apache/opennlp/pull/1290 ([OPENNLP-1937]( > https://issues.apache.org/jira/browse/OPENNLP-1937)) > > - https://github.com/apache/opennlp/pull/1288 ([OPENNLP-1942]( > https://issues.apache.org/jira/browse/OPENNLP-1942)) > > > > **[5] Dependency parser stack** > > - https://github.com/apache/opennlp/pull/1236 ([OPENNLP-547]( > https://issues.apache.org/jira/browse/OPENNLP-547)) > > - https://github.com/apache/opennlp/pull/1237 ([OPENNLP-1919]( > https://issues.apache.org/jira/browse/OPENNLP-1919)) > > - https://github.com/apache/opennlp/pull/1238 ([OPENNLP-1920]( > https://issues.apache.org/jira/browse/OPENNLP-1920)) > > > > **[6] opennlp-addons, with an example add-on PR** > > - https://github.com/apache/opennlp-addons > > - https://github.com/apache/opennlp-addons/pull/184 ([OPENNLP-1885]( > https://issues.apache.org/jira/browse/OPENNLP-1885)) > > > > **[7] Drafts that belong in opennlp-addons** > > - https://github.com/apache/opennlp/pull/1154 ([OPENNLP-1879]( > https://issues.apache.org/jira/browse/OPENNLP-1879)) > > - https://github.com/apache/opennlp/pull/1155 ([OPENNLP-1880]( > https://issues.apache.org/jira/browse/OPENNLP-1880)) > > - https://github.com/apache/opennlp/pull/1167 ([OPENNLP-1887]( > https://issues.apache.org/jira/browse/OPENNLP-1887)) > > - plus the embeddings stack in [2] >
