Sounds good. Agree on most but the stacks...

   - *Hand written guideline tutorial - *I'd volunteer to write up a good
   sorta "tutorial" that goes over "dos" and "do nots" of how you should push
   a PR.
   - *I disagree with not having stacks for complex features, but try to
   minimize the noise and rebase* -Stacks are normal in open source
   projects for larger features.  I think anything on a downstream stack
   should remain in draft.  Github even has a stack feature now. That being
   said, I'm fine with not doing a stack because it's not hard to have it
   otherwise.
   - *Drafts are OK if they start *green but are later converted to draft.
   Otherwise the work should be on a collaboration branch without being in
   draft. * Please let's view drafts as not reviewable* - as the button to
   turn green says "ready for review?" will get users regularly confused about
   the deviation on github.
   - *Limit PRs open - *Unless the team agrees on the number of PRs (i.e.
   the anti-regex push), we should try to limit the amount.  I promised that
   after the current stack is done, I'll limit my PRs to 0-2 with a primary
   focus on bug fixing, backporting, and features to be done collaboratively/
   discussed on.
   - *Discussions on github - *I hope we can bring more discussions to
   github - as design and collaboration have a lot of features you don't get
   in email (i.e. mermaid arch diagrams).


Thanks for bringing it to the thread.




On Thu, Sep 17, 2026 at 9:07 AM Jeff Zemerick <[email protected]> wrote:

> Good points. I agree with putting all of these points in the
> contributor guidelines, and I especially like the expectations for
> committers in the last section. Ultimately, any policy affects the
> committers the most.
>
> It's reassuring to see that OpenNLP is not the only ASF (and I'm sure
> non-ASF) project adapting to the new ways of working. Lucene has
> recently adopted an AI Policy [1] over using AI-generated responses
> for communication. I would favor the same or a similar policy for
> OpenNLP, too. While I don't think OpenNLP is facing this issue to the
> degree Lucene is, I think it would be good to go ahead and get a
> policy in place becase it's probably only a matter of time until
> OpenNLP is.
>
> Thanks,
> Jeff
>
> [1] https://github.com/apache/lucene/blob/main/AI_POLICY.md
>
> On Thu, Sep 17, 2026 at 5:53 AM Richard Zowalla <[email protected]> wrote:
> >
> > Hi all,
> >
> > Over the last few weeks the number and size of open pull requests has
> grown far beyond what we can review, and I think we all share part of the
> blame.
> >
> > So I'd like to agree on a few practices for everyone, contributors and
> committers alike.
> >
> > 1.) Why this matters
> >
> > Every change we merge has to be read and understood by a human
> committer, who then takes responsibility for it. That's how Apache works,
> and a green build or a machine-generated review doesn't replace it. I'm
> including myself here: I've posted LLM-generated reviews on some PRs when I
> was short on time. They can help, but they aren't a review, and we
> shouldn't pretend they are.
> >
> > This has also become a problem for the community itself. Some people
> have told me privately that they can't handle the review load anymore, and
> some have stopped reviewing completely. That's the worst outcome for a
> volunteer project. Reviewers are our scarcest resource, and if we burn them
> out, nothing gets merged, no matter how good the code is. I don't want
> anyone to feel they have to step back because the queue has become
> impossible to keep up with.
> >
> > Right now there are about 24 open PRs [1] adding about 246k lines. Much
> of that code is duplicated across PRs, and nobody can tell which part of a
> diff belongs to which change; description or comments are outdated and a
> human is totally lost. Some examples, only to show the pattern:
> >
> > - Stacked PRs that all target main repeat their parents' diffs: #1213,
> #1214 and #1215 each contain all of #1152 (+30k to +40k lines, 138–165
> commits each) [2], and #1276–#1282 all contain #1275 [3].
> > - The same code is in several PRs at once (the TextEmbedder SPI is in
> #1290 and #1152) [4], so review comments on one copy never reach the others.
> > - Descriptions no longer match the diff, for example PRs that list
> dependencies which were merged long ago, or file counts and scope that have
> since changed [5]. Several of us have left comments saying we simply don't
> understand what a PR is about.
> > - Some PRs have 100+ commits, including merge commits from main.
> > - New scope gets added to PRs while they are under review, and new
> drafts keep opening before the existing ones are done.
> >
> > None of this says the work is bad. Much of it fixes real bugs and adds
> useful features. But a PR that a volunteer can't review in one sitting
> doesn't get a real review.
> > It either sits there (forever) or gets merged on trust, and neither is
> acceptable
> >
> > 2.) Proposal
> >
> > A. One PR, one JIRA, one unit a human can review.
> >
> > Squash to a single commit, or to a few commits that each make sense on
> their own. Rebase on main instead of merging main in. The title and
> description should say what the PR does *now*: rewrite them when the scope
> changes rather than adding "review round" sections, since the history is
> already in the comments.
> >
> > B. No stacking.
> >
> > A PR targets main and contains only its own change. If it depends on
> another open PR, it waits on the contributor's fork until that parent is
> merged. A PR should never show the diff of another open PR.
> >
> > C. Optional modules go to opennlp-addons.
> >
> > New optional modules with their own data, dependencies or follow-up
> plans belong in opennlp-addons [6]. The current embeddings, vector index,
> gazetteer and wordnet drafts are examples [7]. Core only gets the small API
> contracts those modules really need, each in its own focused PR.
> >
> > D. Finish before starting something new.
> >
> > Limit how many PRs one person has ready for review at a time (3–4, say).
> Don't add new scope to a PR that is under review; put new ideas in JIRA
> first.
> > The speed at which we can review sets how fast code gets in, not the
> speed at which code gets written. Five finished PRs are worth more to us
> than 24 open ones.
> >
> > E. Committers, too.
> >
> > We should review in time and reply clearly: approve, request specific
> changes, or say it doesn't fit and close it. We shouldn't post generated
> reviews as if they were our own. If we use tools, we check and sign off on
> every point we post. And it's fine to say "this is too big for me to
> review", which is useful feedback, not a failure.
> >
> >
> > If we agree, we should add this to the contribution guidelines and the
> PR template.
> > For the PRs already open, I suggest we agree on a merge order together,
> and close or park stacked and duplicated ones until their turn comes.
> >
> > And to those who stepped back: thank you for everything you've reviewed
> so far. I hope this makes it possible for you to come back.
> >
> > Thoughts?
> >
> > Gruß
> > Richard
> >
> > ---
> >
> > ## References
> >
> > **[1] Open pull requests**
> > - https://github.com/apache/opennlp/pulls
> >
> > **[2] Static embeddings stack**
> > - https://github.com/apache/opennlp/pull/1152 ([OPENNLP-1877](
> https://issues.apache.org/jira/browse/OPENNLP-1877))
> > - https://github.com/apache/opennlp/pull/1213 ([OPENNLP-1895](
> https://issues.apache.org/jira/browse/OPENNLP-1895))
> > - https://github.com/apache/opennlp/pull/1214 ([OPENNLP-1910](
> https://issues.apache.org/jira/browse/OPENNLP-1910))
> > - https://github.com/apache/opennlp/pull/1215 ([OPENNLP-1911](
> https://issues.apache.org/jira/browse/OPENNLP-1911))
> >
> > **[3] Regex removal**
> > - https://github.com/apache/opennlp/pull/1275 ([OPENNLP-1928](
> https://issues.apache.org/jira/browse/OPENNLP-1928))
> > - https://github.com/apache/opennlp/pull/1276 ([OPENNLP-1930](
> https://issues.apache.org/jira/browse/OPENNLP-1930))
> > - https://github.com/apache/opennlp/pull/1277 ([OPENNLP-1931](
> https://issues.apache.org/jira/browse/OPENNLP-1931))
> > - https://github.com/apache/opennlp/pull/1278 ([OPENNLP-1932](
> https://issues.apache.org/jira/browse/OPENNLP-1932))
> > - https://github.com/apache/opennlp/pull/1279 ([OPENNLP-1933](
> https://issues.apache.org/jira/browse/OPENNLP-1933))
> > - https://github.com/apache/opennlp/pull/1280 ([OPENNLP-1929](
> https://issues.apache.org/jira/browse/OPENNLP-1929))
> > - https://github.com/apache/opennlp/pull/1281 ([OPENNLP-1934](
> https://issues.apache.org/jira/browse/OPENNLP-1934))
> > - https://github.com/apache/opennlp/pull/1282 ([OPENNLP-1935](
> https://issues.apache.org/jira/browse/OPENNLP-1935))
> >
> > **[4] TextEmbedder SPI and shared deep-learning encoder tests**
> > - https://github.com/apache/opennlp/pull/1290 ([OPENNLP-1937](
> https://issues.apache.org/jira/browse/OPENNLP-1937))
> > - https://github.com/apache/opennlp/pull/1288 ([OPENNLP-1942](
> https://issues.apache.org/jira/browse/OPENNLP-1942))
> >
> > **[5] Dependency parser stack**
> > - https://github.com/apache/opennlp/pull/1236 ([OPENNLP-547](
> https://issues.apache.org/jira/browse/OPENNLP-547))
> > - https://github.com/apache/opennlp/pull/1237 ([OPENNLP-1919](
> https://issues.apache.org/jira/browse/OPENNLP-1919))
> > - https://github.com/apache/opennlp/pull/1238 ([OPENNLP-1920](
> https://issues.apache.org/jira/browse/OPENNLP-1920))
> >
> > **[6] opennlp-addons, with an example add-on PR**
> > - https://github.com/apache/opennlp-addons
> > - https://github.com/apache/opennlp-addons/pull/184 ([OPENNLP-1885](
> https://issues.apache.org/jira/browse/OPENNLP-1885))
> >
> > **[7] Drafts that belong in opennlp-addons**
> > - https://github.com/apache/opennlp/pull/1154 ([OPENNLP-1879](
> https://issues.apache.org/jira/browse/OPENNLP-1879))
> > - https://github.com/apache/opennlp/pull/1155 ([OPENNLP-1880](
> https://issues.apache.org/jira/browse/OPENNLP-1880))
> > - https://github.com/apache/opennlp/pull/1167 ([OPENNLP-1887](
> https://issues.apache.org/jira/browse/OPENNLP-1887))
> > - plus the embeddings stack in [2]
>

Reply via email to