It's become standard to have AGENTS.md in an open source project serve as
the "README for agents".

When we release 3.0, we should have one.

Since we were on that subject - anyone object?


On Tue, Sep 22, 2026 at 1:11 PM Richard Zowalla <[email protected]> wrote:

> Hi all,
>
> thanks for the feedback so far. I'd like to follow up on two points.
>
> 1.) Stacked PRs
>
> My "no stacking" in B was too strict as written. I'm fine with stacked PRs
> if they're a deliberate way to break a larger issue into manageable pieces.
> What I object to is stacks where nobody can tell anymore what a single PR
> is about. So I'd rephrase B roughly as:
>
> Each PR in a stack has a concise, self-contained intent and its own issue.
> The initial description is clear to a human reviewer: it says that the PR
> is part of a stack, what it depends on, and what it adds on top.
> The stack stays stable. No new scope while it's under review, and parents
> get reviewed and merged first.
> If a reviewer has to dig through the whole stack to understand one PR, it
> has to be broken down differently.
>
> 2.) GitHub Discussions
>
> GitHub Discussions can be a useful additional channel, similar to users@,
> but it can't and shouldn't replace the mailing list. Mailing lists are a
> core value of the ASF. They're independent of any single vendor, they work
> asynchronously across time zones, they're accessible to everyone, and
> they're archived for the long term. dev@ remains the primary place for
> technical discussions and the sole authority for decisions.
>
> It's fine to gather additional feedback in a GitHub discussion or a GitHub
> issue. But the outcome has to be brought back to the list, and any decision
> is made here. Or, strictly ASF speaking: if it didn't happen on the list,
> it didn't happen.
>
> 3.) Lucene AI policy
>
> I've read the Lucene AI policy and I think it's fine, and a good basis for
> us too. My main concern is simple: I don't want to talk with an AI. When I
> review a PR or reply to a comment, I want to know that there's a human on
> the other side who understands the change and stands behind it. That's the
> core value for me, and the Lucene policy covers it well. Open source is
> about collaboration between people, and tools don't change that.
>
> Gruß
> Richard
>
>
> > Am 22.09.2026 um 18:07 schrieb Kristian Rickert <[email protected]>:
> >
> > Sounds good. Agree on most but the stacks...
> >
> >   - *Hand written guideline tutorial - *I'd volunteer to write up a good
> >   sorta "tutorial" that goes over "dos" and "do nots" of how you should
> push
> >   a PR.
> >   - *I disagree with not having stacks for complex features, but try to
> >   minimize the noise and rebase* -Stacks are normal in open source
> >   projects for larger features.  I think anything on a downstream stack
> >   should remain in draft.  Github even has a stack feature now. That
> being
> >   said, I'm fine with not doing a stack because it's not hard to have it
> >   otherwise.
> >   - *Drafts are OK if they start *green but are later converted to draft.
> >   Otherwise the work should be on a collaboration branch without being in
> >   draft. * Please let's view drafts as not reviewable* - as the button to
> >   turn green says "ready for review?" will get users regularly confused
> about
> >   the deviation on github.
> >   - *Limit PRs open - *Unless the team agrees on the number of PRs (i.e.
> >   the anti-regex push), we should try to limit the amount.  I promised
> that
> >   after the current stack is done, I'll limit my PRs to 0-2 with a
> primary
> >   focus on bug fixing, backporting, and features to be done
> collaboratively/
> >   discussed on.
> >   - *Discussions on github - *I hope we can bring more discussions to
> >   github - as design and collaboration have a lot of features you don't
> get
> >   in email (i.e. mermaid arch diagrams).
> >
> >
> > Thanks for bringing it to the thread.
> >
> >
> >
> >
> > On Thu, Sep 17, 2026 at 9:07 AM Jeff Zemerick <[email protected]>
> wrote:
> >
> >> Good points. I agree with putting all of these points in the
> >> contributor guidelines, and I especially like the expectations for
> >> committers in the last section. Ultimately, any policy affects the
> >> committers the most.
> >>
> >> It's reassuring to see that OpenNLP is not the only ASF (and I'm sure
> >> non-ASF) project adapting to the new ways of working. Lucene has
> >> recently adopted an AI Policy [1] over using AI-generated responses
> >> for communication. I would favor the same or a similar policy for
> >> OpenNLP, too. While I don't think OpenNLP is facing this issue to the
> >> degree Lucene is, I think it would be good to go ahead and get a
> >> policy in place becase it's probably only a matter of time until
> >> OpenNLP is.
> >>
> >> Thanks,
> >> Jeff
> >>
> >> [1] https://github.com/apache/lucene/blob/main/AI_POLICY.md
> >>
> >> On Thu, Sep 17, 2026 at 5:53 AM Richard Zowalla <[email protected]>
> wrote:
> >>>
> >>> Hi all,
> >>>
> >>> Over the last few weeks the number and size of open pull requests has
> >> grown far beyond what we can review, and I think we all share part of
> the
> >> blame.
> >>>
> >>> So I'd like to agree on a few practices for everyone, contributors and
> >> committers alike.
> >>>
> >>> 1.) Why this matters
> >>>
> >>> Every change we merge has to be read and understood by a human
> >> committer, who then takes responsibility for it. That's how Apache
> works,
> >> and a green build or a machine-generated review doesn't replace it. I'm
> >> including myself here: I've posted LLM-generated reviews on some PRs
> when I
> >> was short on time. They can help, but they aren't a review, and we
> >> shouldn't pretend they are.
> >>>
> >>> This has also become a problem for the community itself. Some people
> >> have told me privately that they can't handle the review load anymore,
> and
> >> some have stopped reviewing completely. That's the worst outcome for a
> >> volunteer project. Reviewers are our scarcest resource, and if we burn
> them
> >> out, nothing gets merged, no matter how good the code is. I don't want
> >> anyone to feel they have to step back because the queue has become
> >> impossible to keep up with.
> >>>
> >>> Right now there are about 24 open PRs [1] adding about 246k lines. Much
> >> of that code is duplicated across PRs, and nobody can tell which part
> of a
> >> diff belongs to which change; description or comments are outdated and a
> >> human is totally lost. Some examples, only to show the pattern:
> >>>
> >>> - Stacked PRs that all target main repeat their parents' diffs: #1213,
> >> #1214 and #1215 each contain all of #1152 (+30k to +40k lines, 138–165
> >> commits each) [2], and #1276–#1282 all contain #1275 [3].
> >>> - The same code is in several PRs at once (the TextEmbedder SPI is in
> >> #1290 and #1152) [4], so review comments on one copy never reach the
> others.
> >>> - Descriptions no longer match the diff, for example PRs that list
> >> dependencies which were merged long ago, or file counts and scope that
> have
> >> since changed [5]. Several of us have left comments saying we simply
> don't
> >> understand what a PR is about.
> >>> - Some PRs have 100+ commits, including merge commits from main.
> >>> - New scope gets added to PRs while they are under review, and new
> >> drafts keep opening before the existing ones are done.
> >>>
> >>> None of this says the work is bad. Much of it fixes real bugs and adds
> >> useful features. But a PR that a volunteer can't review in one sitting
> >> doesn't get a real review.
> >>> It either sits there (forever) or gets merged on trust, and neither is
> >> acceptable
> >>>
> >>> 2.) Proposal
> >>>
> >>> A. One PR, one JIRA, one unit a human can review.
> >>>
> >>> Squash to a single commit, or to a few commits that each make sense on
> >> their own. Rebase on main instead of merging main in. The title and
> >> description should say what the PR does *now*: rewrite them when the
> scope
> >> changes rather than adding "review round" sections, since the history is
> >> already in the comments.
> >>>
> >>> B. No stacking.
> >>>
> >>> A PR targets main and contains only its own change. If it depends on
> >> another open PR, it waits on the contributor's fork until that parent is
> >> merged. A PR should never show the diff of another open PR.
> >>>
> >>> C. Optional modules go to opennlp-addons.
> >>>
> >>> New optional modules with their own data, dependencies or follow-up
> >> plans belong in opennlp-addons [6]. The current embeddings, vector
> index,
> >> gazetteer and wordnet drafts are examples [7]. Core only gets the small
> API
> >> contracts those modules really need, each in its own focused PR.
> >>>
> >>> D. Finish before starting something new.
> >>>
> >>> Limit how many PRs one person has ready for review at a time (3–4,
> say).
> >> Don't add new scope to a PR that is under review; put new ideas in JIRA
> >> first.
> >>> The speed at which we can review sets how fast code gets in, not the
> >> speed at which code gets written. Five finished PRs are worth more to us
> >> than 24 open ones.
> >>>
> >>> E. Committers, too.
> >>>
> >>> We should review in time and reply clearly: approve, request specific
> >> changes, or say it doesn't fit and close it. We shouldn't post generated
> >> reviews as if they were our own. If we use tools, we check and sign off
> on
> >> every point we post. And it's fine to say "this is too big for me to
> >> review", which is useful feedback, not a failure.
> >>>
> >>>
> >>> If we agree, we should add this to the contribution guidelines and the
> >> PR template.
> >>> For the PRs already open, I suggest we agree on a merge order together,
> >> and close or park stacked and duplicated ones until their turn comes.
> >>>
> >>> And to those who stepped back: thank you for everything you've reviewed
> >> so far. I hope this makes it possible for you to come back.
> >>>
> >>> Thoughts?
> >>>
> >>> Gruß
> >>> Richard
> >>>
> >>> ---
> >>>
> >>> ## References
> >>>
> >>> **[1] Open pull requests**
> >>> - https://github.com/apache/opennlp/pulls
> >>>
> >>> **[2] Static embeddings stack**
> >>> - https://github.com/apache/opennlp/pull/1152 ([OPENNLP-1877](
> >> https://issues.apache.org/jira/browse/OPENNLP-1877))
> >>> - https://github.com/apache/opennlp/pull/1213 ([OPENNLP-1895](
> >> https://issues.apache.org/jira/browse/OPENNLP-1895))
> >>> - https://github.com/apache/opennlp/pull/1214 ([OPENNLP-1910](
> >> https://issues.apache.org/jira/browse/OPENNLP-1910))
> >>> - https://github.com/apache/opennlp/pull/1215 ([OPENNLP-1911](
> >> https://issues.apache.org/jira/browse/OPENNLP-1911))
> >>>
> >>> **[3] Regex removal**
> >>> - https://github.com/apache/opennlp/pull/1275 ([OPENNLP-1928](
> >> https://issues.apache.org/jira/browse/OPENNLP-1928))
> >>> - https://github.com/apache/opennlp/pull/1276 ([OPENNLP-1930](
> >> https://issues.apache.org/jira/browse/OPENNLP-1930))
> >>> - https://github.com/apache/opennlp/pull/1277 ([OPENNLP-1931](
> >> https://issues.apache.org/jira/browse/OPENNLP-1931))
> >>> - https://github.com/apache/opennlp/pull/1278 ([OPENNLP-1932](
> >> https://issues.apache.org/jira/browse/OPENNLP-1932))
> >>> - https://github.com/apache/opennlp/pull/1279 ([OPENNLP-1933](
> >> https://issues.apache.org/jira/browse/OPENNLP-1933))
> >>> - https://github.com/apache/opennlp/pull/1280 ([OPENNLP-1929](
> >> https://issues.apache.org/jira/browse/OPENNLP-1929))
> >>> - https://github.com/apache/opennlp/pull/1281 ([OPENNLP-1934](
> >> https://issues.apache.org/jira/browse/OPENNLP-1934))
> >>> - https://github.com/apache/opennlp/pull/1282 ([OPENNLP-1935](
> >> https://issues.apache.org/jira/browse/OPENNLP-1935))
> >>>
> >>> **[4] TextEmbedder SPI and shared deep-learning encoder tests**
> >>> - https://github.com/apache/opennlp/pull/1290 ([OPENNLP-1937](
> >> https://issues.apache.org/jira/browse/OPENNLP-1937))
> >>> - https://github.com/apache/opennlp/pull/1288 ([OPENNLP-1942](
> >> https://issues.apache.org/jira/browse/OPENNLP-1942))
> >>>
> >>> **[5] Dependency parser stack**
> >>> - https://github.com/apache/opennlp/pull/1236 ([OPENNLP-547](
> >> https://issues.apache.org/jira/browse/OPENNLP-547))
> >>> - https://github.com/apache/opennlp/pull/1237 ([OPENNLP-1919](
> >> https://issues.apache.org/jira/browse/OPENNLP-1919))
> >>> - https://github.com/apache/opennlp/pull/1238 ([OPENNLP-1920](
> >> https://issues.apache.org/jira/browse/OPENNLP-1920))
> >>>
> >>> **[6] opennlp-addons, with an example add-on PR**
> >>> - https://github.com/apache/opennlp-addons
> >>> - https://github.com/apache/opennlp-addons/pull/184 ([OPENNLP-1885](
> >> https://issues.apache.org/jira/browse/OPENNLP-1885))
> >>>
> >>> **[7] Drafts that belong in opennlp-addons**
> >>> - https://github.com/apache/opennlp/pull/1154 ([OPENNLP-1879](
> >> https://issues.apache.org/jira/browse/OPENNLP-1879))
> >>> - https://github.com/apache/opennlp/pull/1155 ([OPENNLP-1880](
> >> https://issues.apache.org/jira/browse/OPENNLP-1880))
> >>> - https://github.com/apache/opennlp/pull/1167 ([OPENNLP-1887](
> >> https://issues.apache.org/jira/browse/OPENNLP-1887))
> >>> - plus the embeddings stack in [2]
> >>
>
>

Reply via email to