It's become standard to have AGENTS.md in an open source project serve as the "README for agents".
When we release 3.0, we should have one. Since we were on that subject - anyone object? On Tue, Sep 22, 2026 at 1:11 PM Richard Zowalla <[email protected]> wrote: > Hi all, > > thanks for the feedback so far. I'd like to follow up on two points. > > 1.) Stacked PRs > > My "no stacking" in B was too strict as written. I'm fine with stacked PRs > if they're a deliberate way to break a larger issue into manageable pieces. > What I object to is stacks where nobody can tell anymore what a single PR > is about. So I'd rephrase B roughly as: > > Each PR in a stack has a concise, self-contained intent and its own issue. > The initial description is clear to a human reviewer: it says that the PR > is part of a stack, what it depends on, and what it adds on top. > The stack stays stable. No new scope while it's under review, and parents > get reviewed and merged first. > If a reviewer has to dig through the whole stack to understand one PR, it > has to be broken down differently. > > 2.) GitHub Discussions > > GitHub Discussions can be a useful additional channel, similar to users@, > but it can't and shouldn't replace the mailing list. Mailing lists are a > core value of the ASF. They're independent of any single vendor, they work > asynchronously across time zones, they're accessible to everyone, and > they're archived for the long term. dev@ remains the primary place for > technical discussions and the sole authority for decisions. > > It's fine to gather additional feedback in a GitHub discussion or a GitHub > issue. But the outcome has to be brought back to the list, and any decision > is made here. Or, strictly ASF speaking: if it didn't happen on the list, > it didn't happen. > > 3.) Lucene AI policy > > I've read the Lucene AI policy and I think it's fine, and a good basis for > us too. My main concern is simple: I don't want to talk with an AI. When I > review a PR or reply to a comment, I want to know that there's a human on > the other side who understands the change and stands behind it. That's the > core value for me, and the Lucene policy covers it well. Open source is > about collaboration between people, and tools don't change that. > > Gruß > Richard > > > > Am 22.09.2026 um 18:07 schrieb Kristian Rickert <[email protected]>: > > > > Sounds good. Agree on most but the stacks... > > > > - *Hand written guideline tutorial - *I'd volunteer to write up a good > > sorta "tutorial" that goes over "dos" and "do nots" of how you should > push > > a PR. > > - *I disagree with not having stacks for complex features, but try to > > minimize the noise and rebase* -Stacks are normal in open source > > projects for larger features. I think anything on a downstream stack > > should remain in draft. Github even has a stack feature now. That > being > > said, I'm fine with not doing a stack because it's not hard to have it > > otherwise. > > - *Drafts are OK if they start *green but are later converted to draft. > > Otherwise the work should be on a collaboration branch without being in > > draft. * Please let's view drafts as not reviewable* - as the button to > > turn green says "ready for review?" will get users regularly confused > about > > the deviation on github. > > - *Limit PRs open - *Unless the team agrees on the number of PRs (i.e. > > the anti-regex push), we should try to limit the amount. I promised > that > > after the current stack is done, I'll limit my PRs to 0-2 with a > primary > > focus on bug fixing, backporting, and features to be done > collaboratively/ > > discussed on. > > - *Discussions on github - *I hope we can bring more discussions to > > github - as design and collaboration have a lot of features you don't > get > > in email (i.e. mermaid arch diagrams). > > > > > > Thanks for bringing it to the thread. > > > > > > > > > > On Thu, Sep 17, 2026 at 9:07 AM Jeff Zemerick <[email protected]> > wrote: > > > >> Good points. I agree with putting all of these points in the > >> contributor guidelines, and I especially like the expectations for > >> committers in the last section. Ultimately, any policy affects the > >> committers the most. > >> > >> It's reassuring to see that OpenNLP is not the only ASF (and I'm sure > >> non-ASF) project adapting to the new ways of working. Lucene has > >> recently adopted an AI Policy [1] over using AI-generated responses > >> for communication. I would favor the same or a similar policy for > >> OpenNLP, too. While I don't think OpenNLP is facing this issue to the > >> degree Lucene is, I think it would be good to go ahead and get a > >> policy in place becase it's probably only a matter of time until > >> OpenNLP is. > >> > >> Thanks, > >> Jeff > >> > >> [1] https://github.com/apache/lucene/blob/main/AI_POLICY.md > >> > >> On Thu, Sep 17, 2026 at 5:53 AM Richard Zowalla <[email protected]> > wrote: > >>> > >>> Hi all, > >>> > >>> Over the last few weeks the number and size of open pull requests has > >> grown far beyond what we can review, and I think we all share part of > the > >> blame. > >>> > >>> So I'd like to agree on a few practices for everyone, contributors and > >> committers alike. > >>> > >>> 1.) Why this matters > >>> > >>> Every change we merge has to be read and understood by a human > >> committer, who then takes responsibility for it. That's how Apache > works, > >> and a green build or a machine-generated review doesn't replace it. I'm > >> including myself here: I've posted LLM-generated reviews on some PRs > when I > >> was short on time. They can help, but they aren't a review, and we > >> shouldn't pretend they are. > >>> > >>> This has also become a problem for the community itself. Some people > >> have told me privately that they can't handle the review load anymore, > and > >> some have stopped reviewing completely. That's the worst outcome for a > >> volunteer project. Reviewers are our scarcest resource, and if we burn > them > >> out, nothing gets merged, no matter how good the code is. I don't want > >> anyone to feel they have to step back because the queue has become > >> impossible to keep up with. > >>> > >>> Right now there are about 24 open PRs [1] adding about 246k lines. Much > >> of that code is duplicated across PRs, and nobody can tell which part > of a > >> diff belongs to which change; description or comments are outdated and a > >> human is totally lost. Some examples, only to show the pattern: > >>> > >>> - Stacked PRs that all target main repeat their parents' diffs: #1213, > >> #1214 and #1215 each contain all of #1152 (+30k to +40k lines, 138–165 > >> commits each) [2], and #1276–#1282 all contain #1275 [3]. > >>> - The same code is in several PRs at once (the TextEmbedder SPI is in > >> #1290 and #1152) [4], so review comments on one copy never reach the > others. > >>> - Descriptions no longer match the diff, for example PRs that list > >> dependencies which were merged long ago, or file counts and scope that > have > >> since changed [5]. Several of us have left comments saying we simply > don't > >> understand what a PR is about. > >>> - Some PRs have 100+ commits, including merge commits from main. > >>> - New scope gets added to PRs while they are under review, and new > >> drafts keep opening before the existing ones are done. > >>> > >>> None of this says the work is bad. Much of it fixes real bugs and adds > >> useful features. But a PR that a volunteer can't review in one sitting > >> doesn't get a real review. > >>> It either sits there (forever) or gets merged on trust, and neither is > >> acceptable > >>> > >>> 2.) Proposal > >>> > >>> A. One PR, one JIRA, one unit a human can review. > >>> > >>> Squash to a single commit, or to a few commits that each make sense on > >> their own. Rebase on main instead of merging main in. The title and > >> description should say what the PR does *now*: rewrite them when the > scope > >> changes rather than adding "review round" sections, since the history is > >> already in the comments. > >>> > >>> B. No stacking. > >>> > >>> A PR targets main and contains only its own change. If it depends on > >> another open PR, it waits on the contributor's fork until that parent is > >> merged. A PR should never show the diff of another open PR. > >>> > >>> C. Optional modules go to opennlp-addons. > >>> > >>> New optional modules with their own data, dependencies or follow-up > >> plans belong in opennlp-addons [6]. The current embeddings, vector > index, > >> gazetteer and wordnet drafts are examples [7]. Core only gets the small > API > >> contracts those modules really need, each in its own focused PR. > >>> > >>> D. Finish before starting something new. > >>> > >>> Limit how many PRs one person has ready for review at a time (3–4, > say). > >> Don't add new scope to a PR that is under review; put new ideas in JIRA > >> first. > >>> The speed at which we can review sets how fast code gets in, not the > >> speed at which code gets written. Five finished PRs are worth more to us > >> than 24 open ones. > >>> > >>> E. Committers, too. > >>> > >>> We should review in time and reply clearly: approve, request specific > >> changes, or say it doesn't fit and close it. We shouldn't post generated > >> reviews as if they were our own. If we use tools, we check and sign off > on > >> every point we post. And it's fine to say "this is too big for me to > >> review", which is useful feedback, not a failure. > >>> > >>> > >>> If we agree, we should add this to the contribution guidelines and the > >> PR template. > >>> For the PRs already open, I suggest we agree on a merge order together, > >> and close or park stacked and duplicated ones until their turn comes. > >>> > >>> And to those who stepped back: thank you for everything you've reviewed > >> so far. I hope this makes it possible for you to come back. > >>> > >>> Thoughts? > >>> > >>> Gruß > >>> Richard > >>> > >>> --- > >>> > >>> ## References > >>> > >>> **[1] Open pull requests** > >>> - https://github.com/apache/opennlp/pulls > >>> > >>> **[2] Static embeddings stack** > >>> - https://github.com/apache/opennlp/pull/1152 ([OPENNLP-1877]( > >> https://issues.apache.org/jira/browse/OPENNLP-1877)) > >>> - https://github.com/apache/opennlp/pull/1213 ([OPENNLP-1895]( > >> https://issues.apache.org/jira/browse/OPENNLP-1895)) > >>> - https://github.com/apache/opennlp/pull/1214 ([OPENNLP-1910]( > >> https://issues.apache.org/jira/browse/OPENNLP-1910)) > >>> - https://github.com/apache/opennlp/pull/1215 ([OPENNLP-1911]( > >> https://issues.apache.org/jira/browse/OPENNLP-1911)) > >>> > >>> **[3] Regex removal** > >>> - https://github.com/apache/opennlp/pull/1275 ([OPENNLP-1928]( > >> https://issues.apache.org/jira/browse/OPENNLP-1928)) > >>> - https://github.com/apache/opennlp/pull/1276 ([OPENNLP-1930]( > >> https://issues.apache.org/jira/browse/OPENNLP-1930)) > >>> - https://github.com/apache/opennlp/pull/1277 ([OPENNLP-1931]( > >> https://issues.apache.org/jira/browse/OPENNLP-1931)) > >>> - https://github.com/apache/opennlp/pull/1278 ([OPENNLP-1932]( > >> https://issues.apache.org/jira/browse/OPENNLP-1932)) > >>> - https://github.com/apache/opennlp/pull/1279 ([OPENNLP-1933]( > >> https://issues.apache.org/jira/browse/OPENNLP-1933)) > >>> - https://github.com/apache/opennlp/pull/1280 ([OPENNLP-1929]( > >> https://issues.apache.org/jira/browse/OPENNLP-1929)) > >>> - https://github.com/apache/opennlp/pull/1281 ([OPENNLP-1934]( > >> https://issues.apache.org/jira/browse/OPENNLP-1934)) > >>> - https://github.com/apache/opennlp/pull/1282 ([OPENNLP-1935]( > >> https://issues.apache.org/jira/browse/OPENNLP-1935)) > >>> > >>> **[4] TextEmbedder SPI and shared deep-learning encoder tests** > >>> - https://github.com/apache/opennlp/pull/1290 ([OPENNLP-1937]( > >> https://issues.apache.org/jira/browse/OPENNLP-1937)) > >>> - https://github.com/apache/opennlp/pull/1288 ([OPENNLP-1942]( > >> https://issues.apache.org/jira/browse/OPENNLP-1942)) > >>> > >>> **[5] Dependency parser stack** > >>> - https://github.com/apache/opennlp/pull/1236 ([OPENNLP-547]( > >> https://issues.apache.org/jira/browse/OPENNLP-547)) > >>> - https://github.com/apache/opennlp/pull/1237 ([OPENNLP-1919]( > >> https://issues.apache.org/jira/browse/OPENNLP-1919)) > >>> - https://github.com/apache/opennlp/pull/1238 ([OPENNLP-1920]( > >> https://issues.apache.org/jira/browse/OPENNLP-1920)) > >>> > >>> **[6] opennlp-addons, with an example add-on PR** > >>> - https://github.com/apache/opennlp-addons > >>> - https://github.com/apache/opennlp-addons/pull/184 ([OPENNLP-1885]( > >> https://issues.apache.org/jira/browse/OPENNLP-1885)) > >>> > >>> **[7] Drafts that belong in opennlp-addons** > >>> - https://github.com/apache/opennlp/pull/1154 ([OPENNLP-1879]( > >> https://issues.apache.org/jira/browse/OPENNLP-1879)) > >>> - https://github.com/apache/opennlp/pull/1155 ([OPENNLP-1880]( > >> https://issues.apache.org/jira/browse/OPENNLP-1880)) > >>> - https://github.com/apache/opennlp/pull/1167 ([OPENNLP-1887]( > >> https://issues.apache.org/jira/browse/OPENNLP-1887)) > >>> - plus the embeddings stack in [2] > >> > >
