i think simon jakobi's proposal is not installing some label-triggered
auto-review bot on our gitlab instance, it's about him using his own
account to paste llm-generated reviews to patches where authors explicitly
request so.

thanks simon, it's indeed a generous offer, but i don't think the gitlab
instance is the right place for pasting large blobs of llm-generated
reviews. imho it's better to move these to e.g. a self-hosted gerrit
instance, or a discord/matrix channel etc, where people who want this
service can ping you, get a piece of review to take time to digest and
improve their patch accordingly.

i personally find llm-generated reviews very useful in my own work, but the
signal-to-noise ratio of review outputs by even the best models/harnesses
today is simply low, the insight needs to be mined from the text, and this
can be a huge distraction and prevent a human reviewer from entering the
flow when looking at patches.

cheers,
cheng

On Mon, Jul 27, 2026 at 9:40 PM Zubin Duggal via ghc-devs <
[email protected]> wrote:

> I am strongly opposed to this. I also don't see the value in
> it. If you want to ask an LLM to review your code, you can
> probably do this much quicker by just asking one. Today you
> are never more than a few clicks away from a text box that
> will put you in direct contact with one.
>
> If you want to maintain a shared repository of common prompts,
> skills, guidance etc for LLMs regarding work on GHC, I have
> no objection to this (I would prefer to only have documentation
> and prose targetting humans in the main GHC repository though).
>
> In fact I would like to propose a policy in the exact opposite
> direction:
>
> Disallow _purely autonomous_ systems from posting on the GHC
> Gitlab, unless such a system is cleared in advance by the
> administrators of the instance and/or contributors to GHC.
>
> What this won't disallow:
>
> * Getting an LLM to isolate and create a minimal reproducer
> or bug description and posting this.
>
> * Getting an LLM to propose a patch or fix for a particular
> issue
>
> * Asking an LLM about the details of how some particular
> external system works, and then posting the results.
>
> * Asking an LLM to catalogue a bunch of related issues and
> extract a common pattern
>
> * etc.
>
> What this will disallow:
>
> Autonomous systems without humans in the loop writing and
> interacting with the gitlab.
>
> My reasons
>
> 1. Code review is a tool for building a shared understanding
> and a common mental model of the codebase and problem
> domain. An LLM model cannot do that, or even if you believe
> it can, this internal model is lost as soon as that
> particular session is garbage collected. At best you can try
> to reconstruct some approximation of this model later by
> feeding the same inputs into a fresh session on a probably
> different model.
>
> 2. I would find it a very sad if the GHC Gitlab became a
> place where bots talk to other bots who act upon and those
> comments to make addtional changes in response to this in
> an uncontrolled feedback loop. I predict that in the limit
> any such feedback loop becomes quickly detached from the
> original problem and devolves into a lot of noise totally
> divorced from any real human concerns.
>
> We can already see examples of how this goes wrong **today**
> Take a look at a prominent example of a codebase that went
> "all in" on mostly unsupervised autonomous LLM development,
> the javascript runtime Bun: https://github.com/oven-sh/bun/pulls
>
> It would kill any motivation I would have to contribute to
> GHC and collaborate with others if the merge request queue
> looked anything approximating that. I use the gitlab to
> talk to and collaborate with actual humans, who I hope will
> actually engage their brains and help get us to a better
> and more corret mental model of the space.
>
> If I wanted to talk to an LLM, I have countless other
> options for ways in which to do that.
>
> 3. Prompt injection. Gitlab comments are an untrusted
> input source and LLMs fundamentally cannot discriminate
> between instructions and data. You can imagine prompt
> injection attacks where hidden text in a gitlab comment,
> attached bug reproducer, link etc. gets the model to do
> various destructive things, like spamming the gitlab,
> editing/deleting your prior comments, mining bitcoin on the
> CI runners, attempting to inject malicious code into the
> compiler, testcases, reproducers etc.
>
> That doesn't mean that there isn't value for LLM assisted
> autonomous processes on the gitlab, given that they are
> desiged properly keeping in mind the concerns about human
> centeredness, resource usage and prompt injection.
>
> Here are some examples of LLM assisted system/bots that
> could actually be useful if audited and pre-approved with
> the community:
>
> * A bot that automatically bisects regressions and identifies
> the change responsible for it.
> In principle this could be done with a purely deterministic
> script, but often you run into problems with the toolchain
> or commits not being individually buildable etc that require
> some adjustements on the fly.
>
> * A bot that automatically minimises reproducers and posts
> the results after the testcase can been verified on a deterministic
> system
>
> On 26/07/27 19:38, Simon Jakobi via ghc-devs wrote:
> >Hi devs!
> >
> >Inspired by the recent discussion on the LLM policy, I'd like to try
> >introducing LLM-generated code reviews in GHC, as an opt-in service.
> >
> >My motivation:
> >
> >* Catch issues before human reviewers spend time on them.
> >* Reduce the number of bugs merged.
> >
> >
> >What this is — and what it isn't
> >--------------------------------
> >
> >This is tool output: requested on the MR, produced by a model, posted
> verbatim
> >and clearly marked as machine-generated. This is not a review in the
> sense of
> >draft LLM policy.
> >
> >I'm not vouching for any of it. Think of it as a linter the author could
> have
> >run themselves, except on my LLM budget. Nobody is required to act on
> >it even read it.
> >
> >I mention this explicitly because the policy asks reviewers to take full
> >responsibility for each line of their review. I can't do that here — it
> would
> >mean checking every finding myself, which defeats the purpose and would
> >restrict the service to areas of the compiler I already know well.
> >
> >
> >How to request one
> >------------------
> >
> >* Ping me on the MR with a comment like "@sjakobi llm-review please".
> >  * Feel free to request particular aspects the review should cover.
> >* Authors can request a review for their own MRs. Maintainers can request
> one
> >  on any MR.
> >* If I have enough usage left on my plan, I'll confirm that I'm on it.
> >* I'll then prompt the model, and post the result on the MR.
> >* Feedback on the quality of the review is of course very welcome.
> >
> >
> >What I promise
> >--------------
> >
> >* Every review is labelled as machine-generated, with model, effort level
> and
> >  the exact prompt used.
> >* Reviews are posted verbatim. I don't edit or curate them, so what you
> see is
> >  exactly what the model produced.
> >* Only the public MR diff and public repository context go to the model.
> >  Nothing else.
> >* I'll keep working on the prompts and the choice of model to improve the
> >  signal-to-noise ratio.
> >* If the signal-to-noise ratio stays inacceptably bad, we can simply
> >stop this service.
> >
> >
> >Review coverage
> >---------------
> >
> >I'll probably start with a fairly basic review prompt, primarily aimed at
> >correctness issues.
> >
> >Possible extensions:
> >
> >* Documentation consistency: check that documentation stays consistent,
> >  including Notes elsewhere that reference the changed code.
> >* Performance: check for performance issues — potentially including
> checking
> >  the -ddump-simpl output for the changed code for unnecessary
> >  allocation etc.
> >* Feel free to suggest anything else!
> >
> >
> >Format
> >------
> >
> >Initially a single comment per review. Once we reach a good
> signal-to-noise
> >ratio, we can consider inline comments on the diff.
> >
> >
> >Caveats
> >-------
> >
> >This is a volunteer service I intend to provide in my spare time. If I'm
> AFK
> >or on vacation, reviews will take longer. If I can't keep up with
> requests, I
> >will prioritise which MRs I run. I might take a break entirely.
> >
> >
> >Getting involved
> >----------------
> >
> >Is anyone interested in joining this effort? If so we could form a team
> >(@llm-reviewers?), and share prompting techniques, model choices, etc.
> >
> >If the service turns out to be genuinely useful, we can look into
> >automating it and – if necessary – possible funding.
> >
> >
> >Cheers,
> >Simon
> >_______________________________________________
> >ghc-devs mailing list -- [email protected]
> >To unsubscribe send an email to [email protected]
> _______________________________________________
> ghc-devs mailing list -- [email protected]
> To unsubscribe send an email to [email protected]
>
_______________________________________________
ghc-devs mailing list -- [email protected]
To unsubscribe send an email to [email protected]

Reply via email to