Hi Andreas!

Regarding Sashiko, I think the README is quite informative:
https://github.com/sashiko-dev/sashiko#readme
If by "workflow" you mean how developers and maintainers interact with
it, it's apparently mostly via mailing lists:
https://github.com/sashiko-dev/sashiko/blob/main/MAINTAINERS_GUIDE.md.

https://lore.kernel.org/sashiko-reviews/ lists the (some?) review
threads. But the actual reviews are easier to view at
https://sashiko.dev/. (At the bottom of each review page, there's a
link "View Raw Log", where one can read the prompts and the internal
communication of the "subagents" producing the review.)

The claimed false positive rate around 20% sounds pretty good. Initial
findings have to pass through some adversarial filtering
(involving a "false positive guide" like
https://github.com/sashiko-dev/sashiko/blob/main/third_party/prompts/kernel/false-positive-guide.md).

But it's quite clear that a lot of work went into this system. To
adapt it for GHC would be a pretty large task. A mini version for GHC
might be more realistic. I think the largest task would be to create a
good false positive guide.


As for my "service", I may consider the side-channel option.

Cheers,
Simon

Am Mo., 27. Juli 2026 um 23:18 Uhr schrieb Andreas Klebinger via
ghc-devs <[email protected]>:
>
> The noise is mostly inherent to the models. There is no good way to
> "work on it to fix it" imo.
>
> They miss identify things as bugs due to missing context or simply a
> lack of understanding fairly frequently.
> For example I used Opus 5 to audit some GHC code for a specific kind of
> bug today. It flagged up:
>
> * A "bug" that can only happen if there are broken preconditions that
> are already checked by lint.
> * A "bug" that it flagged but marked safe as "only due to" a core
> invariant of GHCs register allocator design.
> * Four "Benign issues" that are all correct and pointless noise.
> * One real bug that I already knew about. (I was just doing a sanity check).
>
> Now I somewhat know my way around the NCG, so it cost me maybe 5 minuts
> to work through those false reports.
> But someone can easily spend an hour or much longer just working through
> those false positives.
>
> If someone who isn't very familiar with the NCG reads through only the
> LLM output but not the actual code their understanding
> of the NCG might actually be worse than before. Because they will
> remember hallucinated reasoning about how things work that
> is embedded in between those flagged issues.
>
> It can still be valuable to a reviewer that is aware of those
> circumstances. So I think people who want to try this should take
> you up on it. But unless one does non-trivial post processing in the
> prompt like spelling out invariants the model ignored,
> certain design decisions etc. that require domain knowledge the reviews
> are typically not worth putting on a MR.
>
> I do find your offer very generous! But for those reasons I think its
> not very useful to embed these artefacts directly on gitlab.
> As others mentioned I think performing this experiment via some side
> channel seems fine to me though.
>
> I would also be interested to hear what the workflow for the review tool
> the kernel uses is in case you have a good writeup readily
> at hand.
>
> Andreas
>
> On 27/07/2026 22:41, Simon Jakobi via ghc-devs wrote:
> > Hi Cheng,
> >
> > Thanks for your kind reply! :)
> >
> > Am Mo., 27. Juli 2026 um 22:18 Uhr schrieb Cheng Shao 
> > <[email protected]>:
> >> i think simon jakobi's proposal is not installing some label-triggered 
> >> auto-review bot on our gitlab instance, it's about him using his own 
> >> account to paste llm-generated reviews to patches where authors explicitly 
> >> request so.
> >>
> >> thanks simon, it's indeed a generous offer, but i don't think the gitlab 
> >> instance is the right place for pasting large blobs of llm-generated 
> >> reviews. imho it's better to move these to e.g. a self-hosted gerrit 
> >> instance, or a discord/matrix channel etc, where people who want this 
> >> service can ping you, get a piece of review to take time to digest and 
> >> improve their patch accordingly.
> > If the problem is with visual noise, I don't mind putting the blobs
> > into `<details>` tags so they don't take up as much space.
> >
> > However if the goal is to reduce maintainers' workload, I think the MR
> > is still the best place, so maintainers can use the presence of the
> > LLM-generated review as a signal.
> >
> > A decent amount of visibility would probably also help to improve the
> > LLM review system more quickly.
> >
> >> i personally find llm-generated reviews very useful in my own work, but 
> >> the signal-to-noise ratio of review outputs by even the best 
> >> models/harnesses today is simply low, the insight needs to be mined from 
> >> the text, and this can be a huge distraction and prevent a human reviewer 
> >> from entering the flow when looking at patches.
> > If the signal-to-noise ratio is indeed so bad, I'm ready to improve
> > the system until the ratio is more acceptable. It would be great to
> > have some collaborators in this, but I'll also give a try on my own.
> >
> > Cheers,
> > Simon
> >
> >> cheers,
> >> cheng
> >>
> >> On Mon, Jul 27, 2026 at 9:40 PM Zubin Duggal via ghc-devs 
> >> <[email protected]> wrote:
> >>> I am strongly opposed to this. I also don't see the value in
> >>> it. If you want to ask an LLM to review your code, you can
> >>> probably do this much quicker by just asking one. Today you
> >>> are never more than a few clicks away from a text box that
> >>> will put you in direct contact with one.
> >>>
> >>> If you want to maintain a shared repository of common prompts,
> >>> skills, guidance etc for LLMs regarding work on GHC, I have
> >>> no objection to this (I would prefer to only have documentation
> >>> and prose targetting humans in the main GHC repository though).
> >>>
> >>> In fact I would like to propose a policy in the exact opposite
> >>> direction:
> >>>
> >>> Disallow _purely autonomous_ systems from posting on the GHC
> >>> Gitlab, unless such a system is cleared in advance by the
> >>> administrators of the instance and/or contributors to GHC.
> >>>
> >>> What this won't disallow:
> >>>
> >>> * Getting an LLM to isolate and create a minimal reproducer
> >>> or bug description and posting this.
> >>>
> >>> * Getting an LLM to propose a patch or fix for a particular
> >>> issue
> >>>
> >>> * Asking an LLM about the details of how some particular
> >>> external system works, and then posting the results.
> >>>
> >>> * Asking an LLM to catalogue a bunch of related issues and
> >>> extract a common pattern
> >>>
> >>> * etc.
> >>>
> >>> What this will disallow:
> >>>
> >>> Autonomous systems without humans in the loop writing and
> >>> interacting with the gitlab.
> >>>
> >>> My reasons
> >>>
> >>> 1. Code review is a tool for building a shared understanding
> >>> and a common mental model of the codebase and problem
> >>> domain. An LLM model cannot do that, or even if you believe
> >>> it can, this internal model is lost as soon as that
> >>> particular session is garbage collected. At best you can try
> >>> to reconstruct some approximation of this model later by
> >>> feeding the same inputs into a fresh session on a probably
> >>> different model.
> >>>
> >>> 2. I would find it a very sad if the GHC Gitlab became a
> >>> place where bots talk to other bots who act upon and those
> >>> comments to make addtional changes in response to this in
> >>> an uncontrolled feedback loop. I predict that in the limit
> >>> any such feedback loop becomes quickly detached from the
> >>> original problem and devolves into a lot of noise totally
> >>> divorced from any real human concerns.
> >>>
> >>> We can already see examples of how this goes wrong **today**
> >>> Take a look at a prominent example of a codebase that went
> >>> "all in" on mostly unsupervised autonomous LLM development,
> >>> the javascript runtime Bun: https://github.com/oven-sh/bun/pulls
> >>>
> >>> It would kill any motivation I would have to contribute to
> >>> GHC and collaborate with others if the merge request queue
> >>> looked anything approximating that. I use the gitlab to
> >>> talk to and collaborate with actual humans, who I hope will
> >>> actually engage their brains and help get us to a better
> >>> and more corret mental model of the space.
> >>>
> >>> If I wanted to talk to an LLM, I have countless other
> >>> options for ways in which to do that.
> >>>
> >>> 3. Prompt injection. Gitlab comments are an untrusted
> >>> input source and LLMs fundamentally cannot discriminate
> >>> between instructions and data. You can imagine prompt
> >>> injection attacks where hidden text in a gitlab comment,
> >>> attached bug reproducer, link etc. gets the model to do
> >>> various destructive things, like spamming the gitlab,
> >>> editing/deleting your prior comments, mining bitcoin on the
> >>> CI runners, attempting to inject malicious code into the
> >>> compiler, testcases, reproducers etc.
> >>>
> >>> That doesn't mean that there isn't value for LLM assisted
> >>> autonomous processes on the gitlab, given that they are
> >>> desiged properly keeping in mind the concerns about human
> >>> centeredness, resource usage and prompt injection.
> >>>
> >>> Here are some examples of LLM assisted system/bots that
> >>> could actually be useful if audited and pre-approved with
> >>> the community:
> >>>
> >>> * A bot that automatically bisects regressions and identifies
> >>> the change responsible for it.
> >>> In principle this could be done with a purely deterministic
> >>> script, but often you run into problems with the toolchain
> >>> or commits not being individually buildable etc that require
> >>> some adjustements on the fly.
> >>>
> >>> * A bot that automatically minimises reproducers and posts
> >>> the results after the testcase can been verified on a deterministic
> >>> system
> >>>
> >>> On 26/07/27 19:38, Simon Jakobi via ghc-devs wrote:
> >>>> Hi devs!
> >>>>
> >>>> Inspired by the recent discussion on the LLM policy, I'd like to try
> >>>> introducing LLM-generated code reviews in GHC, as an opt-in service.
> >>>>
> >>>> My motivation:
> >>>>
> >>>> * Catch issues before human reviewers spend time on them.
> >>>> * Reduce the number of bugs merged.
> >>>>
> >>>>
> >>>> What this is — and what it isn't
> >>>> --------------------------------
> >>>>
> >>>> This is tool output: requested on the MR, produced by a model, posted 
> >>>> verbatim
> >>>> and clearly marked as machine-generated. This is not a review in the 
> >>>> sense of
> >>>> draft LLM policy.
> >>>>
> >>>> I'm not vouching for any of it. Think of it as a linter the author could 
> >>>> have
> >>>> run themselves, except on my LLM budget. Nobody is required to act on
> >>>> it even read it.
> >>>>
> >>>> I mention this explicitly because the policy asks reviewers to take full
> >>>> responsibility for each line of their review. I can't do that here — it 
> >>>> would
> >>>> mean checking every finding myself, which defeats the purpose and would
> >>>> restrict the service to areas of the compiler I already know well.
> >>>>
> >>>>
> >>>> How to request one
> >>>> ------------------
> >>>>
> >>>> * Ping me on the MR with a comment like "@sjakobi llm-review please".
> >>>>   * Feel free to request particular aspects the review should cover.
> >>>> * Authors can request a review for their own MRs. Maintainers can 
> >>>> request one
> >>>>   on any MR.
> >>>> * If I have enough usage left on my plan, I'll confirm that I'm on it.
> >>>> * I'll then prompt the model, and post the result on the MR.
> >>>> * Feedback on the quality of the review is of course very welcome.
> >>>>
> >>>>
> >>>> What I promise
> >>>> --------------
> >>>>
> >>>> * Every review is labelled as machine-generated, with model, effort 
> >>>> level and
> >>>>   the exact prompt used.
> >>>> * Reviews are posted verbatim. I don't edit or curate them, so what you 
> >>>> see is
> >>>>   exactly what the model produced.
> >>>> * Only the public MR diff and public repository context go to the model.
> >>>>   Nothing else.
> >>>> * I'll keep working on the prompts and the choice of model to improve the
> >>>>   signal-to-noise ratio.
> >>>> * If the signal-to-noise ratio stays inacceptably bad, we can simply
> >>>> stop this service.
> >>>>
> >>>>
> >>>> Review coverage
> >>>> ---------------
> >>>>
> >>>> I'll probably start with a fairly basic review prompt, primarily aimed at
> >>>> correctness issues.
> >>>>
> >>>> Possible extensions:
> >>>>
> >>>> * Documentation consistency: check that documentation stays consistent,
> >>>>   including Notes elsewhere that reference the changed code.
> >>>> * Performance: check for performance issues — potentially including 
> >>>> checking
> >>>>   the -ddump-simpl output for the changed code for unnecessary
> >>>>   allocation etc.
> >>>> * Feel free to suggest anything else!
> >>>>
> >>>>
> >>>> Format
> >>>> ------
> >>>>
> >>>> Initially a single comment per review. Once we reach a good 
> >>>> signal-to-noise
> >>>> ratio, we can consider inline comments on the diff.
> >>>>
> >>>>
> >>>> Caveats
> >>>> -------
> >>>>
> >>>> This is a volunteer service I intend to provide in my spare time. If I'm 
> >>>> AFK
> >>>> or on vacation, reviews will take longer. If I can't keep up with 
> >>>> requests, I
> >>>> will prioritise which MRs I run. I might take a break entirely.
> >>>>
> >>>>
> >>>> Getting involved
> >>>> ----------------
> >>>>
> >>>> Is anyone interested in joining this effort? If so we could form a team
> >>>> (@llm-reviewers?), and share prompting techniques, model choices, etc.
> >>>>
> >>>> If the service turns out to be genuinely useful, we can look into
> >>>> automating it and – if necessary – possible funding.
> >>>>
> >>>>
> >>>> Cheers,
> >>>> Simon
> >>>> _______________________________________________
> >>>> ghc-devs mailing list -- [email protected]
> >>>> To unsubscribe send an email to [email protected]
> >>> _______________________________________________
> >>> ghc-devs mailing list -- [email protected]
> >>> To unsubscribe send an email to [email protected]
> > _______________________________________________
> > ghc-devs mailing list -- [email protected]
> > To unsubscribe send an email to [email protected]
> _______________________________________________
> ghc-devs mailing list -- [email protected]
> To unsubscribe send an email to [email protected]
_______________________________________________
ghc-devs mailing list -- [email protected]
To unsubscribe send an email to [email protected]

Reply via email to