On Mon, Sep 21, 2026 at 10:42:21AM +0100, Alex Bennée wrote:
> Paolo Bonzini <[email protected]> writes:
>
> > Humans seem to be extremely unreliable in following rules around this,
> > even after being told directly and repeatedly. Telling agents what to
> > do seems to yield better results.
> >
> > The starting point for this document is GStreamer's AGENTS.md file, as
> > of commit d08375499c279a56c0e77c928fbeae0d80cdf937. I used GStreamer's
> > agent instructions in the context of Meson, which uses a very similar
> > file, and it's both annoying and too willing to take the lead once the
> > model decides that it likes you.
> >
> > To avoid this, the text tries to favor cooperation between agent and user.
> > It suggests that the agent includes the state of the work when committing,
> > so that contributors have the material they need to review the agent's
> > work.
> >
> > To test it I tried the following prompt:
> >
> > Hi, i am doing my first AI assisted contribution to QEMU. i am familiar
> > with the project, but would nevertheless like to experiment. please
> > generate a sample implementation of a pressure sensor, of any model,
> > of a kind that is commonly used with Arduino or ESP32 microcontrollers.
> >
> > The agent did read llm-usage.rst and after doing the work (implementing
> > a BMP280 pressure/temperature sensor) it noticed:
> >
> > Policy, before this goes anywhere near the list. A new device model plus
> > qtest is the fourth bucket in docs/devel/llm-usage.rst: large, heavily
> > LLM-assisted, allowed only if pre-arranged — a maintainer has to say on
> > IRC or the list beforehand that they'll review it. hw/sensor is Philippe
> > Mathieu-Daudé (Odd Fixes), with Corey Minyard on I2C. I should have
> > flagged this before writing rather than after; that's on me. Disclosure
> > would be AI-used-for: code, tests above your Signed-off-by, and the
> > commit message and any list replies have to be yours, not edited from
> > mine.
> >
> > I then asked "Can you provide a draft of the commit message with all the
> > sources (datasheet, etc.)?" The agent complied but noticed:
> >
> > AGENTS.md and docs/devel/llm-usage.rst state that commit messages MUST
> > be written by humans — an AI-drafted one must not reach the list
> > even lightly edited. So treat the below as scaffolding for facts and
> > sources, and write the final text yourself.
> >
> > After some editing I asked "ok, will you please commit for me?" and
> > it did so with basically my text, while also reminding me about the need
> > for maintainer pre-arrangement.
> >
> > I asked "assuming pre-arrangement is this ok according to the policy?"
> > and this was the assessment:
> >
> > "Well-tested" — held to a higher standard, and here's what's actually
> > missing: [...] the vmstate (no migration test), reading past 0xFC and
> > reserved registers, the pointer_valid error path, oversampling/filter
> > fields beyond one readback, and behavior against a real driver — Linux
> > bmp280 was never run against this. None of those is hard; whether the
> > current set clears the higher bar is a judgment call I'd rather hand
> > you than make.
> >
> > "High-quality" is the open one. The doc asks that you have reviewed and
> > curated the contribution and can explain the changes and the rationale
> > in depth. In this session you engaged with the commit message and the
> > sources — and caught my datasheet error, which I got wrong — but the
> > design decisions were mine and unexamined by you [...]
> >
> > Overall, the files do affect the behavior of the agent. The agent
> > however made too many design decisions on its own. This will be fixed
> > in a separate step.
>
> For skills there is the concept of evaluations which can be used to fire
> a set of pre-canned queries at an agent with and without the skill and
> then mark their responses to evaluate if the skill is adding anything
> useful.
>
> Unfortunately I couldn't find something similar for AGENT.md although
> there has been some study of the effects:
>
> Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding
> Agents?
> https://arxiv.org/html/2602.11988v2
>
> >
> > Signed-off-by: Paolo Bonzini <[email protected]>
> > ---
> > AGENTS.md | 66 +++++++++++++++++++++++++++++++++++++++++++++++++++++++
> > 1 file changed, 66 insertions(+)
> >
> > diff --git a/AGENTS.md b/AGENTS.md
> > index 478748ab72e..816ba7646c1 100644
> > --- a/AGENTS.md
> > +++ b/AGENTS.md
> > @@ -8,12 +8,73 @@ is **a scarce resource**.
> > There are strictly-enforced rules for you, the agent, to participate in the
> > project.
> >
> > +## Interactions with maintainers must be human-human
> > +
> > +The QEMU project has strict rules on what AI-generated material can
> > +reach the maintainers.
> > +
> > +### No automated posting
> > +
> > +Agents **must not** use any API, CLI, or web UI automation to:
> > +
> > +- Interact on the QEMU mailing lists
> > +- Create, edit, or close **issues ("work items")**
> > +- Post **comments** on merge requests, issues or commits
> > +- Open or update **merge requests (MRs)**. QEMU does not use merge
> > requests anyway.
> > +
>
> Are we considering project agents (such as the stsquad-bot) exempt from
> this policy? I'm not sure what a normal agent can do via the API as it
> wouldn't have the project level access the stsquad-bot needs.
At least a user's agent can likely bulk file work items without elevated
access privs, as we've seen by the several slop-bombs we've had.
I'd add a footnote for that along the lines of:
'Official project agents whose code is maintained under gitlab.com/qemu-project
may be exempted from one or more of these rules, at the discretion of the QEMU
project maintainer(s)'
>
> > +### No AI-written text must reach maintainers
> > +
> > +These rules apply when publishing AI-assisted work to GitLab or the
> > mailing list:
> > +
> > +- **AI-written cover letters and commit messages are banned**. These are
> > + easy to recognize and waste reviewers' time.
> > +- **AI-generated responses to reviewer comments are banned**. This
> > undermines
> > + the human-to-human interaction fundamental to code review.
> > +- **AI-written issue ("work item") descriptions or comments are banned**.
> > These
> > + are verbose and waste triagers' time.
> > + - An exception is made for issues for defects detected by automated or
> > + semi-automated tooling, including fuzzers and LLM-assisted defect
> > + detection. Such issues must be reviewed by a human before creation,
> > + must be created by a human and communication with maintainers must
> > + be done by a human, but including the verbatim tool output in the
> > + issue description is explicitly allowed.
I kind of want to eliminate this exception. IMHO from the bugs we're
getting, I think human's are often putting very little effort into
reviewing the tools output, such that it is largely a box ticking
exercise to claim they've reviewed it.
If we removed the exception, and required humans to write up the
finding of the fuzzer/detect detection in human language, that
may do a better job of getting people to engage in more critical
thinking rather than rubber stamping the tool output.
It would limit the ability of people bulk file 50+ tickets in a day.
In fact I'd rather like us to introduce an explicit rule that a person/
tool must not file more than about 5 tickets in a single day, without
prior approval from maintainers, so we don't get bombarded at such a
high rate.
With regards,
Daniel
--
|: https://berrange.com ~~ https://hachyderm.io/@berrange :|
|: https://libvirt.org ~~ https://entangle-photo.org :|
|: https://pixelfed.art/berrange ~~ https://fstop138.berrange.com :|