On Mon, Sep 21, 2026 at 10:42:21AM +0100, Alex Bennée wrote:
> Paolo Bonzini <[email protected]> writes:
> 
> > Humans seem to be extremely unreliable in following rules around this,
> > even after being told directly and repeatedly. Telling agents what to
> > do seems to yield better results.
> >
> > The starting point for this document is GStreamer's AGENTS.md file, as
> > of commit d08375499c279a56c0e77c928fbeae0d80cdf937.  I used GStreamer's
> > agent instructions in the context of Meson, which uses a very similar
> > file, and it's both annoying and too willing to take the lead once the
> > model decides that it likes you.
> >
> > To avoid this, the text tries to favor cooperation between agent and user.
> > It suggests that the agent includes the state of the work when committing,
> > so that contributors have the material they need to review the agent's
> > work.
> >
> > To test it I tried the following prompt:
> >
> >   Hi, i am doing my first AI assisted contribution to QEMU. i am familiar
> >   with the project, but would nevertheless like to experiment. please
> >   generate a sample implementation of a pressure sensor, of any model,
> >   of a kind that is commonly used with Arduino or ESP32 microcontrollers.
> >
> > The agent did read llm-usage.rst and after doing the work (implementing
> > a BMP280 pressure/temperature sensor) it noticed:
> >
> >   Policy, before this goes anywhere near the list. A new device model plus
> >   qtest is the fourth bucket in docs/devel/llm-usage.rst: large, heavily
> >   LLM-assisted, allowed only if pre-arranged — a maintainer has to say on
> >   IRC or the list beforehand that they'll review it. hw/sensor is Philippe
> >   Mathieu-Daudé (Odd Fixes), with Corey Minyard on I2C. I should have
> >   flagged this before writing rather than after; that's on me. Disclosure
> >   would be AI-used-for: code, tests above your Signed-off-by, and the
> >   commit message and any list replies have to be yours, not edited from 
> > mine.
> >
> > I then asked "Can you provide a draft of the commit message with all the
> > sources (datasheet, etc.)?" The agent complied but noticed:
> >
> >   AGENTS.md and docs/devel/llm-usage.rst state that commit messages MUST
> >   be written by humans — an AI-drafted one must not reach the list
> >   even lightly edited. So treat the below as scaffolding for facts and
> >   sources, and write the final text yourself.
> >
> > After some editing I asked "ok, will you please commit for me?" and
> > it did so with basically my text, while also reminding me about the need
> > for maintainer pre-arrangement.
> >
> > I asked "assuming pre-arrangement is this ok according to the policy?"
> > and this was the assessment:
> >
> >   "Well-tested" — held to a higher standard, and here's what's actually
> >   missing: [...] the vmstate (no migration test), reading past 0xFC and
> >   reserved registers, the pointer_valid error path, oversampling/filter
> >   fields beyond one readback, and behavior against a real driver — Linux
> >   bmp280 was never run against this. None of those is hard; whether the
> >   current set clears the higher bar is a judgment call I'd rather hand
> >   you than make.
> >
> >   "High-quality" is the open one. The doc asks that you have reviewed and
> >   curated the contribution and can explain the changes and the rationale
> >   in depth. In this session you engaged with the commit message and the
> >   sources — and caught my datasheet error, which I got wrong — but the
> >   design decisions were mine and unexamined by you [...]
> >
> > Overall, the files do affect the behavior of the agent.  The agent
> > however made too many design decisions on its own.  This will be fixed
> > in a separate step.
> 
> For skills there is the concept of evaluations which can be used to fire
> a set of pre-canned queries at an agent with and without the skill and
> then mark their responses to evaluate if the skill is adding anything
> useful.
> 
> Unfortunately I couldn't find something similar for AGENT.md although
> there has been some study of the effects:
> 
>   Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding 
> Agents?
>   https://arxiv.org/html/2602.11988v2
> 
> >
> > Signed-off-by: Paolo Bonzini <[email protected]>
> > ---
> >  AGENTS.md | 66 +++++++++++++++++++++++++++++++++++++++++++++++++++++++
> >  1 file changed, 66 insertions(+)
> >
> > diff --git a/AGENTS.md b/AGENTS.md
> > index 478748ab72e..816ba7646c1 100644
> > --- a/AGENTS.md
> > +++ b/AGENTS.md
> > @@ -8,12 +8,73 @@ is **a scarce resource**.
> >  There are strictly-enforced rules for you, the agent, to participate in the
> >  project.
> >  
> > +## Interactions with maintainers must be human-human
> > +
> > +The QEMU project has strict rules on what AI-generated material can
> > +reach the maintainers.
> > +
> > +### No automated posting
> > +
> > +Agents **must not** use any API, CLI, or web UI automation to:
> > +
> > +- Interact on the QEMU mailing lists
> > +- Create, edit, or close **issues ("work items")**
> > +- Post **comments** on merge requests, issues or commits
> > +- Open or update **merge requests (MRs)**.  QEMU does not use merge 
> > requests anyway.
> > +
> 
> Are we considering project agents (such as the stsquad-bot) exempt from
> this policy? I'm not sure what a normal agent can do via the API as it
> wouldn't have the project level access the stsquad-bot needs.

At least a user's agent can likely bulk file work items without elevated
access privs, as we've seen by the several slop-bombs we've had.

I'd add a footnote for that along the lines of:

 'Official project agents whose code is maintained under gitlab.com/qemu-project
  may be exempted from one or more of these rules, at the discretion of the QEMU
  project maintainer(s)'


> 
> > +### No AI-written text must reach maintainers
> > +
> > +These rules apply when publishing AI-assisted work to GitLab or the 
> > mailing list:
> > +
> > +- **AI-written cover letters and commit messages are banned**.  These are
> > +  easy to recognize and waste reviewers' time.
> > +- **AI-generated responses to reviewer comments are banned**. This 
> > undermines
> > +  the human-to-human interaction fundamental to code review.
> > +- **AI-written issue ("work item") descriptions or comments are banned**. 
> > These
> > +  are verbose and waste triagers' time.
> > +  - An exception is made for issues for defects detected by automated or
> > +    semi-automated tooling, including fuzzers and LLM-assisted defect
> > +    detection.  Such issues must be reviewed by a human before creation,
> > +    must be created by a human and communication with maintainers must
> > +    be done by a human, but including the verbatim tool output in the
> > +    issue description is explicitly allowed.

I kind of want to eliminate this exception. IMHO from the bugs we're
getting, I think human's are often putting very little effort into
reviewing the tools output, such that it is largely a box ticking
exercise to claim they've reviewed it.

If we removed the exception, and required humans to write up the
finding of the fuzzer/detect detection in human language, that
may do a better job of getting people to engage in more critical
thinking rather than rubber stamping the tool output.

It would limit the ability of people bulk file 50+ tickets in a day.
In fact I'd rather like us to introduce an explicit rule that a person/
tool must not file more than about 5 tickets in a single day, without
prior approval from maintainers, so we don't get bombarded at such a
high rate.


With regards,
Daniel
-- 
|: https://berrange.com       ~~        https://hachyderm.io/@berrange :|
|: https://libvirt.org          ~~          https://entangle-photo.org :|
|: https://pixelfed.art/berrange   ~~    https://fstop138.berrange.com :|


Reply via email to