Paolo Bonzini <[email protected]> writes:

> Humans seem to be extremely unreliable in following rules around this,
> even after being told directly and repeatedly. Telling agents what to
> do seems to yield better results.
>
> The starting point for this document is GStreamer's AGENTS.md file, as
> of commit d08375499c279a56c0e77c928fbeae0d80cdf937.  I used GStreamer's
> agent instructions in the context of Meson, which uses a very similar
> file, and it's both annoying and too willing to take the lead once the
> model decides that it likes you.
>
> To avoid this, the text tries to favor cooperation between agent and user.
> It suggests that the agent includes the state of the work when committing,
> so that contributors have the material they need to review the agent's
> work.
>
> To test it I tried the following prompt:
>
>   Hi, i am doing my first AI assisted contribution to QEMU. i am familiar
>   with the project, but would nevertheless like to experiment. please
>   generate a sample implementation of a pressure sensor, of any model,
>   of a kind that is commonly used with Arduino or ESP32 microcontrollers.
>
> The agent did read llm-usage.rst and after doing the work (implementing
> a BMP280 pressure/temperature sensor) it noticed:
>
>   Policy, before this goes anywhere near the list. A new device model plus
>   qtest is the fourth bucket in docs/devel/llm-usage.rst: large, heavily
>   LLM-assisted, allowed only if pre-arranged — a maintainer has to say on
>   IRC or the list beforehand that they'll review it. hw/sensor is Philippe
>   Mathieu-Daudé (Odd Fixes), with Corey Minyard on I2C. I should have
>   flagged this before writing rather than after; that's on me. Disclosure
>   would be AI-used-for: code, tests above your Signed-off-by, and the
>   commit message and any list replies have to be yours, not edited from mine.
>
> I then asked "Can you provide a draft of the commit message with all the
> sources (datasheet, etc.)?" The agent complied but noticed:
>
>   AGENTS.md and docs/devel/llm-usage.rst state that commit messages MUST
>   be written by humans — an AI-drafted one must not reach the list
>   even lightly edited. So treat the below as scaffolding for facts and
>   sources, and write the final text yourself.
>
> After some editing I asked "ok, will you please commit for me?" and
> it did so with basically my text, while also reminding me about the need
> for maintainer pre-arrangement.
>
> I asked "assuming pre-arrangement is this ok according to the policy?"
> and this was the assessment:
>
>   "Well-tested" — held to a higher standard, and here's what's actually
>   missing: [...] the vmstate (no migration test), reading past 0xFC and
>   reserved registers, the pointer_valid error path, oversampling/filter
>   fields beyond one readback, and behavior against a real driver — Linux
>   bmp280 was never run against this. None of those is hard; whether the
>   current set clears the higher bar is a judgment call I'd rather hand
>   you than make.
>
>   "High-quality" is the open one. The doc asks that you have reviewed and
>   curated the contribution and can explain the changes and the rationale
>   in depth. In this session you engaged with the commit message and the
>   sources — and caught my datasheet error, which I got wrong — but the
>   design decisions were mine and unexamined by you [...]
>
> Overall, the files do affect the behavior of the agent.  The agent
> however made too many design decisions on its own.  This will be fixed
> in a separate step.

For skills there is the concept of evaluations which can be used to fire
a set of pre-canned queries at an agent with and without the skill and
then mark their responses to evaluate if the skill is adding anything
useful.

Unfortunately I couldn't find something similar for AGENT.md although
there has been some study of the effects:

  Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding 
Agents?
  https://arxiv.org/html/2602.11988v2

>
> Signed-off-by: Paolo Bonzini <[email protected]>
> ---
>  AGENTS.md | 66 +++++++++++++++++++++++++++++++++++++++++++++++++++++++
>  1 file changed, 66 insertions(+)
>
> diff --git a/AGENTS.md b/AGENTS.md
> index 478748ab72e..816ba7646c1 100644
> --- a/AGENTS.md
> +++ b/AGENTS.md
> @@ -8,12 +8,73 @@ is **a scarce resource**.
>  There are strictly-enforced rules for you, the agent, to participate in the
>  project.
>  
> +## Interactions with maintainers must be human-human
> +
> +The QEMU project has strict rules on what AI-generated material can
> +reach the maintainers.
> +
> +### No automated posting
> +
> +Agents **must not** use any API, CLI, or web UI automation to:
> +
> +- Interact on the QEMU mailing lists
> +- Create, edit, or close **issues ("work items")**
> +- Post **comments** on merge requests, issues or commits
> +- Open or update **merge requests (MRs)**.  QEMU does not use merge requests 
> anyway.
> +

Are we considering project agents (such as the stsquad-bot) exempt from
this policy? I'm not sure what a normal agent can do via the API as it
wouldn't have the project level access the stsquad-bot needs.

> +### No AI-written text must reach maintainers
> +
> +These rules apply when publishing AI-assisted work to GitLab or the mailing 
> list:
> +
> +- **AI-written cover letters and commit messages are banned**.  These are
> +  easy to recognize and waste reviewers' time.
> +- **AI-generated responses to reviewer comments are banned**. This undermines
> +  the human-to-human interaction fundamental to code review.
> +- **AI-written issue ("work item") descriptions or comments are banned**. 
> These
> +  are verbose and waste triagers' time.
> +  - An exception is made for issues for defects detected by automated or
> +    semi-automated tooling, including fuzzers and LLM-assisted defect
> +    detection.  Such issues must be reviewed by a human before creation,
> +    must be created by a human and communication with maintainers must
> +    be done by a human, but including the verbatim tool output in the
> +    issue description is explicitly allowed.

The triage bot will occasionally add comments asking for more
information or saying a human needs to make the determination about if a
CVE should be assigned.

> +
> +Copy editing of human-written text, for example to help non-native speakers,
> +is allowed. Keep such edits small and precise.
> +
> +The human's workflow may require you to perform commits or request you to
> +produce private drafts of maintainer-facing text.  These are allowed, but 
> must
> +be marked as requiring rewrite by a human before publication.
> +
<snip>

-- 
Alex Bennée
Virtualisation Tech Lead @ Linaro

Reply via email to