Paolo Bonzini <[email protected]> writes: > Humans seem to be extremely unreliable in following rules around this, > even after being told directly and repeatedly. Telling agents what to > do seems to yield better results. > > The starting point for this document is GStreamer's AGENTS.md file, as > of commit d08375499c279a56c0e77c928fbeae0d80cdf937. I used GStreamer's > agent instructions in the context of Meson, which uses a very similar > file, and it's both annoying and too willing to take the lead once the > model decides that it likes you. > > To avoid this, the text tries to favor cooperation between agent and user. > It suggests that the agent includes the state of the work when committing, > so that contributors have the material they need to review the agent's > work. > > To test it I tried the following prompt: > > Hi, i am doing my first AI assisted contribution to QEMU. i am familiar > with the project, but would nevertheless like to experiment. please > generate a sample implementation of a pressure sensor, of any model, > of a kind that is commonly used with Arduino or ESP32 microcontrollers. > > The agent did read llm-usage.rst and after doing the work (implementing > a BMP280 pressure/temperature sensor) it noticed: > > Policy, before this goes anywhere near the list. A new device model plus > qtest is the fourth bucket in docs/devel/llm-usage.rst: large, heavily > LLM-assisted, allowed only if pre-arranged — a maintainer has to say on > IRC or the list beforehand that they'll review it. hw/sensor is Philippe > Mathieu-Daudé (Odd Fixes), with Corey Minyard on I2C. I should have > flagged this before writing rather than after; that's on me. Disclosure > would be AI-used-for: code, tests above your Signed-off-by, and the > commit message and any list replies have to be yours, not edited from mine. > > I then asked "Can you provide a draft of the commit message with all the > sources (datasheet, etc.)?" The agent complied but noticed: > > AGENTS.md and docs/devel/llm-usage.rst state that commit messages MUST > be written by humans — an AI-drafted one must not reach the list > even lightly edited. So treat the below as scaffolding for facts and > sources, and write the final text yourself. > > After some editing I asked "ok, will you please commit for me?" and > it did so with basically my text, while also reminding me about the need > for maintainer pre-arrangement. > > I asked "assuming pre-arrangement is this ok according to the policy?" > and this was the assessment: > > "Well-tested" — held to a higher standard, and here's what's actually > missing: [...] the vmstate (no migration test), reading past 0xFC and > reserved registers, the pointer_valid error path, oversampling/filter > fields beyond one readback, and behavior against a real driver — Linux > bmp280 was never run against this. None of those is hard; whether the > current set clears the higher bar is a judgment call I'd rather hand > you than make. > > "High-quality" is the open one. The doc asks that you have reviewed and > curated the contribution and can explain the changes and the rationale > in depth. In this session you engaged with the commit message and the > sources — and caught my datasheet error, which I got wrong — but the > design decisions were mine and unexamined by you [...] > > Overall, the files do affect the behavior of the agent. The agent > however made too many design decisions on its own. This will be fixed > in a separate step.
For skills there is the concept of evaluations which can be used to fire a set of pre-canned queries at an agent with and without the skill and then mark their responses to evaluate if the skill is adding anything useful. Unfortunately I couldn't find something similar for AGENT.md although there has been some study of the effects: Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents? https://arxiv.org/html/2602.11988v2 > > Signed-off-by: Paolo Bonzini <[email protected]> > --- > AGENTS.md | 66 +++++++++++++++++++++++++++++++++++++++++++++++++++++++ > 1 file changed, 66 insertions(+) > > diff --git a/AGENTS.md b/AGENTS.md > index 478748ab72e..816ba7646c1 100644 > --- a/AGENTS.md > +++ b/AGENTS.md > @@ -8,12 +8,73 @@ is **a scarce resource**. > There are strictly-enforced rules for you, the agent, to participate in the > project. > > +## Interactions with maintainers must be human-human > + > +The QEMU project has strict rules on what AI-generated material can > +reach the maintainers. > + > +### No automated posting > + > +Agents **must not** use any API, CLI, or web UI automation to: > + > +- Interact on the QEMU mailing lists > +- Create, edit, or close **issues ("work items")** > +- Post **comments** on merge requests, issues or commits > +- Open or update **merge requests (MRs)**. QEMU does not use merge requests > anyway. > + Are we considering project agents (such as the stsquad-bot) exempt from this policy? I'm not sure what a normal agent can do via the API as it wouldn't have the project level access the stsquad-bot needs. > +### No AI-written text must reach maintainers > + > +These rules apply when publishing AI-assisted work to GitLab or the mailing > list: > + > +- **AI-written cover letters and commit messages are banned**. These are > + easy to recognize and waste reviewers' time. > +- **AI-generated responses to reviewer comments are banned**. This undermines > + the human-to-human interaction fundamental to code review. > +- **AI-written issue ("work item") descriptions or comments are banned**. > These > + are verbose and waste triagers' time. > + - An exception is made for issues for defects detected by automated or > + semi-automated tooling, including fuzzers and LLM-assisted defect > + detection. Such issues must be reviewed by a human before creation, > + must be created by a human and communication with maintainers must > + be done by a human, but including the verbatim tool output in the > + issue description is explicitly allowed. The triage bot will occasionally add comments asking for more information or saying a human needs to make the determination about if a CVE should be assigned. > + > +Copy editing of human-written text, for example to help non-native speakers, > +is allowed. Keep such edits small and precise. > + > +The human's workflow may require you to perform commits or request you to > +produce private drafts of maintainer-facing text. These are allowed, but > must > +be marked as requiring rewrite by a human before publication. > + <snip> -- Alex Bennée Virtualisation Tech Lead @ Linaro
