On 9/25/26 12:05, Kevin Wolf wrote:
What has changed so fundamentally (i.e. not just shifted in numbers)
since then?
I think what has changed is that people have gained more experience in
what LLMs can do, which is hard to put in a commit message, and that it
can do way better than it used to even in areas where the user can drive
the design. These areas are the ones where risk of involuntary
copyright laundering is smallest, because human input is larger.
Basically they became good at areas that are less risky in terms of
copyright, and that's more interesting than areas that are less risky in
terms of consequences (like tests).
We weren't wrong in being conservative, but I do believe the risk
balance has shifted. Contributors can be creative while using AI, and
the possibility of using it for more interesting and creative things
than "give me a device model" already mitigates the risk.
The fact that some risk exists is clear from Conservancy's idea of
treating logs like corresponding source, even though I don't think it's
practical. I think a different and more practical approach is to steer
how people use the tool: the AGENTS.md instructions in this series, and
the requirement that people send well organized patch series, together
should make the output "more human" and less subject to involunyary
repetition.
On top of this, we have the arms race of fixing vs. finding
vulnerabilities, where the fix has limited or no creativity and is way
below the copyright threshold (this does not mean the fix is correct---I
still think LLMs lack taste in many ways---but that's the role of the
maintainer no matter who submits the fix).
With this in mind, it is perhaps not that different to accept a little
bit of uncertainty in risks of LLM unknowingly copying from training
material or 3rd party incompatibly licensed code it happened to find
while walking the web ? The key is that users need to be diligent
in their use of the tools, not reckless. Norms for what that means
are still be established - the so called "clean room" or license
laundering, re-impls of projects are a massively risky activity
but are the exception.
So basically you're saying that it's our interpretation of the DCO that
was wrong from the beginning, and it shouldn't actually be understood as
strict/literally as we said it should?
Not entirely. But we tried to make things black and white, which is
generally an approximation. For example we ignored altogether
incremental improvements, which can be sizeable, for the sake of having
*a* policy: "check if this function call can be moved earlier, and do it
if so", small bug fixes, code that can't really be derivative of
something outside QEMU (e.g. upgrading to a new version of an
in-development kernel API).
As you replied elsewhere, a patch that adds a new device model of
several thousand lines need to be handled with great care. Even if it
was accepted, one might seriously consider taking the output of the LLM
and clean-rooming it *by hand*. A new board that just pulls together
existing devices, or a new accelerator, is considerably less risky even
if it's in the same LOC ballpark.
I trust maintainers to understand where the risk lies. We're still the
same people that accepted the previous policy and I don't see us going
berserk on accepting AI contributions.
Even if people generally agree that that's the case, that should be made
explicit. With the history of our claim that it's impossible to sign the
DCO for AI output, we can't just mention in passing that the usual
requirements like signing the DCO apply as if that is the most natural
and uncontroversial thing to do, when we ourselves are claiming the
opposite today.
I think that's what Paolo tried to address with the paragraph you
suggested to remove. While its wording isn't perfect, I do think we need
something like it.
Yes. The idea that AI-used-for is an attestation of its own, and that
the DCO covers the AI-used-for attestation, is a pretty good
replacement. Something like:
---
In order to certify the provenance of AI/LLM-generated code that you
submit to QEMU, you must follow the policies outlined in this document
and in QEMU's ``AGENTS.md``. **By using 'AI-used-for', a contributor
both identifies areas of the patch that involved use of AI/LLM, and also
attests that their usage of AI/LLMs was in compliance with the policies
outlined in this document.**
---
Paolo