On 9/25/26 12:05, Kevin Wolf wrote:
What has changed so fundamentally (i.e. not just shifted in numbers)
since then?

I think what has changed is that people have gained more experience in what LLMs can do, which is hard to put in a commit message, and that it can do way better than it used to even in areas where the user can drive the design. These areas are the ones where risk of involuntary copyright laundering is smallest, because human input is larger. Basically they became good at areas that are less risky in terms of copyright, and that's more interesting than areas that are less risky in terms of consequences (like tests).

We weren't wrong in being conservative, but I do believe the risk balance has shifted. Contributors can be creative while using AI, and the possibility of using it for more interesting and creative things than "give me a device model" already mitigates the risk.

The fact that some risk exists is clear from Conservancy's idea of treating logs like corresponding source, even though I don't think it's practical. I think a different and more practical approach is to steer how people use the tool: the AGENTS.md instructions in this series, and the requirement that people send well organized patch series, together should make the output "more human" and less subject to involunyary repetition.

On top of this, we have the arms race of fixing vs. finding vulnerabilities, where the fix has limited or no creativity and is way below the copyright threshold (this does not mean the fix is correct---I still think LLMs lack taste in many ways---but that's the role of the maintainer no matter who submits the fix).

With this in mind, it is perhaps not that different to accept a little
bit of uncertainty in risks of LLM unknowingly copying from training
material or 3rd party incompatibly licensed code it happened to find
while walking the web ?  The key is that users need to be diligent
in their use of the tools, not reckless. Norms for what that means
are still be established - the so called "clean room" or license
laundering, re-impls of projects are a massively risky activity
but are the exception.

So basically you're saying that it's our interpretation of the DCO that
was wrong from the beginning, and it shouldn't actually be understood as
strict/literally as we said it should?

Not entirely. But we tried to make things black and white, which is generally an approximation. For example we ignored altogether incremental improvements, which can be sizeable, for the sake of having *a* policy: "check if this function call can be moved earlier, and do it if so", small bug fixes, code that can't really be derivative of something outside QEMU (e.g. upgrading to a new version of an in-development kernel API).

As you replied elsewhere, a patch that adds a new device model of several thousand lines need to be handled with great care. Even if it was accepted, one might seriously consider taking the output of the LLM and clean-rooming it *by hand*. A new board that just pulls together existing devices, or a new accelerator, is considerably less risky even if it's in the same LOC ballpark.

I trust maintainers to understand where the risk lies. We're still the same people that accepted the previous policy and I don't see us going berserk on accepting AI contributions.

Even if people generally agree that that's the case, that should be made
explicit. With the history of our claim that it's impossible to sign the
DCO for AI output, we can't just mention in passing that the usual
requirements like signing the DCO apply as if that is the most natural
and uncontroversial thing to do, when we ourselves are claiming the
opposite today.

I think that's what Paolo tried to address with the paragraph you
suggested to remove. While its wording isn't perfect, I do think we need
something like it.

Yes. The idea that AI-used-for is an attestation of its own, and that the DCO covers the AI-used-for attestation, is a pretty good replacement. Something like:

---
In order to certify the provenance of AI/LLM-generated code that you
submit to QEMU, you must follow the policies outlined in this document
and in QEMU's ``AGENTS.md``.  **By using 'AI-used-for', a contributor
both identifies areas of the patch that involved use of AI/LLM, and also
attests that their usage of AI/LLMs was in compliance with the policies
outlined in this document.**
---

Paolo


Reply via email to