On Thu, Sep 24, 2026 at 12:46:41PM +0200, Kevin Wolf wrote: > Am 24.09.2026 um 11:29 hat Alex Bennée geschrieben: > > Kevin Wolf <[email protected]> writes: > > > > > Am 21.09.2026 um 15:20 hat Daniel P. Berrangé geschrieben: > > >> On Mon, Sep 21, 2026 at 09:52:48AM +0200, Paolo Bonzini wrote: > > >> > +.. note:: **Use of AI does not remove the need for authors to comply > > >> > + with all other requirements for contribution.** In > > >> > particular, > > >> > + the ``Signed-off-by`` label in a patch submission is a > > >> > statement > > >> > + that the author takes responsibility for the entire > > >> > contents of > > >> > + the patch, certifying that their patch submission is made in > > >> > + accordance with the rules of the :ref:`Developer's > > >> > Certificate of > > >> > + Origin (DCO) <dco>`. > > >> > + > > >> > + Since a submitter cannot audit LLM output against its > > >> > training > > >> > + data, the DCO is paired with :ref:`metadata in the commit > > >> > message > > >> > + <ai-used-for>` about AI-generated parts. The DCO still > > >> > certifies > > >> > + that the contributor has the legal right to submit code in > > >> > general. > > >> > > >> I'm not a fan of this second paragraph, as that feels like it is > > >> undermining > > >> the DCO. That first sentence in particular is somewhat saying that the > > >> contributor does not have to think about plagarism and thed project is ok > > >> with that, and also implying that the DCO doesn't apply to the LLM > > >> output. > > >> This is both not OK in its implication of the project accepting some > > >> liability, and then also contradicted by the next sentence. > > >> > > >> IMHO this paragraph should just be removed. The first paragraph clearly > > >> states the DCO applies to the submission as a whole, and leaves all > > >> liability > > >> for infringement on the contributor. > > > > > > I don't think making this the individual contributor's problem is great. > > > It's rather unfriendly towards contributors to expect them to certify > > > something that we all know they can't honestly certify. > > > > This is the current position with the DCO anyway. If the contributor has > > copy and pasted (and maybe adapted) from Stack Overflow they are still > > realistically the only person that can make that judgement. > > Yes, that's the current position with the DCO and the reasoning why we > said LLM output is unacceptable in submissions. > > The big difference is if I copy and pasted from Stack Overflow, I _know_ > that I did that. If the LLM effectively does it for me, I don't know > that. The best I can do is check if I can find similar code elsewhere. > > So while I can certify with 100% confidence that I didn't manually copy > things for code that I wrote myself, I can't certify that code from > other sources, including LLMs, isn't copied from elsewhere. I can only > say that I didn't find any reason to suspect it's copied.
> > > Our current AI policy was built on the assertion that it's impossible to > > > sign the DCO for AI output, and I think that's still right. A developer > > > can't certify the origin of something when they don't really know where > > > it comes from. > > > > It's certainly possible for LLMs to replicate copyright material but by > > far the biggest influence is the code *in the project* that shapes the > > output. It is basically pattern matching and predicting after all. > > If there is one chunk of code that violates the copyright of someone, > you can't heal that just by having 1000 more chunks that are clean. > > > What level of reassurance the submitter needs to be able to certify with > > their DCO is up to them. We can't police it, like every other submission > > we rely in the submitters good faith representation of what they have > > done. > > The DCO doesn't talk about levels of reassurance. It talks about hard > facts, 100% confidence. > > And that I can't give you with LLM generated content. I can maybe give > you 99%, but never 100%. I appreciate you're speaking about licensing/copyright here, but more generally I think we're probably wrong in asserting that a human authored contribution can ever be said to be 100% safe to QEMU with a DCO signoff. It is risk mitigation, but not absolute. >From a copyright POV, if we write the code ourselves we can have confidence we didn't copy it and thus licensing is likely OK (unless you inadvertantly memorized some previous code you saw and reproduced it too closely), but there are other legal aspects. Trademarks are one risk, since we reference lots of vendor's hardware features, but in practice this very rarely arises as a problem for projects, beyond the obvious areas like project name, logo, and other visible branding aspects. Software patents are the really big hidden trap that anyone could unintentionally fall foul of at any time and have caused high profile problems for projects periodically. If someone knows their employer has a patent that is likely to be restrictive for possible use in QEMU, then they shouldn't sign off or contribute code implementing something covered by the patent without the employer granting unrestricted use. We're certainly not expecting people to do searches for 3rd party patents though, so the DCO is not a 100% guarantee of risk elimination to QEMU. We tolerate that risk because there is no practical action that a contributor can take to mitigate. Further many large tech companies have thus formed defensive pacts to de-fang the risk. With this in mind, it is perhaps not that different to accept a little bit of uncertainty in risks of LLM unknowingly copying from training material or 3rd party incompatibly licensed code it happened to find while walking the web ? The key is that users need to be diligent in their use of the tools, not reckless. Norms for what that means are still be established - the so called "clean room" or license laundering, re-impls of projects are a massively risky activity but are the exception. > > > If we decide that we don't care as much about the legal risks any more > > > and that we're willing to accept them to some extent, that should be > > > explicitly reflected in the policy. > > > > > > It seems to me that the cleanest way to do it is to exempt correctly > > > advertised (with 'AI-used-for:') AI output from the DCO requirement and > > > instead add to the 'AI-used-for:' definition some relaxed version of it, > > > e.g. "I have reviewed the contribution for potential licensing issues > > > and haven't found a reason to doubt that I have the rights to submit > > > this AI generated content under the open source license indicated in the > > > file". This is something that could realistically be certified by > > > contributors in good faith and also isn't just "anything goes", but of > > > course it still is weaker than the DCO. I like the idea of associating "AI-used-for" with an explicit statement like: "By using 'AI-used-for', a contributor both identifies areas of the patch that involved use of AI/LLM, and also attests that their usage of AI/LLMs was in compliance with the policies outlined in this document" Then we can add whatever guidance we want in the AI policy that steers people away from LLM usage patterns that we know would expose us to undesirable risk. Explicitly stating that we do not want LLMs used to "clean room" / "license launder" re-impl existing code is something we could include. Or guide that agents should not be given free access to live git repos to reduce change of direct copying ? With regards, Daniel -- |: https://berrange.com ~~ https://hachyderm.io/@berrange :| |: https://libvirt.org ~~ https://entangle-photo.org :| |: https://pixelfed.art/berrange ~~ https://fstop138.berrange.com :|
