On Thu, Sep 24, 2026 at 5:30 AM Daniel P. Berrangé <[email protected]> wrote:
> On Thu, Sep 24, 2026 at 12:46:41PM +0200, Kevin Wolf wrote: > > Am 24.09.2026 um 11:29 hat Alex Bennée geschrieben: > > > Kevin Wolf <[email protected]> writes: > > > > > > > Am 21.09.2026 um 15:20 hat Daniel P. Berrangé geschrieben: > > > >> On Mon, Sep 21, 2026 at 09:52:48AM +0200, Paolo Bonzini wrote: > > > >> > +.. note:: **Use of AI does not remove the need for authors to > comply > > > >> > + with all other requirements for contribution.** In > particular, > > > >> > + the ``Signed-off-by`` label in a patch submission is a > statement > > > >> > + that the author takes responsibility for the entire > contents of > > > >> > + the patch, certifying that their patch submission is > made in > > > >> > + accordance with the rules of the :ref:`Developer's > Certificate of > > > >> > + Origin (DCO) <dco>`. > > > >> > + > > > >> > + Since a submitter cannot audit LLM output against its > training > > > >> > + data, the DCO is paired with :ref:`metadata in the > commit message > > > >> > + <ai-used-for>` about AI-generated parts. The DCO > still certifies > > > >> > + that the contributor has the legal right to submit > code in general. > > > >> > > > >> I'm not a fan of this second paragraph, as that feels like it is > undermining > > > >> the DCO. That first sentence in particular is somewhat saying that > the > > > >> contributor does not have to think about plagarism and thed project > is ok > > > >> with that, and also implying that the DCO doesn't apply to the LLM > output. > > > >> This is both not OK in its implication of the project accepting some > > > >> liability, and then also contradicted by the next sentence. > > > >> > > > >> IMHO this paragraph should just be removed. The first paragraph > clearly > > > >> states the DCO applies to the submission as a whole, and leaves all > liability > > > >> for infringement on the contributor. > > > > > > > > I don't think making this the individual contributor's problem is > great. > > > > It's rather unfriendly towards contributors to expect them to certify > > > > something that we all know they can't honestly certify. > > > > > > This is the current position with the DCO anyway. If the contributor > has > > > copy and pasted (and maybe adapted) from Stack Overflow they are still > > > realistically the only person that can make that judgement. > > > > Yes, that's the current position with the DCO and the reasoning why we > > said LLM output is unacceptable in submissions. > > > > The big difference is if I copy and pasted from Stack Overflow, I _know_ > > that I did that. If the LLM effectively does it for me, I don't know > > that. The best I can do is check if I can find similar code elsewhere. > > > > So while I can certify with 100% confidence that I didn't manually copy > > things for code that I wrote myself, I can't certify that code from > > other sources, including LLMs, isn't copied from elsewhere. I can only > > say that I didn't find any reason to suspect it's copied. > > > > > Our current AI policy was built on the assertion that it's > impossible to > > > > sign the DCO for AI output, and I think that's still right. A > developer > > > > can't certify the origin of something when they don't really know > where > > > > it comes from. > > > > > > It's certainly possible for LLMs to replicate copyright material but by > > > far the biggest influence is the code *in the project* that shapes the > > > output. It is basically pattern matching and predicting after all. > > > > If there is one chunk of code that violates the copyright of someone, > > you can't heal that just by having 1000 more chunks that are clean. > > > > > What level of reassurance the submitter needs to be able to certify > with > > > their DCO is up to them. We can't police it, like every other > submission > > > we rely in the submitters good faith representation of what they have > > > done. > > > > The DCO doesn't talk about levels of reassurance. It talks about hard > > facts, 100% confidence. > > > > And that I can't give you with LLM generated content. I can maybe give > > you 99%, but never 100%. > > I appreciate you're speaking about licensing/copyright here, but more > generally I think we're probably wrong in asserting that a human > authored contribution can ever be said to be 100% safe to QEMU with > a DCO signoff. It is risk mitigation, but not absolute. > Risk mititagion is the best you can hope for. People lie, people accidentally copy, etc. Maybe they infringe on a patent. You can never be 100% sure. > From a copyright POV, if we write the code ourselves we can have > confidence we didn't copy it and thus licensing is likely OK (unless > you inadvertantly memorized some previous code you saw and reproduced > it too closely), but there are other legal aspects. Trademarks are > one risk, since we reference lots of vendor's hardware features, but > in practice this very rarely arises as a problem for projects, beyond > the obvious areas like project name, logo, and other visible branding > aspects. Software patents are the really big hidden trap that anyone > could unintentionally fall foul of at any time and have caused high > profile problems for projects periodically. > Even the memorized stuff is likely OK. It all depends. I have memorized bubble sorts, but there's only so many ways to implement a bubble sort so most bubble sorts don't enjoy much copyright protection. Big tables of numbers don't generally enjoy copyright protection, so are usually 100% OK to copy. Plus, LLMs generally don't copy code from random places. Some early ones were subject to injection attacks, but newer ones tend to look at surrounding code to adopt the repository's conventions. Is that copying? Is that a risk? Probably not because it's likely "scenes a faire" which is the doctrine sometimes the genre dictates that things be a certain way (eg all murder mysteries have a victim and someone trying to solve the crime). If someone knows their employer has a patent that is likely to be > restrictive for possible use in QEMU, then they shouldn't sign off > or contribute code implementing something covered by the patent > without the employer granting unrestricted use. We're certainly > not expecting people to do searches for 3rd party patents though, > so the DCO is not a 100% guarantee of risk elimination to QEMU. > In fact most companies, at least in the US, instruct their engineers to not even bother with a patent search because if you look for it, even if you don't find it, you can be held liable for triple damages because it can become willful. > We tolerate that risk because there is no practical action that a > contributor can take to mitigate. Further many large tech companies > have thus formed defensive pacts to de-fang the risk. > Yes. The actual legal landscape is such that it's hard to know for sure what the level of risk is for any given action. While usually the court will rule in a specific way, there are times it just doesn't, even for what seem like similar facts. Unlike programming, there's an element of randomness that you can't drive to 0. > With this in mind, it is perhaps not that different to accept a little > bit of uncertainty in risks of LLM unknowingly copying from training > material or 3rd party incompatibly licensed code it happened to find > while walking the web ? The key is that users need to be diligent > in their use of the tools, not reckless. Norms for what that means > are still be established - the so called "clean room" or license > laundering, re-impls of projects are a massively risky activity > but are the exception. > And "clean room" procedures are no different whether a human or an LLM performs them. They both produce a new work because the abstraction to the idea (which isn't copyrightable) loses details, creative choices, etc. And the implementation from that spec also makes different creative choices, leaving little in common with the original. > > > > If we decide that we don't care as much about the legal risks any > more > > > > and that we're willing to accept them to some extent, that should be > > > > explicitly reflected in the policy. > > > > > > > > It seems to me that the cleanest way to do it is to exempt correctly > > > > advertised (with 'AI-used-for:') AI output from the DCO requirement > and > > > > instead add to the 'AI-used-for:' definition some relaxed version of > it, > > > > e.g. "I have reviewed the contribution for potential licensing issues > > > > and haven't found a reason to doubt that I have the rights to submit > > > > this AI generated content under the open source license indicated in > the > > > > file". This is something that could realistically be certified by > > > > contributors in good faith and also isn't just "anything goes", but > of > > > > course it still is weaker than the DCO. > > I like the idea of associating "AI-used-for" with an explicit > statement like: > > "By using 'AI-used-for', a contributor both identifies areas of > the patch that involved use of AI/LLM, and also attests that > their usage of AI/LLMs was in compliance with the policies > outlined in this document" > > > Then we can add whatever guidance we want in the AI policy that > steers people away from LLM usage patterns that we know would > expose us to undesirable risk. Explicitly stating that we do > not want LLMs used to "clean room" / "license launder" re-impl > existing code is something we could include. Or guide that > agents should not be given free access to live git repos to > reduce change of direct copying ? > FreeBSD is struggling with these issues as well. There's a huge difference in risk of someone one-shotting the process and sending the results off w/o any critical thought and someone doing a series of a hundred or two hundred prompts with careful thought, arguments about the design trade offs, etc. Even leaving aside the copyright issues, the former has a huge risk, most of it from technical debt, while the latter has a much lower risk: there's less risk of copying (and any copying may fall to de minimis), and the result is much better designed and put together. FreeBSD doesn't yet have any good answers to share. Honestly, the real risk here is that there's a new kind of 'free' software: free as in free puppy, and you want to be careful which ones you accept. The bigger liability to the project is accepting the wrong ones which increase technical debt over the finer points of copyright law that will really only matter if something is actually litigated, and then the thousands of contributors randomly permuting different aspects of the code every few years also launders the code, both in terms of technical debt, and in terms of copyright risk. It's all about mitigating the risks that matter in the end. Warner
