On Thu, Sep 24, 2026 at 11:41:54AM +0200, Kevin Wolf wrote: > Am 24.09.2026 um 04:23 hat Michael S. Tsirkin geschrieben: > > On Wed, Sep 23, 2026 at 06:16:12PM +0200, Kevin Wolf wrote: > > > Am 21.09.2026 um 15:20 hat Daniel P. Berrangé geschrieben: > > > > On Mon, Sep 21, 2026 at 09:52:48AM +0200, Paolo Bonzini wrote: > > > > > +.. note:: **Use of AI does not remove the need for authors to comply > > > > > + with all other requirements for contribution.** In > > > > > particular, > > > > > + the ``Signed-off-by`` label in a patch submission is a > > > > > statement > > > > > + that the author takes responsibility for the entire > > > > > contents of > > > > > + the patch, certifying that their patch submission is made > > > > > in > > > > > + accordance with the rules of the :ref:`Developer's > > > > > Certificate of > > > > > + Origin (DCO) <dco>`. > > > > > + > > > > > + Since a submitter cannot audit LLM output against its > > > > > training > > > > > + data, the DCO is paired with :ref:`metadata in the commit > > > > > message > > > > > + <ai-used-for>` about AI-generated parts. The DCO still > > > > > certifies > > > > > + that the contributor has the legal right to submit code in > > > > > general. > > > > > > > > I'm not a fan of this second paragraph, as that feels like it is > > > > undermining > > > > the DCO. That first sentence in particular is somewhat saying that the > > > > contributor does not have to think about plagarism and thed project is > > > > ok > > > > with that, and also implying that the DCO doesn't apply to the LLM > > > > output. > > > > This is both not OK in its implication of the project accepting some > > > > liability, and then also contradicted by the next sentence. > > > > > > > > IMHO this paragraph should just be removed. The first paragraph clearly > > > > states the DCO applies to the submission as a whole, and leaves all > > > > liability > > > > for infringement on the contributor. > > > > > > I don't think making this the individual contributor's problem is great. > > > It's rather unfriendly towards contributors to expect them to certify > > > something that we all know they can't honestly certify. > > > > > > Our current AI policy was built on the assertion that it's impossible to > > > sign the DCO for AI output, and I think that's still right. A developer > > > can't certify the origin of something when they don't really know where > > > it comes from. > > > > > > If we decide that we don't care as much about the legal risks any more > > > and that we're willing to accept them to some extent, that should be > > > explicitly reflected in the policy. > > > > > > It seems to me that the cleanest way to do it is to exempt correctly > > > advertised (with 'AI-used-for:') AI output from the DCO requirement and > > > instead add to the 'AI-used-for:' definition some relaxed version of it, > > > e.g. "I have reviewed the contribution for potential licensing issues > > > and haven't found a reason to doubt that I have the rights to submit > > > this AI generated content under the open source license indicated in the > > > file". This is something that could realistically be certified by > > > contributors in good faith and also isn't just "anything goes", but of > > > course it still is weaker than the DCO. > > > > > > If we're not willing to take the legal risk from this, we should > > > probably leave the current AI policy unchanged instead of telling > > > contributors that they have to work in a legal grey area. > > > > > > Kevin > > > > Which risk precisely? That a judge will close all LLMs overnight > > declaring their output a derivative of the internet and so illegal to > > use? Please. > > I mean, at least morally that would be the right thing to do. If what > they did isn't illegal, that's a bug in the law.
I'm not prepared to argue morality, and I don't even know who "they" would be. But if an LLM is transformative then it's not a derivative, that's a feature of the law, not a bug. > But no, I don't expect this to happen, but am referring to the more > commonly talked about case of LLMs occasionally reproducing a single > source more or less verbatim, which is what could give you legal trouble > even practically speaking. They don't really seem to do that unless you try to ask them to. > You can't know for sure if this happens, so I don't think you could sign > the DCO (you really can't know for sure if you have "right to submit it > under the open source license indicated in the file"), but you can check > for suspicious things and if you don't find any, say it's probably okay. > > You seem to be saying that it's not only probably okay, but definitely > okay, but then that's even less of a problem. > > > But just like "AI" does not mean "illegal" it does not mean "legal". > > > > For example, if I give an LLM a chunk of code and say "adopt this to qemu" > > then > > an argument can be made that it is a derivative of the code I supplied. > > > > Whether to give an LLM web access so it can pull in random bits of code > > and create derivatives of that is entirely under developer's control. > > > > Whether AI output is a derivative of LLM weights is also an open > > question, mitigated by the fact that many LLM vendors will assign > > copyright to their users. Whether to use an LLM where that is the case > > is also under developer's control. > > > > I'm not sure why we need to get into so much detail, though. > > See, you bring up uncertainties yourself that nobody can answer today. Oh people can easily make sure their LLM is not poking at the web, today. > This is exactly why I wouldn't want to make people lie and say "I know > this is definitely okay", but accept "I tried to be careful and to the > best of my knowledge there is no problem" instead for the licensing of > LLM output. > > This seems reasonable enough to me. I'm just saying that if others > disagree and we're not willing to accept the weaker assertion that > people actually can make in good faith, then we should stay out of LLM > generated content entirely instead of making people lie. > > Kevin
