Am 27.09.2026 um 11:20 hat Michael S. Tsirkin geschrieben: > On Thu, Sep 24, 2026 at 08:06:05PM +0200, Kevin Wolf wrote: > > Am 24.09.2026 um 18:06 hat Michael S. Tsirkin geschrieben: > > > On Thu, Sep 24, 2026 at 12:46:41PM +0200, Kevin Wolf wrote: > > > > Am 24.09.2026 um 11:29 hat Alex Bennée geschrieben: > > > > > Kevin Wolf <[email protected]> writes: > > > > > > > > > > > Am 21.09.2026 um 15:20 hat Daniel P. Berrangé geschrieben: > > > > > >> On Mon, Sep 21, 2026 at 09:52:48AM +0200, Paolo Bonzini wrote: > > > > > >> > +.. note:: **Use of AI does not remove the need for authors to > > > > > >> > comply > > > > > >> > + with all other requirements for contribution.** In > > > > > >> > particular, > > > > > >> > + the ``Signed-off-by`` label in a patch submission is > > > > > >> > a statement > > > > > >> > + that the author takes responsibility for the entire > > > > > >> > contents of > > > > > >> > + the patch, certifying that their patch submission is > > > > > >> > made in > > > > > >> > + accordance with the rules of the :ref:`Developer's > > > > > >> > Certificate of > > > > > >> > + Origin (DCO) <dco>`. > > > > > >> > + > > > > > >> > + Since a submitter cannot audit LLM output against its > > > > > >> > training > > > > > >> > + data, the DCO is paired with :ref:`metadata in the > > > > > >> > commit message > > > > > >> > + <ai-used-for>` about AI-generated parts. The DCO > > > > > >> > still certifies > > > > > >> > + that the contributor has the legal right to submit > > > > > >> > code in general. > > > > > >> > > > > > >> I'm not a fan of this second paragraph, as that feels like it is > > > > > >> undermining > > > > > >> the DCO. That first sentence in particular is somewhat saying that > > > > > >> the > > > > > >> contributor does not have to think about plagarism and thed > > > > > >> project is ok > > > > > >> with that, and also implying that the DCO doesn't apply to the LLM > > > > > >> output. > > > > > >> This is both not OK in its implication of the project accepting > > > > > >> some > > > > > >> liability, and then also contradicted by the next sentence. > > > > > >> > > > > > >> IMHO this paragraph should just be removed. The first paragraph > > > > > >> clearly > > > > > >> states the DCO applies to the submission as a whole, and leaves > > > > > >> all liability > > > > > >> for infringement on the contributor. > > > > > > > > > > > > I don't think making this the individual contributor's problem is > > > > > > great. > > > > > > It's rather unfriendly towards contributors to expect them to > > > > > > certify > > > > > > something that we all know they can't honestly certify. > > > > > > > > > > This is the current position with the DCO anyway. If the contributor > > > > > has > > > > > copy and pasted (and maybe adapted) from Stack Overflow they are still > > > > > realistically the only person that can make that judgement. > > > > > > > > Yes, that's the current position with the DCO and the reasoning why we > > > > said LLM output is unacceptable in submissions. > > > > > > > > The big difference is if I copy and pasted from Stack Overflow, I _know_ > > > > that I did that. If the LLM effectively does it for me, I don't know > > > > that. The best I can do is check if I can find similar code elsewhere. > > > > > > You can block your LLM from accessing stack overflow. > > > > Given that training my own model from scratch isn't a realistic option, > > I can't block the LLM from being trained on data from Stack Overflow. > > > > Kevin > > But again, that is beside the point, unless you believe that a judge > will (or should?) close all LLMs overnight declaring their output a > derivative of the internet and so illegal to use. If not, training is > transformative, and LLM output is not a derivative of the internet > *unless it is input with Stack Overflow*.
You're treating things as black and white here, which is maybe how things should be, but not how they are in practice today. At least my understanding is that the common interpretation is that LLM output is generally okay, but if it reproduces an existing work too closely, then you're still infringing on its copyright and that's your own problem. Just because in >99% of cases the model doesn't reproduce things from the training data too closely, that doesn't automatically make everything it produces from training data clean. And this is the part that courts in different jurisdictions have even decided on, so it's not just speculation. I'm not aware of any court declaring that LLMs are purely transformative and everything that comes from training data is automatically good (though I'm also not aware of a court deciding the opposite; it seems to be always just about the result in the individual case and how much it resembles existing works). Kevin
