Am 24.09.2026 um 11:29 hat Alex Bennée geschrieben: > Kevin Wolf <[email protected]> writes: > > > Am 21.09.2026 um 15:20 hat Daniel P. Berrangé geschrieben: > >> On Mon, Sep 21, 2026 at 09:52:48AM +0200, Paolo Bonzini wrote: > >> > +.. note:: **Use of AI does not remove the need for authors to comply > >> > + with all other requirements for contribution.** In > >> > particular, > >> > + the ``Signed-off-by`` label in a patch submission is a > >> > statement > >> > + that the author takes responsibility for the entire contents > >> > of > >> > + the patch, certifying that their patch submission is made in > >> > + accordance with the rules of the :ref:`Developer's > >> > Certificate of > >> > + Origin (DCO) <dco>`. > >> > + > >> > + Since a submitter cannot audit LLM output against its training > >> > + data, the DCO is paired with :ref:`metadata in the commit > >> > message > >> > + <ai-used-for>` about AI-generated parts. The DCO still > >> > certifies > >> > + that the contributor has the legal right to submit code in > >> > general. > >> > >> I'm not a fan of this second paragraph, as that feels like it is > >> undermining > >> the DCO. That first sentence in particular is somewhat saying that the > >> contributor does not have to think about plagarism and thed project is ok > >> with that, and also implying that the DCO doesn't apply to the LLM output. > >> This is both not OK in its implication of the project accepting some > >> liability, and then also contradicted by the next sentence. > >> > >> IMHO this paragraph should just be removed. The first paragraph clearly > >> states the DCO applies to the submission as a whole, and leaves all > >> liability > >> for infringement on the contributor. > > > > I don't think making this the individual contributor's problem is great. > > It's rather unfriendly towards contributors to expect them to certify > > something that we all know they can't honestly certify. > > This is the current position with the DCO anyway. If the contributor has > copy and pasted (and maybe adapted) from Stack Overflow they are still > realistically the only person that can make that judgement.
Yes, that's the current position with the DCO and the reasoning why we said LLM output is unacceptable in submissions. The big difference is if I copy and pasted from Stack Overflow, I _know_ that I did that. If the LLM effectively does it for me, I don't know that. The best I can do is check if I can find similar code elsewhere. So while I can certify with 100% confidence that I didn't manually copy things for code that I wrote myself, I can't certify that code from other sources, including LLMs, isn't copied from elsewhere. I can only say that I didn't find any reason to suspect it's copied. > > Our current AI policy was built on the assertion that it's impossible to > > sign the DCO for AI output, and I think that's still right. A developer > > can't certify the origin of something when they don't really know where > > it comes from. > > It's certainly possible for LLMs to replicate copyright material but by > far the biggest influence is the code *in the project* that shapes the > output. It is basically pattern matching and predicting after all. If there is one chunk of code that violates the copyright of someone, you can't heal that just by having 1000 more chunks that are clean. > What level of reassurance the submitter needs to be able to certify with > their DCO is up to them. We can't police it, like every other submission > we rely in the submitters good faith representation of what they have > done. The DCO doesn't talk about levels of reassurance. It talks about hard facts, 100% confidence. And that I can't give you with LLM generated content. I can maybe give you 99%, but never 100%. > > If we decide that we don't care as much about the legal risks any more > > and that we're willing to accept them to some extent, that should be > > explicitly reflected in the policy. > > > > It seems to me that the cleanest way to do it is to exempt correctly > > advertised (with 'AI-used-for:') AI output from the DCO requirement and > > instead add to the 'AI-used-for:' definition some relaxed version of it, > > e.g. "I have reviewed the contribution for potential licensing issues > > and haven't found a reason to doubt that I have the rights to submit > > this AI generated content under the open source license indicated in the > > file". This is something that could realistically be certified by > > contributors in good faith and also isn't just "anything goes", but of > > course it still is weaker than the DCO. > > > > If we're not willing to take the legal risk from this, we should > > probably leave the current AI policy unchanged instead of telling > > contributors that they have to work in a legal grey area. > > I don't think we are asking them to work in a legally grey area, we are > insisting they are confident there is no legal risk in what they submit. If anyone is 100% "confident there is no legal risk", then I think they're simply wrong (unless the contribution is obviously not copyrightable anyway). The real question is if the risk is low enough. But I suspect in most cases it's actually more like they're only 99% confident and think "ah well, I'll sign it anyway, I probably won't get into trouble". That gives people who are willing to bend the rules a little an advantage over more conscientious contributors, which I don't think is justified, because in the end it's the same 99% certainty for both. Kevin
