On Thu, Sep 24, 2026 at 12:46:41PM +0200, Kevin Wolf wrote:
> Am 24.09.2026 um 11:29 hat Alex Bennée geschrieben:
> > Kevin Wolf <[email protected]> writes:
> > 
> > > Am 21.09.2026 um 15:20 hat Daniel P. Berrangé geschrieben:
> > >> On Mon, Sep 21, 2026 at 09:52:48AM +0200, Paolo Bonzini wrote:
> > >> > +.. note:: **Use of AI does not remove the need for authors to comply
> > >> > +          with all other requirements for contribution.**  In 
> > >> > particular,
> > >> > +          the ``Signed-off-by`` label in a patch submission is a 
> > >> > statement
> > >> > +          that the author takes responsibility for the entire 
> > >> > contents of
> > >> > +          the patch, certifying that their patch submission is made in
> > >> > +          accordance with the rules of the :ref:`Developer's 
> > >> > Certificate of
> > >> > +          Origin (DCO) <dco>`.
> > >> > +
> > >> > +          Since a submitter cannot audit LLM output against its 
> > >> > training
> > >> > +          data, the DCO is paired with :ref:`metadata in the commit 
> > >> > message
> > >> > +          <ai-used-for>` about AI-generated parts.  The DCO still 
> > >> > certifies
> > >> > +          that the contributor has the legal right to submit code in 
> > >> > general.
> > >> 
> > >> I'm not a fan of this second paragraph, as that feels like it is 
> > >> undermining
> > >> the DCO. That first sentence in particular is somewhat saying that the
> > >> contributor does not have to think about plagarism and thed project is ok
> > >> with that, and also implying that the DCO doesn't apply to the LLM 
> > >> output.
> > >> This is both not OK in its implication of the project accepting some
> > >> liability, and then also contradicted by the next sentence.
> > >> 
> > >> IMHO this paragraph should just be removed.  The first paragraph clearly
> > >> states the DCO applies to the submission as a whole, and leaves all 
> > >> liability
> > >> for infringement on the contributor.
> > >
> > > I don't think making this the individual contributor's problem is great.
> > > It's rather unfriendly towards contributors to expect them to certify
> > > something that we all know they can't honestly certify.
> > 
> > This is the current position with the DCO anyway. If the contributor has
> > copy and pasted (and maybe adapted) from Stack Overflow they are still
> > realistically the only person that can make that judgement.
> 
> Yes, that's the current position with the DCO and the reasoning why we
> said LLM output is unacceptable in submissions.
> 
> The big difference is if I copy and pasted from Stack Overflow, I _know_
> that I did that. If the LLM effectively does it for me, I don't know
> that. The best I can do is check if I can find similar code elsewhere.

You can block your LLM from accessing stack overflow.


> 
> So while I can certify with 100% confidence that I didn't manually copy
> things for code that I wrote myself, I can't certify that code from
> other sources, including LLMs, isn't copied from elsewhere. I can only
> say that I didn't find any reason to suspect it's copied.
> 
> > > Our current AI policy was built on the assertion that it's impossible to
> > > sign the DCO for AI output, and I think that's still right. A developer
> > > can't certify the origin of something when they don't really know where
> > > it comes from.
> > 
> > It's certainly possible for LLMs to replicate copyright material but by
> > far the biggest influence is the code *in the project* that shapes the
> > output. It is basically pattern matching and predicting after all.
> 
> If there is one chunk of code that violates the copyright of someone,
> you can't heal that just by having 1000 more chunks that are clean.
> 
> > What level of reassurance the submitter needs to be able to certify with
> > their DCO is up to them. We can't police it, like every other submission
> > we rely in the submitters good faith representation of what they have
> > done.
> 
> The DCO doesn't talk about levels of reassurance. It talks about hard
> facts, 100% confidence.
> 
> And that I can't give you with LLM generated content. I can maybe give
> you 99%, but never 100%.
> 
> > > If we decide that we don't care as much about the legal risks any more
> > > and that we're willing to accept them to some extent, that should be
> > > explicitly reflected in the policy.
> > >
> > > It seems to me that the cleanest way to do it is to exempt correctly
> > > advertised (with 'AI-used-for:') AI output from the DCO requirement and
> > > instead add to the 'AI-used-for:' definition some relaxed version of it,
> > > e.g. "I have reviewed the contribution for potential licensing issues
> > > and haven't found a reason to doubt that I have the rights to submit
> > > this AI generated content under the open source license indicated in the
> > > file". This is something that could realistically be certified by
> > > contributors in good faith and also isn't just "anything goes", but of
> > > course it still is weaker than the DCO.
> > >
> > > If we're not willing to take the legal risk from this, we should
> > > probably leave the current AI policy unchanged instead of telling
> > > contributors that they have to work in a legal grey area.
> > 
> > I don't think we are asking them to work in a legally grey area, we are
> > insisting they are confident there is no legal risk in what they submit.
> 
> If anyone is 100% "confident there is no legal risk", then I think
> they're simply wrong (unless the contribution is obviously not
> copyrightable anyway). The real question is if the risk is low enough.
> 
> But I suspect in most cases it's actually more like they're only 99%
> confident and think "ah well, I'll sign it anyway, I probably won't get
> into trouble". That gives people who are willing to bend the rules a
> little an advantage over more conscientious contributors, which I don't
> think is justified, because in the end it's the same 99% certainty for
> both.
> 
> Kevin


Reply via email to