On Thu, Sep 24, 2026 at 11:41:54AM +0200, Kevin Wolf wrote:
> Am 24.09.2026 um 04:23 hat Michael S. Tsirkin geschrieben:
> > On Wed, Sep 23, 2026 at 06:16:12PM +0200, Kevin Wolf wrote:
> > > Am 21.09.2026 um 15:20 hat Daniel P. Berrangé geschrieben:
> > > > On Mon, Sep 21, 2026 at 09:52:48AM +0200, Paolo Bonzini wrote:
> > > > > +.. note:: **Use of AI does not remove the need for authors to comply
> > > > > +          with all other requirements for contribution.**  In 
> > > > > particular,
> > > > > +          the ``Signed-off-by`` label in a patch submission is a 
> > > > > statement
> > > > > +          that the author takes responsibility for the entire 
> > > > > contents of
> > > > > +          the patch, certifying that their patch submission is made 
> > > > > in
> > > > > +          accordance with the rules of the :ref:`Developer's 
> > > > > Certificate of
> > > > > +          Origin (DCO) <dco>`.
> > > > > +
> > > > > +          Since a submitter cannot audit LLM output against its 
> > > > > training
> > > > > +          data, the DCO is paired with :ref:`metadata in the commit 
> > > > > message
> > > > > +          <ai-used-for>` about AI-generated parts.  The DCO still 
> > > > > certifies
> > > > > +          that the contributor has the legal right to submit code in 
> > > > > general.
> > > > 
> > > > I'm not a fan of this second paragraph, as that feels like it is 
> > > > undermining
> > > > the DCO. That first sentence in particular is somewhat saying that the
> > > > contributor does not have to think about plagarism and thed project is 
> > > > ok
> > > > with that, and also implying that the DCO doesn't apply to the LLM 
> > > > output.
> > > > This is both not OK in its implication of the project accepting some
> > > > liability, and then also contradicted by the next sentence.
> > > > 
> > > > IMHO this paragraph should just be removed.  The first paragraph clearly
> > > > states the DCO applies to the submission as a whole, and leaves all 
> > > > liability
> > > > for infringement on the contributor.
> > > 
> > > I don't think making this the individual contributor's problem is great.
> > > It's rather unfriendly towards contributors to expect them to certify
> > > something that we all know they can't honestly certify.
> > > 
> > > Our current AI policy was built on the assertion that it's impossible to
> > > sign the DCO for AI output, and I think that's still right. A developer
> > > can't certify the origin of something when they don't really know where
> > > it comes from.
> > > 
> > > If we decide that we don't care as much about the legal risks any more
> > > and that we're willing to accept them to some extent, that should be
> > > explicitly reflected in the policy.
> > > 
> > > It seems to me that the cleanest way to do it is to exempt correctly
> > > advertised (with 'AI-used-for:') AI output from the DCO requirement and
> > > instead add to the 'AI-used-for:' definition some relaxed version of it,
> > > e.g. "I have reviewed the contribution for potential licensing issues
> > > and haven't found a reason to doubt that I have the rights to submit
> > > this AI generated content under the open source license indicated in the
> > > file". This is something that could realistically be certified by
> > > contributors in good faith and also isn't just "anything goes", but of
> > > course it still is weaker than the DCO.
> > > 
> > > If we're not willing to take the legal risk from this, we should
> > > probably leave the current AI policy unchanged instead of telling
> > > contributors that they have to work in a legal grey area.
> > > 
> > > Kevin
> > 
> > Which risk precisely? That a judge will close all LLMs overnight
> > declaring their output a derivative of the internet and so illegal to
> > use? Please.
> 
> I mean, at least morally that would be the right thing to do. If what
> they did isn't illegal, that's a bug in the law.

I'm not prepared to argue morality, and I don't even know who "they"
would be. But if an LLM is transformative then it's not a derivative,
that's a feature of the law, not a bug.


> But no, I don't expect this to happen, but am referring to the more
> commonly talked about case of LLMs occasionally reproducing a single
> source more or less verbatim, which is what could give you legal trouble
> even practically speaking.

They don't really seem to do that unless you try to ask them to.

> You can't know for sure if this happens, so I don't think you could sign
> the DCO (you really can't know for sure if you have "right to submit it
> under the open source license indicated in the file"), but you can check
> for suspicious things and if you don't find any, say it's probably okay.
> 
> You seem to be saying that it's not only probably okay, but definitely
> okay, but then that's even less of a problem.
> 
> > But just like "AI" does not mean "illegal" it does not mean "legal".
> > 
> > For example, if I give an LLM a chunk of code and say "adopt this to qemu" 
> > then
> > an argument can be made that it is a derivative of the code I supplied.
> > 
> > Whether to give an LLM web access so it can pull in random bits of code
> > and create derivatives of that is entirely under developer's control.
> > 
> > Whether AI output is a derivative of LLM weights is also an open
> > question, mitigated by the fact that many LLM vendors will assign
> > copyright to their users. Whether to use an LLM where that is the case
> > is also under developer's control.
> > 
> > I'm not sure why we need to get into so much detail, though.
> 
> See, you bring up uncertainties yourself that nobody can answer today.

Oh people can easily make sure their LLM is not poking at the web, today.

> This is exactly why I wouldn't want to make people lie and say "I know
> this is definitely okay", but accept "I tried to be careful and to the
> best of my knowledge there is no problem" instead for the licensing of
> LLM output.
> 
> This seems reasonable enough to me. I'm just saying that if others
> disagree and we're not willing to accept the weaker assertion that
> people actually can make in good faith, then we should stay out of LLM
> generated content entirely instead of making people lie.
> 
> Kevin


Reply via email to