On Thu, Sep 24, 2026 at 12:46:41PM +0200, Kevin Wolf wrote:
> Am 24.09.2026 um 11:29 hat Alex Bennée geschrieben:
> > Kevin Wolf <[email protected]> writes:
> > 
> > > Am 21.09.2026 um 15:20 hat Daniel P. Berrangé geschrieben:
> > >> On Mon, Sep 21, 2026 at 09:52:48AM +0200, Paolo Bonzini wrote:
> > >> > +.. note:: **Use of AI does not remove the need for authors to comply
> > >> > +          with all other requirements for contribution.**  In 
> > >> > particular,
> > >> > +          the ``Signed-off-by`` label in a patch submission is a 
> > >> > statement
> > >> > +          that the author takes responsibility for the entire 
> > >> > contents of
> > >> > +          the patch, certifying that their patch submission is made in
> > >> > +          accordance with the rules of the :ref:`Developer's 
> > >> > Certificate of
> > >> > +          Origin (DCO) <dco>`.
> > >> > +
> > >> > +          Since a submitter cannot audit LLM output against its 
> > >> > training
> > >> > +          data, the DCO is paired with :ref:`metadata in the commit 
> > >> > message
> > >> > +          <ai-used-for>` about AI-generated parts.  The DCO still 
> > >> > certifies
> > >> > +          that the contributor has the legal right to submit code in 
> > >> > general.
> > >> 
> > >> I'm not a fan of this second paragraph, as that feels like it is 
> > >> undermining
> > >> the DCO. That first sentence in particular is somewhat saying that the
> > >> contributor does not have to think about plagarism and thed project is ok
> > >> with that, and also implying that the DCO doesn't apply to the LLM 
> > >> output.
> > >> This is both not OK in its implication of the project accepting some
> > >> liability, and then also contradicted by the next sentence.
> > >> 
> > >> IMHO this paragraph should just be removed.  The first paragraph clearly
> > >> states the DCO applies to the submission as a whole, and leaves all 
> > >> liability
> > >> for infringement on the contributor.
> > >
> > > I don't think making this the individual contributor's problem is great.
> > > It's rather unfriendly towards contributors to expect them to certify
> > > something that we all know they can't honestly certify.
> > 
> > This is the current position with the DCO anyway. If the contributor has
> > copy and pasted (and maybe adapted) from Stack Overflow they are still
> > realistically the only person that can make that judgement.
> 
> Yes, that's the current position with the DCO and the reasoning why we
> said LLM output is unacceptable in submissions.
> 
> The big difference is if I copy and pasted from Stack Overflow, I _know_
> that I did that. If the LLM effectively does it for me, I don't know
> that. The best I can do is check if I can find similar code elsewhere.
> 
> So while I can certify with 100% confidence that I didn't manually copy
> things for code that I wrote myself, I can't certify that code from
> other sources, including LLMs, isn't copied from elsewhere. I can only
> say that I didn't find any reason to suspect it's copied.

> > > Our current AI policy was built on the assertion that it's impossible to
> > > sign the DCO for AI output, and I think that's still right. A developer
> > > can't certify the origin of something when they don't really know where
> > > it comes from.
> > 
> > It's certainly possible for LLMs to replicate copyright material but by
> > far the biggest influence is the code *in the project* that shapes the
> > output. It is basically pattern matching and predicting after all.
> 
> If there is one chunk of code that violates the copyright of someone,
> you can't heal that just by having 1000 more chunks that are clean.
> 
> > What level of reassurance the submitter needs to be able to certify with
> > their DCO is up to them. We can't police it, like every other submission
> > we rely in the submitters good faith representation of what they have
> > done.
> 
> The DCO doesn't talk about levels of reassurance. It talks about hard
> facts, 100% confidence.
> 
> And that I can't give you with LLM generated content. I can maybe give
> you 99%, but never 100%.

I appreciate you're speaking about licensing/copyright here, but more
generally I think we're probably wrong in asserting that a human
authored contribution can ever be said to be 100% safe to QEMU with
a DCO signoff. It is risk mitigation, but not absolute.
 
>From a copyright POV, if we write the code ourselves we can have
confidence we didn't copy it and thus licensing is likely OK (unless
you inadvertantly memorized some previous code you saw and reproduced
it too closely), but there are other legal aspects. Trademarks are
one risk, since we reference lots of vendor's hardware features, but
in practice this very rarely arises as a problem for projects, beyond
the obvious areas like project name, logo, and other visible branding
aspects. Software patents are the really big hidden trap that anyone
could unintentionally fall foul of at any time and have caused high
profile problems for projects periodically.

If someone knows their employer has a patent that is likely to be
restrictive for possible use in QEMU, then they shouldn't sign off
or contribute code implementing something covered by the patent
without the employer granting unrestricted use. We're certainly
not expecting people to do searches for 3rd party patents though,
so the DCO is not a 100% guarantee of risk elimination to QEMU.

We tolerate that risk because there is no practical action that a
contributor can take to mitigate. Further many large tech companies
have thus formed defensive pacts to de-fang the risk.


With this in mind, it is perhaps not that different to accept a little
bit of uncertainty in risks of LLM unknowingly copying from training
material or 3rd party incompatibly licensed code it happened to find
while walking the web ?  The key is that users need to be diligent
in their use of the tools, not reckless. Norms for what that means
are still be established - the so called "clean room" or license
laundering, re-impls of projects are a massively risky activity
but are the exception.



> > > If we decide that we don't care as much about the legal risks any more
> > > and that we're willing to accept them to some extent, that should be
> > > explicitly reflected in the policy.
> > >
> > > It seems to me that the cleanest way to do it is to exempt correctly
> > > advertised (with 'AI-used-for:') AI output from the DCO requirement and
> > > instead add to the 'AI-used-for:' definition some relaxed version of it,
> > > e.g. "I have reviewed the contribution for potential licensing issues
> > > and haven't found a reason to doubt that I have the rights to submit
> > > this AI generated content under the open source license indicated in the
> > > file". This is something that could realistically be certified by
> > > contributors in good faith and also isn't just "anything goes", but of
> > > course it still is weaker than the DCO.

I like the idea of associating "AI-used-for" with an explicit
statement like:

  "By using 'AI-used-for', a contributor both identifies areas of
   the patch that involved use of AI/LLM, and also attests that
   their usage of AI/LLMs was in compliance with the policies
   outlined in this document"


Then we can add whatever guidance we want in the AI policy that
steers people away from LLM usage patterns that we know would
expose us to undesirable risk. Explicitly stating that we do
not want LLMs used to "clean room" / "license launder" re-impl
existing code is something we could include. Or guide that 
agents should not be given free access to live git repos to
reduce change of direct copying ?


With regards,
Daniel
-- 
|: https://berrange.com       ~~        https://hachyderm.io/@berrange :|
|: https://libvirt.org          ~~          https://entangle-photo.org :|
|: https://pixelfed.art/berrange   ~~    https://fstop138.berrange.com :|


Reply via email to