Am 28.09.2026 um 10:23 hat Michael S. Tsirkin geschrieben:
> On Mon, Sep 28, 2026 at 09:58:15AM +0200, Kevin Wolf wrote:
> > Am 27.09.2026 um 11:20 hat Michael S. Tsirkin geschrieben:
> > > On Thu, Sep 24, 2026 at 08:06:05PM +0200, Kevin Wolf wrote:
> > > > Am 24.09.2026 um 18:06 hat Michael S. Tsirkin geschrieben:
> > > > > On Thu, Sep 24, 2026 at 12:46:41PM +0200, Kevin Wolf wrote:
> > > > > > Am 24.09.2026 um 11:29 hat Alex Bennée geschrieben:
> > > > > > > Kevin Wolf <[email protected]> writes:
> > > > > > > 
> > > > > > > > Am 21.09.2026 um 15:20 hat Daniel P. Berrangé geschrieben:
> > > > > > > >> On Mon, Sep 21, 2026 at 09:52:48AM +0200, Paolo Bonzini wrote:
> > > > > > > >> > +.. note:: **Use of AI does not remove the need for authors 
> > > > > > > >> > to comply
> > > > > > > >> > +          with all other requirements for contribution.**  
> > > > > > > >> > In particular,
> > > > > > > >> > +          the ``Signed-off-by`` label in a patch submission 
> > > > > > > >> > is a statement
> > > > > > > >> > +          that the author takes responsibility for the 
> > > > > > > >> > entire contents of
> > > > > > > >> > +          the patch, certifying that their patch submission 
> > > > > > > >> > is made in
> > > > > > > >> > +          accordance with the rules of the 
> > > > > > > >> > :ref:`Developer's Certificate of
> > > > > > > >> > +          Origin (DCO) <dco>`.
> > > > > > > >> > +
> > > > > > > >> > +          Since a submitter cannot audit LLM output against 
> > > > > > > >> > its training
> > > > > > > >> > +          data, the DCO is paired with :ref:`metadata in 
> > > > > > > >> > the commit message
> > > > > > > >> > +          <ai-used-for>` about AI-generated parts.  The DCO 
> > > > > > > >> > still certifies
> > > > > > > >> > +          that the contributor has the legal right to 
> > > > > > > >> > submit code in general.
> > > > > > > >> 
> > > > > > > >> I'm not a fan of this second paragraph, as that feels like it 
> > > > > > > >> is undermining
> > > > > > > >> the DCO. That first sentence in particular is somewhat saying 
> > > > > > > >> that the
> > > > > > > >> contributor does not have to think about plagarism and thed 
> > > > > > > >> project is ok
> > > > > > > >> with that, and also implying that the DCO doesn't apply to the 
> > > > > > > >> LLM output.
> > > > > > > >> This is both not OK in its implication of the project 
> > > > > > > >> accepting some
> > > > > > > >> liability, and then also contradicted by the next sentence.
> > > > > > > >> 
> > > > > > > >> IMHO this paragraph should just be removed.  The first 
> > > > > > > >> paragraph clearly
> > > > > > > >> states the DCO applies to the submission as a whole, and 
> > > > > > > >> leaves all liability
> > > > > > > >> for infringement on the contributor.
> > > > > > > >
> > > > > > > > I don't think making this the individual contributor's problem 
> > > > > > > > is great.
> > > > > > > > It's rather unfriendly towards contributors to expect them to 
> > > > > > > > certify
> > > > > > > > something that we all know they can't honestly certify.
> > > > > > > 
> > > > > > > This is the current position with the DCO anyway. If the 
> > > > > > > contributor has
> > > > > > > copy and pasted (and maybe adapted) from Stack Overflow they are 
> > > > > > > still
> > > > > > > realistically the only person that can make that judgement.
> > > > > > 
> > > > > > Yes, that's the current position with the DCO and the reasoning why 
> > > > > > we
> > > > > > said LLM output is unacceptable in submissions.
> > > > > > 
> > > > > > The big difference is if I copy and pasted from Stack Overflow, I 
> > > > > > _know_
> > > > > > that I did that. If the LLM effectively does it for me, I don't know
> > > > > > that. The best I can do is check if I can find similar code 
> > > > > > elsewhere.
> > > > > 
> > > > > You can block your LLM from accessing stack overflow.
> > > > 
> > > > Given that training my own model from scratch isn't a realistic option,
> > > > I can't block the LLM from being trained on data from Stack Overflow.
> > > > 
> > > > Kevin
> > > 
> > > But again, that is beside the point, unless you believe that a judge
> > > will (or should?) close all LLMs overnight declaring their output a
> > > derivative of the internet and so illegal to use. If not, training is
> > > transformative, and LLM output is not a derivative of the internet
> > > *unless it is input with Stack Overflow*.
> > 
> > You're treating things as black and white here, which is maybe how
> > things should be, but not how they are in practice today. At least my
> > understanding is that the common interpretation is that LLM output is
> > generally okay, but if it reproduces an existing work too closely, then
> > you're still infringing on its copyright and that's your own problem.
> 
> As far as I understand, it's not true that you are "still infringing" -
> rather, the courts did not yet decide this question.
> Correct me if I am wrong.

I'm not following the legal developments too closely, so other may know
more, but there have been court decisions in the news at least in the
context of music. Memorised song lyrics are probably quite comparable
to memorised code.

> They might in theory decide that yes.  We can decide to wait until some
> court decides that no, and decline AI contributions until then.  Though
> I note that there's no limit here - there's always a chance some court
> decides differently. Traditionally, the free software community used
> e.g. the GPL without waiting for it to be "tested in courts".

"training data is always clean" isn't like licensing your own code under
the GPL when it hadn't been "tested in courts" yet, but more like
copying GPL code into incompatibly licensed software with the
justification that courts hadn't ruled on it yet.

We don't have to wait for courts to change their mind on memorisation,
but we can simply require that people be careful and do their best to
avoid reproducing existing code in their contributions. Because if you
don't have a copy of something in your contribution, that's the best way
to know that you're not infringing on its copyright.

> > Just because in >99% of cases the model doesn't reproduce things from
> > the training data too closely, that doesn't automatically make
> > everything it produces from training data clean. And this is the part
> > that courts in different jurisdictions have even decided on, so it's not
> > just speculation. I'm not aware of any court declaring that LLMs are
> > purely transformative and everything that comes from training data is
> > automatically good (though I'm also not aware of a court deciding the
> > opposite; it seems to be always just about the result in the individual
> > case and how much it resembles existing works).
> 
> Really, QEMU internal APIs and coding style being as unique as they
> are, I have a lot of trouble even imagining any code that is
> substantially similar to anything on the internet and still being
> accepted into QEMU.

It depends. A patch that converts an error_report() deep in a call stack
of void functions to passing an Error ** everywhere seems pretty safe. A
patch that adds a new device model or backends with 1000 LOC or more can
easily copy an algorithm from somewhere in a static function. If you
want to submit something like the latter, it's probably not asking too
much that you search the internet for similar code.

Kevin


Reply via email to