Ian Eure <[email protected]> writes:
> It matters because the provenence of the code is different. Prior > to LLMs, "tool" meant, say, nano, or vi, or Emacs, or a > typewriter, or some other thing which a human used to directly > create the code. In that scenario, I agree, the tool does not > matter. This situation is entirely different, as the provenance > of the code no longer ends at the contributor, because they did > not create it. Thus, there is a substantial risk of bugs due to > inattention, or of the code not being copyrightable, or containing > infringing material. Those risks exist for hand-written code too. Inattentive developers, uncopyrightable trivial code, and infringing copy-paste predate LLMs. Review is how we catch them, regardless of origin. > It is both process and artifact, as the one produces the other. > If my process for creating software is to reuse as much infringing > code as possible, the created software is tainted because of that, > and its users bear some risk -- even if the risk is merely that > they can no longer use it. By this logic any artifact produced through a process touching copyrighted material is tainted. The English I'm writing this in was learned from copyrighted books and films. Is this email tainted? At some point process-purity arguments collapse under their own weight. > This is a false comparison, because those are tools used by > humans. Humans do not create LLM output. That’s the whole point > of them. RMS disagrees. In March he asked emacs-devel to drop support for every LLM in core and ELPA *except* Apertus, which he considered actually free[1]. The FSF position is not "LLMs taint software" but "non-free LLMs are the problem, free ones are fine." That distinction is fatal to the labelling proposal. If some LLMs meet free-software standards and others don't, what is the label tracking? Freedom? We already track that. Tool category? That's not a freedom concern. The metadata field has no coherent referent. > People who like LLMs seem to really like them, so it strikes me as > quite odd that they would be against a property indicating that > they were used. Such a thing seems like it would be a mark of > pride. I'm not against the property because I'm hiding something. I'm against it because it doesn't track anything Guix users need to know, and because labelling software by properties of its authors rather than properties of itself is a road I don't think Guix should start down. If users want to filter packages by upstream tooling, that's a legitimate preference and Guix's open data model already supports it; someone can build a third-party tool that maintains such a list, the same way people maintain databases for F-Droid downstream. That doesn't require Guix itself to take a position, collect the metadata, or enforce disclosure. [1] https://yhetil.org/emacs-devel/[email protected]/ -- Thanos Apollo ☧ https://thanosapollo.org
signature.asc
Description: PGP signature
