On 8/28/2026 07:36 AM, Jaroslaw Rafa via Postfix-users wrote:
> AI does not copy/paste code it finds somewhere on the Internet.
> It doesn't work like that (someone already mentioned this here).
> Being a probabilistic machine, it's hard to even expect that it
> copies anything literally.

Sorry to send one additional message. I initially just asked about what the policy is and didn't want to get involved otherwise. However, I wanted to leave some sources regarding this one point just in case it helps. (I will not respond further since I don't want to debate.)

Case study 1: https://blog.gdeltproject.org/do-llms-truly-create-or-merely-arrange-just-how-much-of-an-llms-writing-is-original/

Quote: "The differences between human and machine-generated text overlap support the image of LLMs as more "arrangers" than "creators" of text."

Case study 2: https://lcamtuf.substack.com/p/large-language-models-and-plagiarism

Quote: "Our results suggest that (1) three types of plagiarism widely exist in LMs beyond memorization, (2) [...] These findings overall cast doubt on the practicality of current LMs in mission-critical writing tasks [...]"

Study exploring the plagiarism rate: https://dl.acm.org/doi/10.1145/3543507.3583199 (Seems to be around 2-5%+)

Another newer study exploring how newer LLMs seem to still copy: https://www.sciencedirect.com/science/article/pii/S2949719123000213#b7

Quote: "In this work we explored the relationship between discourse quality and memorization for LLMs. We found that the models that consistently output the highest-quality text are also the ones that have the highest memorization rate."

Study testing out 2026 safeguards to reduce direct copies: https://arxiv.org/abs/2601.02671

Quote: "In some cases, jailbroken Claude 3.7 Sonnet outputs entire books near-verbatim [...] Taken together, our work highlights that, even with model- and system-level safeguards, extraction of (in-copyright) training data remains a risk for production LLMs."

Apple study showing LLMs apparently can't reason on their own: https://www.forbes.com/sites/corneliawalther/2025/06/09/intelligence-illusion-what-apples-ai-study-reveals-about-reasoning/ (Which I feel like might prompt questions how they would give correct responses without copying.)

Recent high-profile plagiarism incident regarding math proof: https://xcancel.com/ValerioCapraro/status/2097791836269977996

High-profile plagiarism incident regarding visual graph: https://www.pcgamer.com/software/ai/microsoft-uses-plagiarized-ai-slop-flowchart-to-explain-how-github-works-removes-it-after-original-creator-calls-it-out-careless-blatantly-amateuristic-and-lacking-any-ambition-to-put-it-gently/

Plagiarism incident shown with code, triggered by "function isEven", with lawyer commentary: https://github.com/mastodon/mastodon/issues/38072#issuecomment-4105681567

I hope somebody finds this helpful. Sorry again for the noise.

Regards,

Ellie
_______________________________________________
Postfix-users mailing list -- [email protected]
To unsubscribe send an email to [email protected]

Reply via email to