On 8/28/2026 07:36 AM, Jaroslaw Rafa via Postfix-users wrote: > AI does not copy/paste code it finds somewhere on the Internet. > It doesn't work like that (someone already mentioned this here). > Being a probabilistic machine, it's hard to even expect that it > copies anything literally.
Sorry to send one additional message. I initially just asked about what the policy is and didn't want to get involved otherwise. However, I wanted to leave some sources regarding this one point just in case it helps. (I will not respond further since I don't want to debate.)
Case study 1: https://blog.gdeltproject.org/do-llms-truly-create-or-merely-arrange-just-how-much-of-an-llms-writing-is-original/
Quote: "The differences between human and machine-generated text overlap support the image of LLMs as more "arrangers" than "creators" of text."
Case study 2: https://lcamtuf.substack.com/p/large-language-models-and-plagiarism
Quote: "Our results suggest that (1) three types of plagiarism widely exist in LMs beyond memorization, (2) [...] These findings overall cast doubt on the practicality of current LMs in mission-critical writing tasks [...]"
Study exploring the plagiarism rate: https://dl.acm.org/doi/10.1145/3543507.3583199 (Seems to be around 2-5%+)
Another newer study exploring how newer LLMs seem to still copy: https://www.sciencedirect.com/science/article/pii/S2949719123000213#b7
Quote: "In this work we explored the relationship between discourse quality and memorization for LLMs. We found that the models that consistently output the highest-quality text are also the ones that have the highest memorization rate."
Study testing out 2026 safeguards to reduce direct copies: https://arxiv.org/abs/2601.02671
Quote: "In some cases, jailbroken Claude 3.7 Sonnet outputs entire books near-verbatim [...] Taken together, our work highlights that, even with model- and system-level safeguards, extraction of (in-copyright) training data remains a risk for production LLMs."
Apple study showing LLMs apparently can't reason on their own: https://www.forbes.com/sites/corneliawalther/2025/06/09/intelligence-illusion-what-apples-ai-study-reveals-about-reasoning/ (Which I feel like might prompt questions how they would give correct responses without copying.)
Recent high-profile plagiarism incident regarding math proof: https://xcancel.com/ValerioCapraro/status/2097791836269977996
High-profile plagiarism incident regarding visual graph: https://www.pcgamer.com/software/ai/microsoft-uses-plagiarized-ai-slop-flowchart-to-explain-how-github-works-removes-it-after-original-creator-calls-it-out-careless-blatantly-amateuristic-and-lacking-any-ambition-to-put-it-gently/
Plagiarism incident shown with code, triggered by "function isEven", with lawyer commentary: https://github.com/mastodon/mastodon/issues/38072#issuecomment-4105681567
I hope somebody finds this helpful. Sorry again for the noise. Regards, Ellie _______________________________________________ Postfix-users mailing list -- [email protected] To unsubscribe send an email to [email protected]
