I have used the new template below to show how it looks. We can add more details to the linked CONTRIBUTING.md in the pop repo.
I have not included any LLM (also known as AI) generated content in...
i think the copyright question is still unanswered: it could turn out that any LLM-generated code is a copyright violation by definition unless trained exclusively on a clean, legitimately-obtained dataset (which few of the major models are).
Will projects that allow LLM contributions have to roll back years of progress when the other shoe finally drops? Seems like a huge risk, especially for FOSS and copyleft. I think disallowing LLM-written contributions until this is all sorted out in the courts is the pragmatic move from a legal perspective.
If it turns out that it gets ruled as copyright violations, you can bet your ass they’re going to reform copyright law instead of rolling everything back.
I don’t disagree that it is a risk, and I am trying to move toward projects that do have err on the side of avoiding that risk. I have NetBSD on my laptop, and when I get a little more comfortable with it, I intend to convert the other Linux installations I maintain.
BUT, I believe the BSDs already went through a situation where some of their source was possibly under restrictive copyright and rather than “rolling back”, they “simply” identified the possibly infringing code and re-wrote those sections to have the same function (which can’t be copyrighted) without sharing any creative expression (which is). So, even the projects that are taking the risk that an LLM (or other generative AI) generates infringing code might not have quite as much cleanup / lost effort as you describe.
Also, LLM out isn’t automatically a derivative work of the training data. I’d have to dig through some other messages to find an exact quote from their documents, but I believe they (EDIT: the U.S. copyright office) said only output that is “significantly similar” to training data is potentially infringing. That does further limit the risk.
I still think it’s too high of a risk because well-meaning contributors might incorrectly introduce infringing code, since for models that don’t disclose their training data (Claude, Copilot, Gemini, etc.) even dedicated contributors don’t have the information they need to discover the output is infringing. In that past, that result (introducing infringing code) was generally limited to the acts of malicious actors that are submitting code they know to be infringing to poison a project and open it to legal action.
But, I can’t ask that someone (i.e. a project maintainer) substitute my risk/reward judgement for theirs, and I have no experience maintaining a large project. All of my code contributions are to either projects others maintain, or my own hobby projects that I doubt have any users other than myself (and I don’t even use all the published/available ones anymore).
i think the copyright question is still unanswered: it could turn out that any LLM-generated code is a copyright violation by definition unless trained exclusively on a clean, legitimately-obtained dataset (which few of the major models are).
I think this argument is BS, there are several remix/sample based albums that count as derived works and AFAIK, no one is getting paid.
Not exactly the same, and the music industry has had plenty of lawsuits going both ways on that kind of thing establishing a status quo for remixes and samples in music
IBM Granite (EDIT: and Apertus) do disclose all their training data and claim that all their training data is effectively free of copyright (highly permissively licensed). I have not been able to verify that, due to my lack of skills with the conventions and tools of LLM / Agent training and publishing.
The Apertus Swiss AI unfortunately doesn’t seem to live up to its claims, and they have been silent on the issue brought up there. I honestly suspect the same of IBM’s Granite, but have not investigated.
Thank you for the link! It does look like Apertus itself might be Free Software (the U.S. copyright office says training can infringe, but is usually fair use), but it can still output derivative works of copyrighted inputs that might prevent them from being distributed as-is (for example, requiring attribution) – at all, much less under a strong copyleft.
i think the copyright question is still unanswered: it could turn out that any LLM-generated code is a copyright violation by definition unless trained exclusively on a clean, legitimately-obtained dataset (which few of the major models are).
Will projects that allow LLM contributions have to roll back years of progress when the other shoe finally drops? Seems like a huge risk, especially for FOSS and copyleft. I think disallowing LLM-written contributions until this is all sorted out in the courts is the pragmatic move from a legal perspective.
If it turns out that it gets ruled as copyright violations, you can bet your ass they’re going to reform copyright law instead of rolling everything back.
I don’t disagree that it is a risk, and I am trying to move toward projects that do have err on the side of avoiding that risk. I have NetBSD on my laptop, and when I get a little more comfortable with it, I intend to convert the other Linux installations I maintain.
BUT, I believe the BSDs already went through a situation where some of their source was possibly under restrictive copyright and rather than “rolling back”, they “simply” identified the possibly infringing code and re-wrote those sections to have the same function (which can’t be copyrighted) without sharing any creative expression (which is). So, even the projects that are taking the risk that an LLM (or other generative AI) generates infringing code might not have quite as much cleanup / lost effort as you describe.
Also, LLM out isn’t automatically a derivative work of the training data. I’d have to dig through some other messages to find an exact quote from their documents, but I believe they (EDIT: the U.S. copyright office) said only output that is “significantly similar” to training data is potentially infringing. That does further limit the risk.
I still think it’s too high of a risk because well-meaning contributors might incorrectly introduce infringing code, since for models that don’t disclose their training data (Claude, Copilot, Gemini, etc.) even dedicated contributors don’t have the information they need to discover the output is infringing. In that past, that result (introducing infringing code) was generally limited to the acts of malicious actors that are submitting code they know to be infringing to poison a project and open it to legal action.
But, I can’t ask that someone (i.e. a project maintainer) substitute my risk/reward judgement for theirs, and I have no experience maintaining a large project. All of my code contributions are to either projects others maintain, or my own hobby projects that I doubt have any users other than myself (and I don’t even use all the published/available ones anymore).
It’s answered, unless all big tech is going down suddenly, they are allowed, copyright violations are for the poor anyway.
I think this argument is BS, there are several remix/sample based albums that count as derived works and AFAIK, no one is getting paid.
Not exactly the same, and the music industry has had plenty of lawsuits going both ways on that kind of thing establishing a status quo for remixes and samples in music
“Most” music is also under a compulsory licensing system, while virtually no code, prose, or visual art is.
SilvaGunner 🫡
Few? Do you have some example, maybe two?
IBM Granite (EDIT:
and Apertus) do disclose all their training data and claim that all their training data is effectively free of copyright (highly permissively licensed). I have not been able to verify that, due to my lack of skills with the conventions and tools of LLM / Agent training and publishing.So, yeah, probably (EDIT:
twoone).for those who said the deets, thank you!
The Apertus Swiss AI unfortunately doesn’t seem to live up to its claims, and they have been silent on the issue brought up there. I honestly suspect the same of IBM’s Granite, but have not investigated.
Thank you for the link! It does look like Apertus itself might be Free Software (the U.S. copyright office says training can infringe, but is usually fair use), but it can still output derivative works of copyrighted inputs that might prevent them from being distributed as-is (for example, requiring attribution) – at all, much less under a strong copyleft.
Few could be zero lol, i have no idea.
IBM Granite (EDIT:
and Apertus) probably qualify.