I have used the new template below to show how it looks. We can add more details to the linked CONTRIBUTING.md in the pop repo.
I have not included any LLM (also known as AI) generated content in...
I don’t disagree that it is a risk, and I am trying to move toward projects that do have err on the side of avoiding that risk. I have NetBSD on my laptop, and when I get a little more comfortable with it, I intend to convert the other Linux installations I maintain.
BUT, I believe the BSDs already went through a situation where some of their source was possibly under restrictive copyright and rather than “rolling back”, they “simply” identified the possibly infringing code and re-wrote those sections to have the same function (which can’t be copyrighted) without sharing any creative expression (which is). So, even the projects that are taking the risk that an LLM (or other generative AI) generates infringing code might not have quite as much cleanup / lost effort as you describe.
Also, LLM out isn’t automatically a derivative work of the training data. I’d have to dig through some other messages to find an exact quote from their documents, but I believe they (EDIT: the U.S. copyright office) said only output that is “significantly similar” to training data is potentially infringing. That does further limit the risk.
I still think it’s too high of a risk because well-meaning contributors might incorrectly introduce infringing code, since for models that don’t disclose their training data (Claude, Copilot, Gemini, etc.) even dedicated contributors don’t have the information they need to discover the output is infringing. In that past, that result (introducing infringing code) was generally limited to the acts of malicious actors that are submitting code they know to be infringing to poison a project and open it to legal action.
But, I can’t ask that someone (i.e. a project maintainer) substitute my risk/reward judgement for theirs, and I have no experience maintaining a large project. All of my code contributions are to either projects others maintain, or my own hobby projects that I doubt have any users other than myself (and I don’t even use all the published/available ones anymore).
I don’t disagree that it is a risk, and I am trying to move toward projects that do have err on the side of avoiding that risk. I have NetBSD on my laptop, and when I get a little more comfortable with it, I intend to convert the other Linux installations I maintain.
BUT, I believe the BSDs already went through a situation where some of their source was possibly under restrictive copyright and rather than “rolling back”, they “simply” identified the possibly infringing code and re-wrote those sections to have the same function (which can’t be copyrighted) without sharing any creative expression (which is). So, even the projects that are taking the risk that an LLM (or other generative AI) generates infringing code might not have quite as much cleanup / lost effort as you describe.
Also, LLM out isn’t automatically a derivative work of the training data. I’d have to dig through some other messages to find an exact quote from their documents, but I believe they (EDIT: the U.S. copyright office) said only output that is “significantly similar” to training data is potentially infringing. That does further limit the risk.
I still think it’s too high of a risk because well-meaning contributors might incorrectly introduce infringing code, since for models that don’t disclose their training data (Claude, Copilot, Gemini, etc.) even dedicated contributors don’t have the information they need to discover the output is infringing. In that past, that result (introducing infringing code) was generally limited to the acts of malicious actors that are submitting code they know to be infringing to poison a project and open it to legal action.
But, I can’t ask that someone (i.e. a project maintainer) substitute my risk/reward judgement for theirs, and I have no experience maintaining a large project. All of my code contributions are to either projects others maintain, or my own hobby projects that I doubt have any users other than myself (and I don’t even use all the published/available ones anymore).