I have used the new template below to show how it looks. We can add more details to the linked CONTRIBUTING.md in the pop repo.
I have not included any LLM (also known as AI) generated content in...
There’s some serious concerns with corporate AI that, IMO, make it a no-go under almost any scenario.
Corporate AI Concerns
Simply using an AI that’s run on a corporate data centers encourages the construction of yet more data centers, with all of the environmental/climate negatives they bring, as well as local harms they induce on the people living near them, such as increased electricity rates.
Using corporate AI directly helps the financial situations of those giant corporations (by boosting usage/user numbers, they are able to attract more investment capital), most of which are ran by right-wing CEOs who are more than willing to collaborate with and fund fascist governments to ensure that they are not regulated in search of both maximum profits. Some of these companies, such as Nvidia, Palantir and Oracle, genuinely appear to be seeking to use these tools for what would previously be considered crackpot conspiracy theory levels of public control and surveillance.
But if we ignore those and instead focus purely on self-hosted AI, we’re still going to run into some roadblocks :\
Unfortunately, there does not currently exist a self-hosted LLM model that wasn’t trained unethically on copyrighted code (the ones that do claim to be ethical have not lived up to that claim).
Realistically, even with an unethically trained local LLM, I don’t think using the available self-hosted LLMs to try to find bugs in some code is terribly unethical or damaging, that’s probably the only use-case I would sign-off on within a FLOSS project. I also don’t think it would matter too much if someone wanted to use a self-hosted AI to code some changes to a personal project, like modding an old C64 game source code with some new features. Go for it.
But once we start getting into distributing software, or contributing code to existing FLOSS projects, none of the current models are safe to use, as they all are prone to generating parts of the copyrighted code they were trained on (potentially each prompt can have a 3 to 10% chance of generating a perfect copy of copyrighted code). That is a legal minefield waiting to be stepped on, and could be disastrous for projects that accept a large portion of LLM code, as it may become unfeasible to re-write that code, or to survive a legal battle from a company claiming damages.
It’s also quite unsettled on whether or not LLM output can even be copyrighted at all, stripping any LLM generated code of the ability to be protected with copyleft licenses like the GPL (this would in effect allow companies to take FLOSS project and close source it, using the free labor of the project, but without any ability for the FLOSS project to do the same back, it would be a one-way pillaging).
Until there’s a a truly ethical and non-infringing self-hosted LLM available, there really is no middle ground for a FLOSS project to take. And if/when one is available, there will still be a concern of skill loss/critical thinking atrophy that seems to be likely with extended LLM use, like a sort’ve Wall-E effect (example 1, example 2).
There’s some serious concerns with corporate AI that, IMO, make it a no-go under almost any scenario.
Corporate AI Concerns
Simply using an AI that’s run on a corporate data centers encourages the construction of yet more data centers, with all of the environmental/climate negatives they bring, as well as local harms they induce on the people living near them, such as increased electricity rates.
Using corporate AI directly helps the financial situations of those giant corporations (by boosting usage/user numbers, they are able to attract more investment capital), most of which are ran by right-wing CEOs who are more than willing to collaborate with and fund fascist governments to ensure that they are not regulated in search of both maximum profits. Some of these companies, such as Nvidia, Palantir and Oracle, genuinely appear to be seeking to use these tools for what would previously be considered crackpot conspiracy theory levels of public control and surveillance.
But if we ignore those and instead focus purely on self-hosted AI, we’re still going to run into some roadblocks :\
Unfortunately, there does not currently exist a self-hosted LLM model that wasn’t trained unethically on copyrighted code (the ones that do claim to be ethical have not lived up to that claim).
Realistically, even with an unethically trained local LLM, I don’t think using the available self-hosted LLMs to try to find bugs in some code is terribly unethical or damaging, that’s probably the only use-case I would sign-off on within a FLOSS project. I also don’t think it would matter too much if someone wanted to use a self-hosted AI to code some changes to a personal project, like modding an old C64 game source code with some new features. Go for it.
But once we start getting into distributing software, or contributing code to existing FLOSS projects, none of the current models are safe to use, as they all are prone to generating parts of the copyrighted code they were trained on (potentially each prompt can have a 3 to 10% chance of generating a perfect copy of copyrighted code). That is a legal minefield waiting to be stepped on, and could be disastrous for projects that accept a large portion of LLM code, as it may become unfeasible to re-write that code, or to survive a legal battle from a company claiming damages.
It’s also quite unsettled on whether or not LLM output can even be copyrighted at all, stripping any LLM generated code of the ability to be protected with copyleft licenses like the GPL (this would in effect allow companies to take FLOSS project and close source it, using the free labor of the project, but without any ability for the FLOSS project to do the same back, it would be a one-way pillaging).
Until there’s a a truly ethical and non-infringing self-hosted LLM available, there really is no middle ground for a FLOSS project to take. And if/when one is available, there will still be a concern of skill loss/critical thinking atrophy that seems to be likely with extended LLM use, like a sort’ve Wall-E effect (example 1, example 2).