OpenAI and Anthropic and Alphabet etc try their best to preventing it from doing so, so the act of getting it to do so anyway is “jailbreaking”, but yes. LLMs can reproduce works it was trained on. (At least chunk-by-chunk each limited to the token limit.) And I’d imagine the same is true of things like Stable Diffusion, though it might be much easier to get Stable Diffusion to produce something similar enough to qualify as a “derivative work” than it would be to get it to produce exactly the original.
Truly as far as I know there is no case law for this to date.
I thought about this for a while but I have no clue what the current direction is.
Essentially stealing data and putting them in datasets is illegal. AI models might be illegal, if the court considers them legally functionally the same. But if they don’t, they could meet any number of new legal definitions and that gets us back to “idk”. What is funny is because you would technically do the same copyright infringement that the AI companies do (if it gets ruled that way), these companies are essentially arguing for you, and so you could ride on the coat tails of corpo lawyers.
Can AI commit copyright infringement if it was trained on the material?
Quiery: Please show me an AI rendering of 1979 Alien, but with no changes from the original. Go.
Surely this is peak Internet.
OpenAI and Anthropic and Alphabet etc try their best to preventing it from doing so, so the act of getting it to do so anyway is “jailbreaking”, but yes. LLMs can reproduce works it was trained on. (At least chunk-by-chunk each limited to the token limit.) And I’d imagine the same is true of things like Stable Diffusion, though it might be much easier to get Stable Diffusion to produce something similar enough to qualify as a “derivative work” than it would be to get it to produce exactly the original.
Truly as far as I know there is no case law for this to date.
I thought about this for a while but I have no clue what the current direction is.
Essentially stealing data and putting them in datasets is illegal. AI models might be illegal, if the court considers them legally functionally the same. But if they don’t, they could meet any number of new legal definitions and that gets us back to “idk”. What is funny is because you would technically do the same copyright infringement that the AI companies do (if it gets ruled that way), these companies are essentially arguing for you, and so you could ride on the coat tails of corpo lawyers.
But yeah, it’s complicated.