Fingers crossed Gnome follows suit! :)

  • Peasley@lemmy.world
    link
    fedilink
    arrow-up
    43
    arrow-down
    3
    ·
    edit-2
    17 小时前

    i think the copyright question is still unanswered: it could turn out that any LLM-generated code is a copyright violation by definition unless trained exclusively on a clean, legitimately-obtained dataset (which few of the major models are).

    Will projects that allow LLM contributions have to roll back years of progress when the other shoe finally drops? Seems like a huge risk, especially for FOSS and copyleft. I think disallowing LLM-written contributions until this is all sorted out in the courts is the pragmatic move from a legal perspective.

    • bss03@infosec.pub
      cake
      link
      fedilink
      English
      arrow-up
      12
      ·
      edit-2
      14 小时前

      I don’t disagree that it is a risk, and I am trying to move toward projects that do have err on the side of avoiding that risk. I have NetBSD on my laptop, and when I get a little more comfortable with it, I intend to convert the other Linux installations I maintain.

      BUT, I believe the BSDs already went through a situation where some of their source was possibly under restrictive copyright and rather than “rolling back”, they “simply” identified the possibly infringing code and re-wrote those sections to have the same function (which can’t be copyrighted) without sharing any creative expression (which is). So, even the projects that are taking the risk that an LLM (or other generative AI) generates infringing code might not have quite as much cleanup / lost effort as you describe.

      Also, LLM out isn’t automatically a derivative work of the training data. I’d have to dig through some other messages to find an exact quote from their documents, but I believe they (EDIT: the U.S. copyright office) said only output that is “significantly similar” to training data is potentially infringing. That does further limit the risk.

      I still think it’s too high of a risk because well-meaning contributors might incorrectly introduce infringing code, since for models that don’t disclose their training data (Claude, Copilot, Gemini, etc.) even dedicated contributors don’t have the information they need to discover the output is infringing. In that past, that result (introducing infringing code) was generally limited to the acts of malicious actors that are submitting code they know to be infringing to poison a project and open it to legal action.

      But, I can’t ask that someone (i.e. a project maintainer) substitute my risk/reward judgement for theirs, and I have no experience maintaining a large project. All of my code contributions are to either projects others maintain, or my own hobby projects that I doubt have any users other than myself (and I don’t even use all the published/available ones anymore).

    • esc@piefed.social
      link
      fedilink
      English
      arrow-up
      2
      arrow-down
      1
      ·
      10 小时前

      It’s answered, unless all big tech is going down suddenly, they are allowed, copyright violations are for the poor anyway.

    • Skullgrid@lemmy.world
      link
      fedilink
      arrow-up
      9
      arrow-down
      6
      ·
      16 小时前

      i think the copyright question is still unanswered: it could turn out that any LLM-generated code is a copyright violation by definition unless trained exclusively on a clean, legitimately-obtained dataset (which few of the major models are).

      I think this argument is BS, there are several remix/sample based albums that count as derived works and AFAIK, no one is getting paid.

      • Peasley@lemmy.world
        link
        fedilink
        arrow-up
        9
        ·
        14 小时前

        Not exactly the same, and the music industry has had plenty of lawsuits going both ways on that kind of thing establishing a status quo for remixes and samples in music

      • bss03@infosec.pub
        cake
        link
        fedilink
        English
        arrow-up
        7
        ·
        edit-2
        13 小时前

        IBM Granite (EDIT: and Apertus) do disclose all their training data and claim that all their training data is effectively free of copyright (highly permissively licensed). I have not been able to verify that, due to my lack of skills with the conventions and tools of LLM / Agent training and publishing.

        So, yeah, probably (EDIT: two one).

          • bss03@infosec.pub
            cake
            link
            fedilink
            English
            arrow-up
            4
            ·
            edit-2
            13 小时前

            Thank you for the link! It does look like Apertus itself might be Free Software (the U.S. copyright office says training can infringe, but is usually fair use), but it can still output derivative works of copyrighted inputs that might prevent them from being distributed as-is (for example, requiring attribution) – at all, much less under a strong copyleft.