Is it April 1st again already? That sounds like BS when applied to text. For code, this doesn’t work at all. A linter will kill any statistical shenanigans they would introduce. And if they have a mechanism whatsoever to see the watermark, then others will see it too, which will make it easy to remove it.
EDIT there is a paper about how this is supposed to work. They acknowledge that an attack on the algorithm can only be prevented when the algorithm is not disclosed to the public. They also explain how it is basically useless to detect AI generated code. And on top of it all, one would have to force the same non-public watermarking algorithm on every LLM out there to be effective. Which seems unlikely to happen. While it might be a fancy way for an owner of an LLM to detect whether their product was used, it is totally useless as a measure of detecting AI generated texts in general.
Is it April 1st again already? That sounds like BS when applied to text. For code, this doesn’t work at all. A linter will kill any statistical shenanigans they would introduce. And if they have a mechanism whatsoever to see the watermark, then others will see it too, which will make it easy to remove it.
EDIT there is a paper about how this is supposed to work. They acknowledge that an attack on the algorithm can only be prevented when the algorithm is not disclosed to the public. They also explain how it is basically useless to detect AI generated code. And on top of it all, one would have to force the same non-public watermarking algorithm on every LLM out there to be effective. Which seems unlikely to happen. While it might be a fancy way for an owner of an LLM to detect whether their product was used, it is totally useless as a measure of detecting AI generated texts in general.