Discussion about this post

User's avatar
Simon's avatar

This article is overstating the issue in a way that it s very misleading. Text watermarking works by changing the source of randomness that is an inherent aspect of token generation. A succession of random choices that match the models own randomness have a higher probability of having been generated by that model.

By definition, then, you need a large number of such choices to have high confidence that AI made them. Anthropic themselves admit that "detecting a watermark.. doesn’t work well on small samples, where there are fewer word choices and thus less information to go on"

In other words, neither a single fixed comma nor a corrected spelling mistake is rich enough to be watermarked.

A translation of a substantial piece of text would be susceptible to watermarking, but that's entirely valid. As any translator will tell you, translation is a form of interpretation. Offloading that interpretation to AI is a significant decision that is worth flagging.

More here:

https://www.anthropic.com/news/claude-text-watermark

3 more comments...

No posts

Ready for more?