As the others explained, it’s mainly statistics regarding usage of specific words. This also means anyone can make a tool that uses a thesaurus to mess with the watermark (these are already available online - they have the disclaimer that they can’t guarantee that it works every time tho), and short texts might not provide enough data for the watermark to be reliably identified at all.
deleted by creator
As the others explained, it’s mainly statistics regarding usage of specific words. This also means anyone can make a tool that uses a thesaurus to mess with the watermark (these are already available online - they have the disclaimer that they can’t guarantee that it works every time tho), and short texts might not provide enough data for the watermark to be reliably identified at all.
If someone’s interested in more details: https://youtu.be/BFksx2M93sw
Based on “controlling” of a sampling set for a token prediction . Nice summary is here https://sebastianraschka.com/blog/2026/claude-text-watermarking.html eg
This is also a really good explainer: https://declaude.org/watermarking/
An algorithm that prefers certain words over others, for specific inputs. The same way we encrypt data really.
A pattern of invisible characters perhaps.