Summary
From the article:
Scott Aaronson, while working at OpenAI, largely solved AI text watermarking together with Hendrik Kirchner.
Here is how his solution works, or see Tenobrus’s version.
AI outputs are not deterministic. The AI’s job is to pick the probability of each potential next token. The token is then chosen at random.
By default you use a source of pseudo-randomness for each choice, since actual true randomness is annoying.
To apply the watermark, you use an otherwise identical private source of pseudo-randomness derived from a secret key.
Then, given enough text, a score is derived for howe well the choices fit with that particular pseudo-randomness source, versus a different source.
You provide an API that lets anyone check for the watermark.
[...]
Google implemented this, including for Gemini 3.7 Flash, and they have been rolling out this feature since 2024. Google has done, for over two years, the exact thing Anthropic is now doing, except with a public detector, and Google confirmed in a test (n = 20 million) that there is no difference in user feedback.
Anthropic quietly announced a week ago they were rolling out watermarking to comply with the EU Code of Practice. Since they don’t want to have to differentiate traffic sources, the marginal cost is zero, and watermarking is pro-social, this will apply to everyone. They then offered an FAQ of how it works.
[...]
The entire practical effect is: There will be an API that will tell you if a given piece of writing comes from Claude. That’s it. And yet.
[...]
The rest of this post is about exploring why people are Big Mad about this, in large part as a worked example of how people get worked up over approximately nothing.