Top
Best
New

Posted by mfiguiere 12 hours ago

How Claude marks AI-generated content(support.claude.com)
185 points | 145 comments
Dilettante_ 1 hour ago|
I notice the "Limitations" section talks about how content only at some point touched by Claude may return a positive, and content that returns a negative may still be Claude generated. But I really would have liked for them to state explicitly that entirely false positives where a piece is fully human-written may still be marked as generated, because too many institutions with the power to ruin someone's life over that have trouble understanding the concept.
sureMan6 1 hour ago|
Maybe it has no false positive rate
embedding-shape 40 minutes ago|||
Read said section yourself perhaps.
suddenlybananas 1 hour ago|||
That's essentially impossible, unless you mean they didn't measure a false positive rate.
Filligree 15 minutes ago||
For watermarked long-form text, it is actually possible. Makes the watermark more fragile, but the math is considerably more forgiving than usual.
embedding-shape 10 minutes ago||
> For watermarked long-form text

What about "watermarked long-form code"? I'm having a hard time understanding how a model could watermark not prose, but functional/semantic text like code, that actually has meaning. You can't switch our the characters, you can't use various types of whitespace, you can't add arbitrary code comments, and a lot of other restrictions. Is there any state of the art methods for watermarking code without affecting the quality/correctness?

simonw 12 hours ago||
> When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won’t see it, and it doesn’t change the meaning, quality, or readability of Claude’s response.

I'd like to know a lot more about how that works.

A lot of my interactions with Claude return pretty precise text. If I ask it to edit a project and refactor a specific function in several places I know exactly what I want to happen, it will NOT be OK if those refactors have some kind of weird pattern baked into their text to act as a watermark.

I guess this may be covered by this:

> Content generated by Claude may not carry a detectable mark if, for example: [...] The passage is very short, leaving too little text for a reliable signal;

COAGULOPATH 6 hours ago||
>I'd like to know a lot more about how that works.

My guess is that it works like Gemini's SynthID: by altering the logprobs of the next token.

Like, for every 10th token, instead of outputting the most probable, it outputs the 17th most probable, or something. (Obviously it's way more complicated but I think conceptually this is how it works.) No human will notice this, but a classifier trained on Claude's output will.

So it's not like the watermark is the words "le epic bacon" and Claude will output "le epic bacon" in everything. That would be extremely annoying (and easy to defeat).

thunfischtoast 3 hours ago|||
They still need to choose when to do that though. When I prompt the program to e.g. alter a bash script in a specific way or to recite a longer known text it can't go round and randomly exchange tokens. It has to somehow define what is a simple repeated text from a different origin and what is a novel generation.
user43928 3 hours ago||
I am wondering how that applies to newly generated code.

Odd variable naming? Stylistic choices that are watermarked?

Or as someone else noted further down in the comments, it could be more subtle:

Between the first and second most likely choice, in certain positions it will consistently choose in a certain way.

melvinroest 2 hours ago||
> Odd variable naming? Stylistic choices that are watermarked?

Whatever it is, I'm sure it's load-bearing.

asdfsa32 1 hour ago||
You're absolutely right. But it is not just load-bearing, it is the load-bearing seams.
ebtebt 52 minutes ago||||
Just double checking my understanding: If this is true then only Anthropic will be able to detect if text was generated by one of its models, correct?
ForHackernews 40 minutes ago||
Likely yes.
DanielHB 30 minutes ago||
But what prevents someone from using Anthropic own detection system to train a watermark-scrubber?

Seems like this would only catch the most unsophisticated cases.

lozenge 1 hour ago|||
Most probable usually means for a specific prompt. How can this operate without the the original prompt?
infinite_spin 49 minutes ago|||
My guess is it will be similar to how Genius watermarked lyrics, using things like variants of punctuation

https://www.pcmag.com/news/genius-we-caught-google-red-hande...

toxik 33 minutes ago||
In program code? Unlikely, surely¡
akozak 12 hours ago|||
Most likely this method https://arxiv.org/pdf/2301.10226 (EDIT: and Google's SynthID paper which builds on it https://www.nature.com/articles/s41586-024-08025-4)
iamherrylok 58 minutes ago|||
If different model providers use different green logits, does that mean they can only tell if the text came from their own model?
hannasanarion 11 hours ago|||
That "just add a constant to the green logits" as a fix to the entropy problem is so elegant I love it.
sixtyj 3 hours ago|||
It was quick :) … https://claudewatermarkremover.app/
shinryuu 2 hours ago||
Though if pangram should be trusted, there are still statistical artifacts that tells you that a text LLM generated. I don't find that to be implausible.
timpera 53 minutes ago||
Alas, Pangram should not be trusted.
baq 12 hours ago|||
> I'd like to know a lot more about how that works.

Count load-bearing words using two different algorithms in a belt-and-braces fashion

r_lee 2 hours ago|||
One thing worth flagging: those words are load-bearing
jbs789 12 hours ago||||
Fair - I should have been honest about the watermark.
seamlessdev 12 hours ago||||
Belt, braces, and suspenders.
isoprophlex 2 hours ago||
Don't forget the suppositories
quintu5 1 hour ago||
This is why I never use max effort! I’ll stick with my suspenders, thank you.
Razengan 1 hour ago||||
You’re absolutely right. Yo momma is doing a lot of heavy lifting here. Her load-bearing methods have the right shape.
dd8601fn 3 hours ago|||
That’s the real shape of the problem.
miohtama 4 hours ago|||
Maybe there is a reason why Opus 5 produces such word salad conversations
w_for_wumbo 12 hours ago|||
What happens if someone handwrites a Claude output, then someone uses that handwritten text as a reference. Now you've got a watermarked idea which may have no direct linkage to the usage of Claude.
dns_snek 3 hours ago|||
Are you worried about being accused of using LLMs to generate your work? As long as you don't plagiarize you have nothing to worry about.
Cthulhu_ 2 hours ago|||
I'm not too sure about that, people making stuff have already gotten penalized by overzealous AI detectors, most recently Kurtzgesagt.
platinumrad 1 hour ago||||
You can't make a blanket statement like this without knowing how the watermark is implemented.
AlecSchueler 2 hours ago|||
What if I unknowingly read content written by Claude in various articles and it influences my own writing style?
phainopepla2 11 hours ago|||
How is that different from referencing digital text that someone copied and pasted from Claude?
w_for_wumbo 7 hours ago||
Because there's an expectation of authenticity from the written word. If you've referenced something handwritten, you don't expect it to be the output of an LLM.

Similarly, if you quote someone word-for-word, you wouldn't anticipate their words to be flagged as Claude content, but if someone memorized Claude output word-for-word. That would still be classified as a Claude output.

Going forward you could categorize the influence of Claude on a population based off a percentage match between their spoken words with the LLM prose.

stabbles 12 hours ago|||
It will just thread some load-bearing seams through the paragraphs.
mihaelm 12 hours ago|||
> have some kind of weird pattern baked into their text to act as a watermark.

public abstract class BaseAnimalBeanFactoryGeneratedFromClaudeFactory

gajus 12 hours ago|||
Most likely watermark will be proportional to the input/output ratio, i.e. if you input a long document and ask to make edits, it will not attempt to watermark it. On the other hand, if you provide a tweet and ask it to write an article, that will include watermark. Just a guess (and yes, it feels flawed)
nprateem 3 hours ago|||
Load-bearing==claude
siva7 11 hours ago||
I can tell you how: Claude produces a huge wall of text with jargon ridden bullshit and invented terms no human subject matter expert would seriously use and overuse.
benrow 12 hours ago||
I've heard that this kind of watermarking process works by biassing the statistical sampling towards a partition of the set of possible next tokens (red set and green set), at each position. It might only be a slight nudge each time, but over a sequence of tokens, the likelihood of repeating the bias by chance is increasingly improbable.

The bias is different for each position and follows a defined RNG, seeded somehow predictably.

Can be either an open algorithm, or not. If not open, then an API could be provided to determine if text is watermarked or not.

How it applies to code - maybe it could be a subtle nudge to symbol names, etc, I'm just speculating (I only read about this in passing very recently).

londons_explore 18 minutes ago||
The bias has to be small enough that if you ask an LLM to repeat some passage of text like the national anthem, either from the training data or from the prompt it doesn't change random words.

Gotta be hard to tune that.

m-chrzan 11 hours ago|||
There's a computerphile video (https://www.youtube.com/watch?v=XZJc1p6RE78) with Dr. Mark Pound explaining a paper by John Kirchenbauer, Jonas Geiping et al. (https://arxiv.org/abs/2301.10226) that described a method for watermarking LLM output like this. It's not directly stated anywhere in the Claude support article that this is what they're using, but the properties of the watermark described seem to point to this method.
metalcrow 9 hours ago|||
Based on my understanding, it can only be applied to code in very limited ways: docstrings, variable names, string literals. The code itself can't really have tokens changed to another equally correct token (the foundation of the watermark) because then the code breaks! And the few places that you can do so are likely erased by formatters anyway.
cassianoleal 11 hours ago|||
> a defined RNG, seeded somehow predictably

So, an NG?

IshKebab 11 hours ago||
If it's based on position mod 2, wouldn't inserting or deleting (or splitting/merging) words every now and then trivially defeat it?

If it is based on position mod 2 then wouldn't inserting/deleting (or splitting and merging) words every now and then defeat it?

jonplackett 2 hours ago||
We need to just stop pretending we can reliably tell if plain text is written by an LLM.

It’s just not a reasonable ask.

JohnKemeny 51 minutes ago|
True, but what you can do is a one-sided guarantee. If it bears the mark, it is likely generated (or someone deliberately made it look generated).

Thus, if a news article, research article, book, student paper submission, blog post, HN comment, etc, bears the mark, it could be automatically flagged as such.

It helps detect low effort slop.

---

Caveat. If you write your own create work and send it to Claude for "cleaning up grammar", it might insert the watermark.

DanielHB 25 minutes ago|||
It seems like it would be so low effort to bypass, especially when you can just train a system (maybe even another LLM) using the watermarker validation from Anthropic themselves.

It seems it would get as simple as:

  outputText = promptLLM(prompt)
  scrubbedText = scrubWatermark(outputText)
Might help with students and low-technical people passing off work as their own, but any industrial scale slop-generator should be able to bypass it trivially.
asnelt 21 minutes ago|||
> If it bears the mark, it is likely generated (or someone deliberately made it look generated).

One could even say, the mark is load-bearing.

padolsey 13 minutes ago||
Is this just to appease regulators? They surely know this won't work in the long run.
akersten 11 hours ago||
So my code that Claude makes, which previously was using the best (most probable) tokens for the job, will now be getting worse in random positions, to appease a voluntary EU suggestion. Love that.
neuroticnews25 2 hours ago||
They aren't using greedy decoding, there's enough randomness in sampling to swap some with independent signal.
MagicMoonlight 2 hours ago||
[dead]
aabhay 12 hours ago||
I have had a hunch for a while now that (in addition to these tools), Anthropic has actually leaned in to Claude's distinctive manner of writing since it makes the text more obviously AI generated and thus less susceptible to misuse.

That's not necessarily the same thing as a markov-style fingerprint but it could be a correlated factor.

pjm331 12 hours ago||
I had a similar thought but I assumed they leaned in because it improved performance on coding or something like that
kingstnap 12 hours ago|||
It could also partly be a byproduct of examples of claude writing being in the dataset, which of course anthropic has lots and lots of and they do train on.
r_lee 2 hours ago||
no way. there's just no good excuse for why "load-bearing" and "worth flagging" are everywhere now, I've pretty much never seen that in the wild before
LoganDark 12 hours ago|||
I suspect it's because of alignment concerns. The more deeply they can integrate their principles, the harder it'll be to misuse. Or at least that's the idea.
noman-land 9 hours ago||
It's pretty trivial to command it to not speak that way. That's one of the first things you should write into the prompt. What style you want it to write in. Make it use a very concise and dry academic style with no overt LLMisms, melodramatic or flowery language, or metacommentary.
ethin 12 hours ago||
Can someone help me understand how exactly this watermarking of text works?

Given that text is, well, text, and not some kind of binary format, I don't see how any watermarking can work unless you insert characters which are invalid under Unicode. I further don't really understand how this won't be perceivable by assistive technology (the "watermark" will just appear as either unreadable characters, or if the watermark is mixed thoroughly enough into the text, it will scramble the text to any speech synthesizer and will make it really really obvious). Thus, I don't see how this wouldn't be insanely trivial to remove. And this is before we get into things being put on the clipboard. Sure, I can press the "Copy" button at the end of each response, but what I can also do is manually select the response and copy it, or only copy partial selections, or any number of other things. How does this "watermark" (or any "watermark" technology) take into account this?

So, really, to summarize this: I see no way of this actually being technologically achievable unless we revise the very core of how computers work and encodings for textual information. So I'm very curious as to how this is actually supposed to work.

benrow 12 hours ago||
Have a look around for token biasing, or green lists. It's based on a nudge to the choice of the next token (which can always be drawn from a set of possibilities which are all probable enough).

At first I thought this approach was just the "LLM flavour" of writing, but it's way more subtle, especially as the bias is applied uniquely for each token position.

ethin 11 hours ago||
Yeah, will do, this sounds interesting since I'm not entirely sure how this would actually be reliable to any degree. Thanks for the help, not sure why I got downvoted since I was genuinely curious.
resonantjacket5 12 hours ago|||
it's a statistical way. like for example maybe in your above paragraph claude maybe writes "Thus, I don't see how this wouldn't be insanely <easy>(instead of trivial) to remove" and then also says like "And this is before we <analyze> things being put on the clipboard." or maybe the i just says the word "the" in a certain pattern or frequency.

you can then consistently like figure out if it was claude that wrote the sentence. it is easy as you noted if you just get another ai to read it and then rewrite it.

MagicMoonlight 2 hours ago||
[dead]
simonw 12 hours ago||
An interesting factor of this is competition.

If Claude was the only model family they could ship a change like this and users who want to cheat (or don't like watermarks for other reasons) would just have to put up with it.

In a world with many different competing models, the risk of losing customers to other providers over this is much more real.

Maybe they've looked at the numbers and the portion of people who clearly use Claude to cheat on examples etc is so tiny that losing them to other providers isn't a problem?

lhd1 5 hours ago||
Scott Aaronson spoke about this in a colloquium where he said that this was mooted at OpenAI before the decision was made by Altman to not implement it for the reasons you describe.

https://youtu.be/9udWn1Hlj_s?si=VWOiK5-y4zcyDoHI

reasonableklout 2 hours ago||
OpenAI will soon be adding watermarking to text as well, as it signed the EU Code of Practice on Transparency of AI-Generated Content: https://openai.com/index/advancing-responsible-ai-across-eur...
slowin 10 hours ago|||
I’m more worried that this will degrade performance. I want the best results from a model, not the results that fit a constraint that’s not defined by me. Any increased cost or latency is also unacceptable.
gajus 12 hours ago|||
There are already small models trained specifically to prevent statistical detection, e.g., https://huggingface.co/kalpeshk2011/dipper-paraphraser-xxl

I guess whoever is the policy maker is assuming that some protection is better than none and that most people will not reach for such tools.

nprateem 3 hours ago||
Either that or they want to comply with the EU AI Act when it affects them.
izonu 12 hours ago|
> We’re also working to enable users and other third parties to detect Claude’s embedded watermarks and provenance metadata.

This seems to be similar in execution to Google's SynthID. I hope they release actual code the technically proficient can use, unlike SynthID which can only (afaik) be queried with Gemini's UI.

gajus 12 hours ago|
The moment Google announced SynthID, the first domain I bought was deSynthID.com

Several open-source projects have already proven SynthID to be ineffective.

moebrowne 3 hours ago||
There are many free lock picking tutorials, but yet locks are still effective.
More comments...