Top
Best
New

Posted by ropbear 15 hours ago

Anthropic's 'watermark' text adulteration in Claude is a perversion of writing(daringfireball.net)
370 points | 352 commentspage 2
andOlga 7 hours ago|
What a truly bizarre article. Arguments about pre-existing randomness, temperature and whatnot aside, I simply cannot comprehend what the author here really thinks the "best word" is. There's no such thing. We humans fall on familiar patterns of writing ourselves, so we may forego something with a flourish in favor of a more commonly-used word unless we put in effort to be "special", which should be used sparingly. That is to say, human writers are likely to choose a "worse" word in far more than the supposed 51% of cases, and that has no effect on the actual quality of writing in the end.

But even if there were such a thing as a truly "best word", for some context, what are the examples here? Mango vs pineapple? Gray vs overcast? In what case is one of these better, that AI would normally infer but would suddenly be "perverted" by SynthID? Do you think your emotional state and preferences are being evaluated if they aren't explicitly in memory? And if they are there, do you think that the generator will bypass those instructions in favor of the watermark instead of placing it somewhere you won't care? I just. Genuinely don't get it. There may be words that matter in specific contexts or to you as a reader, so you should bloody well put them there.

its-summertime 1 hour ago||
> My error was believing Anthropic that their system wouldn’t adulterate and corrupt the semantics of the text their models generate.

Has that ever been the case? Are they not actively tweaking their models, their fine tuning, the system prompts, the tool definitions and implementations, the guard rails, tool calls, instant responses. There are hundreds of knobs that they can change daily, or between each prompt, or even half way through a generation.

Imnimo 14 hours ago||
>I want any LLM I use to choose the very best, most precise words at every single decision point.

Does the author think he is currently getting T=0 output from Claude? Is he under the impression that T=0 produces the "best" writing?

This entire article just seems so detached from the basics of how LLMs work.

dofm 13 hours ago||
> Does the author think he is currently getting T=0 output from Claude? Is he under the impression that T=0 produces the "best" writing?

No and no. I am not sure I agree with his point but I know he is not ill-informed on either of these points, because I mentioned them to him a couple of days ago.

brookst 13 hours ago||
It didn’t take, apparently.
dofm 13 hours ago||
It’s fully possible I didn’t explain it very well in the first place, but he is making a wider point.

The point I made (quite briefly) is that watermarking is only feasible because for good writing it is necessary to use T>0, or the writing will never explore a more creative choice, and that at T=0 you don’t even need a watermark to spot LLM-generated text.

The point he is making is consistent with this, isn’t it? Either you allow temperature to drive creativity, consistently in a way that can be influenced and analysed, or you adulterate that process for the purposes of meeting a corporate/legal directive, in a way that is proprietary and obscure. These are ethically distinct approaches, and since he disagrees with the EU objective he comes down on one side I guess.

Me, I don’t care about the hypothetical enough.

Not least because I think Claude writes depressingly badly and I doubt any steganographic change will enrage me less.

wasabi991011 11 hours ago|||
The point he is making is not consistent with understanding how temperature influences LLM text generation, no.

He repeatedly states that choosing "the best word" is the most important thing to him. I don't know how you reconcile that with creativity itself, let alone probabilistic sampling.

eru 4 hours ago|||
Yes, naively understood in his sense 'best' word means that you pick the word with the maximum score, instead of sampling from the distribution.

That doesn't actually give you the 'best' text in any human sense of the word. Just like playing the 'best' move in Poker without sampling leads you to lose a lot of money.

dofm 11 hours ago|||
I mean creativity there in the LLM sense (temperature driving more creative solutions), not the human sense, and in the context of its consistency, configurability and being amenable to analysis. The watermarking approach makes that non-reproducible, yes? Because Anthropic can and will change it as they see fit.
Imnimo 10 hours ago||||
>Either you allow temperature to drive creativity, consistently in a way that can be influenced and analysed, or you adulterate that process for the purposes of meeting a corporate/legal directive, in a way that is proprietary and obscure

This framing does not make sense to me. What do you mean by "influenced and analysed"? How have you or anyone been influencing or analyzing the randomness behind the sampling process to create better writing? What makes the unadulterated randomness "driving creativity" but a different random choice uncreative?

dofm 1 hour ago||
> How have you or anyone been influencing or analyzing the randomness behind the sampling process to create better writing?

You're mischaracterising or misunderstanding my point, or I mangled it.

I mean it is possible to analyse, control, monitor, study the impact of changing temperature on the writing, yes?

The point about watermarking is that this relationship — change the temperature, see the effect — is now being adjusted by an unstated, secret process you explicitly can't control.

(I gather Anthropic have recently taken away this setting anyway; that was news to me.)

beering 12 hours ago||||
No, you are jumping to conclusions about how watermarking works. This is some audiophile thinking that because your RNG is “pure”, you get text with an expansive soundstage or whatever. Intuitively this may be true or false depending on your personal prior but you’d need to show it mathematically. The overall token distribution shouldn’t change and the frequency at which you see the word “load-bearing” will remain the same.
dofm 12 hours ago||
> This is some audiophile thinking that because your RNG is “pure”, you get text with an expansive soundstage or whatever.

You are projecting that onto me, and I cannot tell you how comically poorly aimed it is.

inigyou 12 hours ago|||
What does he think of all the other adulterations of LLMs that already happen?
Gigachad 13 hours ago||
I think the author is just mad people will be able to detect and filter out their AI slop writing in the future.
kalleboo 9 hours ago|||
Gruber just hates any kind of EU regulation of US tech companies ever since they started making what he calls "product decisions" for Apple.
llm_nerd 2 hours ago||||
Gruber has long been a talented, excellent writer. I hugely doubt he uses AI at all, nor does he plan to.

His first take on this situation was cutely naive, thinking they were going to inject secret hidden unicode characters. But ultimately he has a massive hate on for the EU -- they were mean to Apple once -- and it comes out in any topic that overlaps.

dofm 13 hours ago||||
This is not it, no. He is not using AI and it is not I think remotely in his nature to surrender that control. He is engaging with this on principle. Again I am not sure I agree with him, but then it’s a hypothetical because I am not going to get an LLM to write for me either.
NitpickLawyer 3 hours ago|||
> He is not using AI

That's ... even worse? So we're all here in the comments trying to figure out what the author means, and what their overall point is, while clearly they don't even use the damn thing? Oof... What a waste of time for everyone involved.

dofm 2 hours ago||
Why? I really don't understand this. Why can't a tech writer take a deep but neutral intellectual interest in something? Isn't it important that some do? Do people have to be stakeholders or clearly on one given team, pro- or anti-, for their opinion to matter? Is it that tribal?

It seems fully logical to me that someone who writes for a living (who, as it happens, developed the very markup language LLMs use for everything) should be invested in understanding the automatic plagiarism and word calculating machine from an intellectually honest position.

I personally am pretty severely big-two-AI-firms, increasingly anti-big-tech, but I am learning and researching uses of LLMs because for myself I really need to understand how to use them in an intellectually and (as far as is possible) ethically sound way. Learning because as a boring old freelance programmer I have to; foolish to pretend otherwise.

So I completely understand his position — that the AI industry is hot air and crooked and scammy and weird, and some of the people involved genuinely rather dark-sided, but the technology exists and if it hints at threatening your livelihood, you need to understand it.

From reading his work for the best part of twenty years or so (and emailing him intermittently over that time) it would seem to me that he's a lot less bearish on the tech industry than me, and a lot less fond of the EU than I am; he's more optimistic than I am. But he writes because he has to write. I should think that would make him highly invested in understanding what LLMs do.

breezybottom 22 minutes ago||
If he writes then he should have no investment in LLMs. Humans have been able to write for thousands of years.
beering 12 hours ago|||
Well, it cant be that he is super worried on behalf of people who publish AI slop. That’s not a credible motivation. In fact, he complained a lot about the new ChatGPT app so I can’t believe your claim that he is not using AI.

Seems like he really likes to use LLMs and is worried that quality will be degraded. But he will never demonstrate such degradation scientifically, we don’t have anecdotes even.

dofm 11 hours ago||
His complaint about the ChatGPT app is that it’s a shitty non-Mac-ish Mac app. Complaining about shitty non-Mac-ish Mac apps to people who hate shitty non-Mac-ish Mac apps is more or less how he became a full time writer.
cryptonector 8 hours ago|||
People already are, and do.
smallerize 15 hours ago||
Translation: No one can ever again use Claude for proofreading their own prose unless they’re willing to risk that the whole thing might be flagged as having been generated by Claude.

I think that was intended, yes.

epihelix 14 hours ago||
You know, back in the era when proofreaders were human, I never one met a proofreader who rewrote my text afresh, rather than annotating the text with a pen.

It's still possible to use Claude to proofread - highlight grammatical, flow, structure, logic errors and make simple suggestions for you to pick and choose or adapt as you wish. No watermarking will flag your text. No flaw accusations of LLM authorship will haunt you. All will be fine.

But if you want an LLM to rewrite your text, that's (a) not proofreading, and (b) should be flagged as LLM generated ... because it is.

wuschel 4 hours ago|||
Where is the problem with using LLM generated text?

You could use your own hypothetical house elf to do it for you, or pay someone to do it. LLMs are just cheaper for a certain set of problems.

People will find ways to circumvent this, so this limitation will only hit the technically less adept people.

zahlman 4 hours ago|||
> Where is the problem with using LLM generated text?

In the fact that you didn't write it.

> You could use your own hypothetical house elf to do it for you, or pay someone to do it.

Yes, and those would be similarly problematic (and more expensive).

ivan_gammel 3 hours ago||
> In the fact that you didn't write it.

This is a fact and this is generally not a problem. Customer support guy Joe did not write that email to you with a refund: someone else did it and Joe did pick the template. Alice did not write that post card to Bob, someone else did and she just googled some nice text. We deal with a lot of content that wasn’t written by the person who signed it. That content, when written by LLM, may indeed contain watermarks and nobody will care about the choice of words, because only the meaning matters in such communications.

People pay too much attention to authenticity here, which is no more than a demonstration of an effort. LLM can and should write scientific articles because the real effort is in directing research, not summarizing it. LLMs can and should write news, because it is cheap and efficient, and real reporting is in discovery. LLM can and should write fiction and make movies, because there is no reason why creators of various junk should earn their money easily. LLMs do not replace real talent. They just emphasize for an average person how easily replaceable they are. And that‘s ok. Creative industry is a blue collar job now.

Alpha3031 2 hours ago||
If it's not a problem, then there should be no issue with not concealing the fact, no? Lying about things one considers inconsequential is a useful signal about one's willingness to lie with little benefit.
NicuCalcea 41 minutes ago||||
If there was a ghostwriting detector, I would use that too.
vrganj 3 hours ago|||
Nobody says it's a problem. We'd just like to know.

Factory farming also makes meat cheaper than organic practices. I'd just like to know which one I'm getting.

richardatlarge 12 hours ago|||
Nonsense. Forget proofreaders. Think editors. In publishing some editors practically wrote the books. And then theres ghostwriting ! Think of that!
kalleboo 9 hours ago|||
Part of me wishes we had the same regulation for ghostwriting etc. Nobody should be claiming to have written a book they didn't.
Apocryphon 10 hours ago|||
Okay, then it should be acknowledged if a work was AI-edited-written, or AI-ghostwritten.
demetrius 14 hours ago|||
I'm not sure the quoted statement is true. Proofreading like "point to problems in the text", if you fix the problems yourself and don't copy-paste the solutions given to you, should still be safe, shouldn't it? So, human-written text should not be falsely flagged if you use LLM for proofreading.

And if you copy-paste the answers from LLM, I think it's only fair the end result gets flagged. You're not writing it yourself.

190n 13 hours ago||
> And if you copy-paste the answers from LLF, I think it's only fair the end result gets flagged. You're not writing it yourself.

I wonder if it would even get flagged in that case, because wouldn't the probability distribution of a token when the LLM is suggesting an edit to your writing be different than the distribution of that token once it is in the context of the text it's editing?

beej71 10 hours ago||
I don't even ask LLMs to go that far. Tell me if I've made a spelling, punctuation, or grammatical error, period. Don't rewrite a thing, because LLMs suck at that.
zmmmmm 12 hours ago|||
The question is, when is the "pro writer" version coming that lets you control this behaviour but costs more? Like night follows day, this will happen.

They will need to dodge around the EU requirements but it will probably just come down to an alternative method to watermark or a contractual assurance you won't mis-represent the source of the text.

wuschel 4 hours ago||
I had the same thought.

I hope the other providers will add a geographical limitation on this EU rule.

(On a side note, I wish they would replace those EU beauracts with LLms).

skew-aberration 14 hours ago|||
Can't the LLM just generate e.g diffs? Or some other intermediate language. Then the watermark is lost when the translation step is applied.
jleyank 14 hours ago|||
Rands made this point a few days ago as I recall. Worries about having his tool corrupt his writing during editing, etc.
smb06 12 hours ago|||
The thing that could change is interpreting "the whole thing as generated by Claude"
inigyou 12 hours ago|||
Well yeah, if it's output from Claude it's likely to get detected as being output from Claude.
beering 12 hours ago|||
It’s a strawman argument because if the LLM is really just “proofreading” for you, there will be little or no watermarked text in your writing. Not enough to trip the watermark detector.
KaiserPro 2 hours ago|||
wait what?

but proof reading is a linter, not a writer. the proof reader will say "I think this is clumsy can you try x,y & z"

ButlerianJihad 14 hours ago||
It is quite just, if you think about it. Human works are copyrighted and protected at the moment of creation. All rights reserved. Yet, LLM outputs are uncopyrightable. Therefore, if Claude or any AI has processed my copyrighted work, the end result is uncopyrightable and in the Public Domain. The public has a right to know: is this a human copyrighted work, an LLM PD work, or is the human falsely claiming authorship in order to retain copyright?

A point of confusion for me, however: is every watermark unique? Is every algorithm for watermarking going to vary amongst models and amongst model versions? Will each model publisher keep this watermarking as a trade secret, that they alone can detect? If so, this can't scale! How do you detect "JoeBob 4.3 LLM" output? By querying every single model's watermark-detector? And if they all work by re-running the model and using tokens anew? That is extraordinarily wasteful.

If a watermark is not self-evident, or universally detectable, then it is no good. Take, for example, US currency. The security measures are published and well known. Any count-out room in retail has a big poster indicating how you can detect authentic US bills. Nobody has to accept non-US currency in the US, and so the only authenticity you need to worry about is your US bills alone. LLM watermarking has none of this in common. Currently sounding like a shitshow, if you ask me.

dare944 14 hours ago|||
As I understand it, the current watermarking methods rely on a secret key, making the detection schemes a black box to anyone not in possession of the key. This means organizations like Anthropic are free to make any claim about authorship they want, true or not, and no one can call them on it.
demibabs 12 hours ago||
Part of the legislation requires them to make a public AI text detector (ala GPTZero I assume).

Wouldn’t having that be enough to eventually reverse engineer the key?

dare944 12 hours ago|||
Not if they designed the algorithm right.
inigyou 12 hours ago|||
Probably not to get the key, but you could certainly use it adversarially to remove the watermark.

Removal may come down to changing every third token to a different one.

fwipsy 14 hours ago|||
Perhaps LLM outputs are uncopyrightable, but derivative works of copyrighted works are not automatically in the public domain.
ButlerianJihad 14 hours ago||
That's an intriguing twist, isn't it? It could lead to a tug-of-war.

Working backwards: if it is possible to confirm 100% confidence that a chunk of text is LLM output, then it is "PD until proven otherwise". How can a human reliably assert human authorship of their source text? When all watermark tests fail? Is that proof of humanity now?

If a human proves human authorship, and LLM watermarking tests positive, then is that going to be considered a "derivative work" or not? What if there is an applicable license for the source work, such as "CC-BY-ND" that prohibits derivative works?

This has not been court-tested, and I expect that it will need testing at that level before we can have any assurances.

inigyou 12 hours ago|||
How do you assert it now? I post some text on the internet, you claim you have copyright, how do you prove that?
pizzly 12 hours ago||
Some camera pointing at you while you work. Guessing the camera needs to be watermarked itself, which can be done by adding some watermark on the chip level (each camera will have different watermark), manufacturer can confirm the watermark.
inigyou 12 hours ago||
Is that realistically how that is proven in court today?
pessimizer 13 hours ago|||
The watermark, if I'm understanding correctly, only proves that something has been touched at some point by a particular model, not that it was entirely generated by a particular model.

As for it being "uncopyrightable" if it were the output of an LLM, I think computer people are making a very aspie interpretation of a single decision. I think it's more that the LLM (and thereby its owners) cannot itself hold a copyright on its output, that output has to be touched by a person before it is copyrightable. A particular view from the top of a mountain can't be copyrighted, for example, but a photograph of that view can be.

I'm not sure it at all precludes a "robosigning"* sort of situation, where machines generate output, hired temps sign and claim that output, and immediately sign it over to the people who hired them (as a work-for-hire.) Copyright is stupid, artificial law, not logical.

-----

* https://www.mortgageauditsonline.com/what-are-robo-signers/

demibabs 12 hours ago||
> The watermark, if I'm understanding correctly, only proves that something has been touched at some point by a particular model, not that it was entirely generated by a particular model.

For the watermark to be detectable, the text needs to be like 75% AI generated.

If you have an LLM “touch” one section of the article, it’s not gonna be detectable.

inigyou 13 hours ago||
> One of my fundamental problem with this is that no two synonyms carry the exact same meaning. “He leaped at the chance” and “He jumped at the opportunity” are very similar sentences expressing the same general sentiment, but they are not the same. The exact words we choose when writing matter.

Then why are you using an LLM to write? They're not capable of understanding such nuance. They do pick randomly between two synonymous phrases, they do not use some super smart algorithm to pick the one that sounds the best.

This excuse doesn't hold any water at all - Occam's razor says the author is just super annoyed that his AI writing will be identifiable as AI writing.

FeteCommuniste 12 hours ago|
Yeah, I snorted at the sentence "The exact words we choose when writing matter." Well, then why the heck are you using an LLM to "write," man?
mrweasel 4 hours ago||
This seems like a non-issue, or maybe I have the wrong expectations about writing. You write a text, ask Claude to proof-read it, but then you wholesale just copy Claudes output and use that as the final text? Wouldn't you review the changes it suggests and only take those you agree with, there by completely bypassing the watermarking?

Alternatively, you ask Claude to write the whole thing and proof read it yourself. In that case I'd like to know how much you'd need to change to break the watermarking, i.e. how much of a text would you need to change for it to be considered your work and not that of Claude?

aselimov3 14 hours ago||
This article feels slightly incoherent. You want high quality precise writing and to use an LLM to generate it? Feels like those are diametrically opposed
beering 13 hours ago||
Exactly. The watermark is proportional to how much text is AI generated. Either the AI really just “fixed some typos” (not enough AI content to hide a watermark) or the AI did most of the writing (enough AI content to hide a watermark).
breezybottom 12 hours ago||
This feels like the inevitable outcome of a STEM-only education system. Now people think there's a mathematical formula for picking the "best" words, instead of having to be thoughtful and creative.
amanzi 13 hours ago||
I was initially surprised that Gruber was so invested in the "quality" of AI-generated text, which in my mind is an oxymoron. But really, Gruber's interest here is with the EU. This forms part of his ongoing attacks on the EU, all because they have been forcing Apple to align with regulations.
inigyou 12 hours ago||
Can we install random unapproved apps on our iPhones yet, or is Apple aiming to just be fined a trillion dollars because they make more than that from the 30% cut?
rsynnott 6 hours ago||
Yes: https://support.apple.com/en-mk/117767
inigyou 2 hours ago||
This says we can only install Apple-approved apps.
rimliu 2 hours ago||
I am from EU. Alas it has a tendency to produce some idiotic regulations. Cookie banner, new packaging fee, etc. I genuinely think some Apple related ones hurt customers more than help them.
egypturnash 13 hours ago||
LLMs are already perversions of writing, so what else is new. Oh no, the over-long circumlocution generated by three autocorrects in a trenchcoat might be slightly longer because of this and maybe people will start noticing the subtle rhythms of vaguely peculiar word choices as yet another cue that you are wasting their time with machine-generated wordslop, what a terrible fate. Your long rambling walls of machine-waffling might be 37.05% longer than they need to be instead of the mere 36.58% longer they are now.
_joel 1 hour ago|
So the watermark can be removed by rearranging words and choice of words. This seems trivial to bypass with a local model. If I understand this correctly.
More comments...