Top
Best
New

Posted by ropbear 1 day ago

Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing(daringfireball.net)
752 points | 662 commentspage 11
mdavid626 13 hours ago|
Just copy Claude’s output and shove it into Gemini and ask to rephrase.
solid_fuel 7 hours ago||
What a stupid take. Generating text with an LLM is already a ‘perversion’ of writing. Tweaking the last random-choice step doesn’t meaningfully change that at all.
masswerk 11 hours ago||
Mind that even in their first example, "The results of the study were quite (important | significant | substantial | notable)", the meaning is by no means interchangeable.

"Important" refers to impact, "significant" to the statistical qualities of the underlying hypothesis, "substantial" to the work involved, and "notable" is a referential judgement by the speaker. The implied normalization of words and their respective meaning also marks one of the mechanisms how "slop" is typically creeping into the productions of "broad verbose interchange replicas" (it's all interchangeable, and a choice isn't really that, a choice, isn't it?).

analog31 11 hours ago||
The first thing that came to my mind was "security theater." Making people think that AI is detectable could have the same effect as actually making it detectable.
khelavastr 5 hours ago||
It's because the individuals who write those laws are literally trying to shove neo-Nazi policies into EU practice. I wonder how many people would make policies like this if they and their families were publicly identified and criticized as neo-Nazi elements in society.

You think someone will write Nazi-promoting AI policy like this when society is encouraged to look at their families as examples of neo-Nazi corruption? When their wives' and kids' friends spurn them while their families engage in obvious criminal activity to harm world productivity?

Critics need to be more precise.

tantalor 11 hours ago||
Fix it for you:

Generative AI is a perversion of writing

0x_rs 23 hours ago||
I'd encourage reading this paper, and literature on scaling laws in autoregressive models: https://arxiv.org/abs/2303.11156

Total variation distance has been measured to decrease as you scale a model, and that is the primary mechanism "watermarking" as discussed in the Anthropic announcement relies on. It becomes more difficult to reliably detect text as a fixed sample count without tweaking the distribution further. Either way, it's a minor problem that will be addressed over time, compared to the issue of who can detect this without guessing or developing their own sets: providers not releasing a way to detect any such watermarks without going through them makes this entire approach hostile to the public. The EU regulation on this subject is interesting, although again most certainly not the primary driver for these practices:

"1.1.2: Signatories will ensure that AI-generated or manipulated content is marked with an imperceptible watermark, with the exception of very short text. For free-form text longer than 200 tokens, watermarking still needs to be applied, even though it may have lower reliability compared to that of watermarking very long text"

A proper, effective and useful law would have required providers to regularly release datasets to run your own verification on any text released within a fixed interval of time, presumably once out of rotation. Instead, it only talks about exposing an user interface going through their own services:

"Signatories will ensure access to their detection solution through a user interface appropriate for the audience of end-users that may eventually be exposed to the content generated or manipulated by their AI system. [...] Any restriction to the access will be limited in time until more reliable and robust detection mechanisms have emerged and have been adopted as the state of the art for detection mechanisms for the watermarking of free-form text evolves."

Most interestingly, in line with the EU's mass-surveillance program, an alternative solution to watermarking where it may not be sufficient is also suggested, although only optional for now:

"Where appropriate and taking into account potential trade-offs related to privacy and security, as well as scalability challenges and costs, Signatories may implement as an optional supplementary measure fingerprinting or logging solutions for AI-generated or manipulated content which allow for checking whether content has been generated or manipulated by their AI system. For example, direct logging may be appropriate for text content, whereas fingerprinting approaches may be preferable for audio and visual content."

zoobab 10 hours ago||
"logging solutions"

In 2005, FFII predicted that the Data Retention directive would not pass the Courts.

It just took 10 years for the CJEU to strike down the measure, as it is "mass surveillance".

Logging everything your bot does by law is of the same dimension?

troupo 13 hours ago||
EU is a capitalist union first and foremost. So their regulations try to straddle the thin line between "regulate to death and tell companies exactly how to do things" and "the companies can do whatever the hell they want".

Since most Americans are in the latter camp, anything that even hints at making companies responsible for anything is viewed as being squarely the former.

Whereas most EU regs are "play nice, be responsible, behave like adults. If not, this can always turn ugly". Same here.

1vuio0pswjnm7 4 hours ago||
"My writing is my work, and Anthropic's current strategy is aggressively writer-hostile."
Retr0id 12 hours ago||
What a strange take. LLMs themselves are a perversion of writing. I couldn't care less about the implementation details of the PRNG they use for next-token sampling (well, as long as they're not stuffing a user ID in there).
adamretter 9 hours ago|
I almost never post a comment here, but everything about the author's position is offensive and self entitled. I am so enraged that I can't even beging to formulate a response without resorting to very bad language. It's sad, because until now I respected the author. But clearly, and sadly, he has been afflicted with AI brain rot, and is likely in some stage of withdrawal. I wish him a speedy and safe recovery.
More comments...