Posted by ropbear 1 day ago
"Important" refers to impact, "significant" to the statistical qualities of the underlying hypothesis, "substantial" to the work involved, and "notable" is a referential judgement by the speaker. The implied normalization of words and their respective meaning also marks one of the mechanisms how "slop" is typically creeping into the productions of "broad verbose interchange replicas" (it's all interchangeable, and a choice isn't really that, a choice, isn't it?).
You think someone will write Nazi-promoting AI policy like this when society is encouraged to look at their families as examples of neo-Nazi corruption? When their wives' and kids' friends spurn them while their families engage in obvious criminal activity to harm world productivity?
Critics need to be more precise.
Generative AI is a perversion of writing
Total variation distance has been measured to decrease as you scale a model, and that is the primary mechanism "watermarking" as discussed in the Anthropic announcement relies on. It becomes more difficult to reliably detect text as a fixed sample count without tweaking the distribution further. Either way, it's a minor problem that will be addressed over time, compared to the issue of who can detect this without guessing or developing their own sets: providers not releasing a way to detect any such watermarks without going through them makes this entire approach hostile to the public. The EU regulation on this subject is interesting, although again most certainly not the primary driver for these practices:
"1.1.2: Signatories will ensure that AI-generated or manipulated content is marked with an imperceptible watermark, with the exception of very short text. For free-form text longer than 200 tokens, watermarking still needs to be applied, even though it may have lower reliability compared to that of watermarking very long text"
A proper, effective and useful law would have required providers to regularly release datasets to run your own verification on any text released within a fixed interval of time, presumably once out of rotation. Instead, it only talks about exposing an user interface going through their own services:
"Signatories will ensure access to their detection solution through a user interface appropriate for the audience of end-users that may eventually be exposed to the content generated or manipulated by their AI system. [...] Any restriction to the access will be limited in time until more reliable and robust detection mechanisms have emerged and have been adopted as the state of the art for detection mechanisms for the watermarking of free-form text evolves."
Most interestingly, in line with the EU's mass-surveillance program, an alternative solution to watermarking where it may not be sufficient is also suggested, although only optional for now:
"Where appropriate and taking into account potential trade-offs related to privacy and security, as well as scalability challenges and costs, Signatories may implement as an optional supplementary measure fingerprinting or logging solutions for AI-generated or manipulated content which allow for checking whether content has been generated or manipulated by their AI system. For example, direct logging may be appropriate for text content, whereas fingerprinting approaches may be preferable for audio and visual content."
In 2005, FFII predicted that the Data Retention directive would not pass the Courts.
It just took 10 years for the CJEU to strike down the measure, as it is "mass surveillance".
Logging everything your bot does by law is of the same dimension?
Since most Americans are in the latter camp, anything that even hints at making companies responsible for anything is viewed as being squarely the former.
Whereas most EU regs are "play nice, be responsible, behave like adults. If not, this can always turn ugly". Same here.