Top
Best
New

Posted by dopamine_daddy 1 day ago

How we measured AI writing across arXiv, and where the measurement breaks(unslop.run)
228 points | 156 commentspage 3
themeiguoren 21 hours ago|
I ran two papers and three blog posts of mine through here, and all but one (correctly) flagged as 0-1% machine written. The other (human written) blog post was 20%. Pretty good afaict!
ergl 21 hours ago||
On my knees begging that the developers of AI detectors use them on their own writing and that at least pretended to mask the obvious LLM cliches in their blog posts.
WhyIsItAlwaysHN 23 hours ago||
Awesome work, the detector seems accurate on a bunch of texts I tried, surprisingly even human/ai hybrid texts
dopamine_daddy 23 hours ago|
That's honestly so good to hear, thank you.
cat-whisperer 23 hours ago||
the funniest part of these AI detectors is that if I were to upload any of einstin's paper's they will all be flagged as AI-written. it makes sense because it's part of their training data.

but this post makes me wonder, if more papers' are written with AI, or the shape of knowledge of converging?

n_e 22 hours ago|
> the funniest part of these AI detectors is that if I were to upload any of einstin's paper's they will all be flagged as AI-written.

Have you tried doing that or even read the article?

The article says that their detector flags 0.4% of pre-AI papers as AI-written.

If I paste the first page from this paper (https://www.fourmilab.ch/etexts/einstein/specrel/specrel.pdf) in https://unslop.run/app, I get a 0% chance that it was AI-written.

consp 22 hours ago||
Even the translated version of the field equations paper has a 1% match only. Which I expected to be a tiny bit higher but still near 0.
wxw 19 hours ago||
Not convinced that this slop measurement is useful.

I asked Codex to generate an article with a high score and then asked Codex to (reverse?) hill climb that score. The original generation scored 97% and then the optimized one scored 1%. Both are pretty bad and read like slop.

https://gist.github.com/wbew/8a2bd6686bf875210f2244ac8ea65bf...

luciana1u 20 hours ago||
the measurement breaks because the detector is probably also AI, and at that point you are just watching two language models argue about who wrote what
arjunvrofficial 22 hours ago||
Good contents should always thrive, with or without AI
IshKebab 21 hours ago||
> and an honest account of the limitations.

Gotta be trolling :-D

digitalPhonix 19 hours ago||
> Here is the method, the results, and an honest account of the limitations.

Pot meet kettle?

guywithahat 22 hours ago|
People are saying this is a bad thing but is it really a problem? The compelling aspect of research is the data and/or description of work, not the writing. Papers probably should be written by AI so that they're clear and well presented, while the researchers should focus on generating good data. If there is no data or work behind the paper, we should question whether the research group needs funding.
epq22 20 hours ago|
but the assumption you are making is that the underlying ideas are compelling. In practice, people often decide to publish incremental and/or mediocre work for various reasons, and then dress these ideas up to seem as compelling as possible to get past peer review or make a press release.

at the least, this is problematic for peer-review because the absolute number of submissions outpaces the time availability of a finite number of expert reviewers. We cannot quickly generate expert human reviewers, and so the community might converge towards half-baked solutions (AI-generated reviews or rejection systems, vastly expanded referee pools, etc.) that tend to erode trust and and make scientific communities more adversarial.

More comments...