Posted by Eloissssss 1 day ago
This was hilighted to me by a librarian friend if mine, She said they destroy books all the time in her job, and got rather angry when people suggested that it was intrinsicly bad, because there is nothing sacred about simply being a book. It is about the replacibility. They had a program where, if they destroyed a book they would receive a replacement from the publisher for a tiny cost. If a book got damaged it would be much cheaper to get a freshly minted replacement than it would be to spend time repairing.
If there was a glitch in the matrix and 30 billion Gutenberg bibles suddenly fell in the middle of the Amazon(the wet one), most would be destroyed as an environmental hazard.
I see a lot of comparisons to book burning when the AI scanning is mentioned, but the point of book burning is a symbolic act to indicate that the information within the book should not be shared.
Destroying the book to capture the information within sits at the polar opposite reason for destroying a book.
If Anthropic subsequently released the digitized version of it we’d be having a different discussion, but for obvious reasons they’re not.
That is also why ChatGPT blocks it now. The plagiarism is still there of course, just hidden.
This simply isn't true. An LLM can do a weak parody of a sufficiently-famous author, but they're extremely poor at sustained fiction writing even without trying to emulate a specific style.
Perfect grammar is only a small part of writing well.
You see a lot of crap for the same reason most code is crap - the human who did the thing sucked. Language models aren't magic, they're just a compiler that guarantees an output regardless of input quality and structure.
That doesn't mean much wrt to the capability of the model.
The world is filed with crap literature, movies, stories, and advertisements written by honest to God human professionals in their respective field, and we trained these things in that environment. It's not surprising to me that they need a nudge to get to Hemingway when the corpus average is presumably so low.
What does this mean?
I just looked into it and I think I can still find GPT 3.5, but wonder if it too has been trained out of usefulness.
The characters a styles were quirky and not always perfect, but added a lot more flavor and character nuance that could be corrected with editing.
> I would like to get the old models back. If anyone can show me the way, I would be grateful. I liked them for NPC dialogue, for TTRPG campaigns and the like.
> I looked into it. GPT 3.5 is still around, I think. But perhaps they have trained it out of usefulness. It is difficult to say.
> They were quirky, the old ones, and not always perfect. But they had flavor, and their characters had nuance, and what they got wrong you could fix in the editing.
Full conversation with thinking: https://paste.ononoki.org/?2fc049fe1ae7302b#GjzuEn8VmaYJ565M... Password: kimitest
I just can't stand to see any more "Why It Matters.", and this trick seems to strip out the worst offenders
Works unfinished at the time of the author's death?
Abandoned works?
Lost works? Prompt: Aristotle's work on comedy has been lost. Write a short treatise on comedy, using Aristotle's methods as shown in his Corpus and especially his treatises on Rhetoric and Poetics.
No amount of prompting would get you Aristotle's actual lost work.
2. No one who appreciates those books wants this.
wouldn't be surprised if pretty much all the training data is legal now, at least retroactively from having obtained copies later.
I don't know what would happen when DeepSeek inevitably catch up though. Perhaps that's going to be how this wave of AI hype ends?
Open weights are a one-way ratchet.