Deflate (as used in gzip) uses a Huffman coder. LZMA (as used by xz) uses a predictive range coder. Zstandard can use either Huffman or FSE. Some high-speed compressors like LZ4 skip the entropy coding stage entirely at the expense of compression ratio.
Bzip2 is an interesting aversion of this pattern - it uses the Burrows-Wheeler transform as a first pass instead of LZ. Unfortunately, this is one of the major reasons why it's so slow.
Practical Prefetching via Data Compression; Vitter, Krishnam, Curewitz. 1993
The page addresses ('names') were the characters and the built-up LZ dictionary used to predict which "characters" → pages would come next.
https://www.ittc.ku.edu/~jsv/Papers/CKV93.practical-prefetch...
Optimal Prediction for Prefecting in the Worst Case; Vitter, Krishnan
https://dl.acm.org/doi/pdf/10.5555/314464.314575
Apparently the same trick was later rediscovered for web-pages.
Hutter Prize being where you are paid if you can compress wikipedia small enough. LLMs do very well at that, if, big if, you ignore the cost of initial weights.
A cool Claude Shannon story:
Shannon wanted to measure how much information is actually contained in ordinary
English text. His 1948 theory said such a number must exist, but he had no way to
calculate it, because the patterns in English reach across dozens of letters and no
equation or frequency table captures all of them at once.
So instead of calculating it, he ran an experiment on a person.
He took a passage from a novel that the subject had not read, and covered it with a
card so only the text already guessed was visible. He asked the subject to name
the first letter. If the guess was wrong, he asked again, and kept asking until the
subject named the correct letter. He wrote down how many guesses it had taken,
revealed the letter, and moved the card one position to the right. Then he repeated
the process for the next letter, and the next, through the whole passage.
What this produced was not a sequence of letters but a sequence of numbers — one
number per letter, recording how many guesses that letter required. Most of the
numbers were 1, because someone fluent in English, seeing the preceding text,
usually names the next letter correctly on the first attempt.
Shannon then argued that this sequence of numbers contains exactly as much
information as the original passage.
Sounds a lot like next token prediction to me.https://corecursive.com/the-hutter-prize/
Many of the debates in the comments seem to come down to whether people believe prediction and extrapolation are synonymous.
I found this interesting and wonder whether LLMs have a higher density ceiling, since training and inference don't rely on a fixed lattice and can instead learn their own representational geometry.
This reads as written by someone who just happened to understand what LLMs do, so I totally fail to understand how anything further said can have any real credibility…
As a matter of fact the best compression by Fabrice Bellard’s models have been achieved with NNs long before LLMs.
And also MP3 and MPEG in general are very apparent neural networks, yet not deep as in modern VLMs
Another thought that came from the same post is that, insofar as we see LLMs as human-style intelligence, they're more like stream of consciousness devices. Essentially incessant talking and buying enough time until you get to a usable answer. I think I associate some subset of intelligence with what you don't say, which is impossible with the SOC-style outputs, so this is something I think about a fair bit.
What could maybe differentiate current gen models from next gen is the ability to call tools modeled within the layers themselves, not externally. I think as far as I understand it, model trainers expect the model to do this itself in a way we don't understand or control, like a version of the bitter lesson. But I posit we can model many determinate tools as NNs themselves and figure out how to get the internal states of the LLM to make use of them during inference, e.g. calculators, indexes, citations. Just an enthusiast though, so grain of salt and all.
I'm not a LessWrong^TM rationalist guy, but one really good thought experiment I always keep in the back of my mind from them is Solomonoff induction. AIT people take it as a framework to work with - it's pretty cool, I agree. But I (and some other people, such as certain AI execs at Amazon - according to my interpretation of their public interviews) think it just highlights the trap - given an arbitrarily powerful oracle, you can get compression down pat. Like, if you assume the source is generatable with a turing machine, and you write a function to brute force over all turing machines, then whoa, your compression works. You will necessarily find the optimal compression at some point because your search function is literally searching over all possible turing machines that could've generated the input sequence, anyways (because the input sequence was generated by a turing machine)
These are the kinds of results you can get if you don't have any actual constraints on what the compressor can do.
(Of course, again - this is not the point of solomonoff induction - it's to use this as a base truth, to then layer parsimony on top of that. There are infinite number of turing machines that could match your prefix, parsimony filters, throw some bayesian inference on top of that, and you get Solomonoff induction. They constrain it afterwards. But I think to that intuition as a base whenever people claim new results.).
But I see in casual conversation, people constantly making claims like, "LLM's are so good because they compress a model of the world". What is that model then? Scott Aaronson has made points like this before - your "model" could just be a massive lookup table, so you can't just claim "compression" and win - the compressor must be reasonably small, too.
I don't object to the notion that LLM's have some notion of world models more sophisticated than memorization. That's proven by actual interventional experiments, such as the ones that actual interperability researchers do. But mere compression is vacuously powerful. "Vacuous" not in the sense that "oh, you might be suboptimal and be a little more complex", vacuous as in "the philosophical point you were trying to make is vacuous because you make a vacuously powerful statement".
(I'm not a total fan of intervention either, as an end-all gospel as some people use, but it's far, far better than not having it).
If you want to best predict what state comes next from a space of possibilities, you have to figure out how probable each next state is and pick the most probable one.
If you want to compress something, you have to figure out how probable each next next state is and assign the smallest code to the most probable state.
These two processes are essentially the same, and the resulting structure of a system which regulates either process will be similar, approaching the same structure at high confidence.
Driven by Compression Progress: A Simple Principle Explains Essential Aspects of Subjective Beauty, Novelty, Surprise, Interestingness, Attention, Curiosity, Creativity, Art, Science, Music, Jokes
https://simplicitytheory.telecom-paris.fr/
page created 8 days after Schmidhuber's paper.