Top
Best
New

Posted by AnhTho_FR 5 hours ago

A walk through of the DeltaNet family of linear attention variants(blog.doubleword.ai)
271 points | 111 commentspage 2
neutrinobro 5 hours ago|
You know its a doozy when the author writes a disclaimer at the top saying that bra-ket notation was chosen in order to make the algorithm and data structures clearer.
CodesInChaos 4 hours ago|
One of the more annoying parts of my physics study was getting used to the new matrix multiplication notation they came up with every semester.
kurthr 4 hours ago||
bra-ket is the (most?) general form of tensor manipulation.

Raising and lowering operators for summation notation are the beginner tools for covariant derivatives of the metric tensor.

Christoffel symbols are where it's at, if you need to write out the Ricci tensor. The more constrained the space the more concise the notation can be.

Note that MechE tensor notation has an even more compact (eigen) form for principal stresses.

LogicFailsMe 3 hours ago|||
All of this is true, but I don't believe and I want to be wrong about this that there is something in this notation that starts at an ELI 5 level and gently guides you to physicist level expertise. I all but majored in math (deriving back prop was trivial once it was clear it was the chain rule as one example) but I have never been able to keep bra ket notation straight in my head for the more exotic operations. Einsteinian notation on the other hand is a few minutes of furled brows and then all is clear.

It is what has separated me from being able to code just about anything on a GPU and being known for some of that work and coming up with a better way to run ab initio quantum chemistry on them.

It truly has been my Waterloo for many years. So make me wrong.

kurthr 3 hours ago||
Yeah, bra-ket is arbitrary tensors (inner and outer multiplication) rather than the nice 4D of space-time (with derivatives).

I will say that seeing transformers written this way gives me a bit more intuition for what is going on (being able to identify correct equations), but there's enough complexity in actual transformer implementations, that it still feels like I'm fooling myself.

Conceivably, I think you could use Feynman diagrams to talk about phonon dispersion in (eg asymetric crystaline) solids, but even though they're a "simplification", they're overkill for the problem.

neutrinobro 2 hours ago|||
I don't think there is anything in this article that actually demands bra-ket notation (a state in some Hilbert space), that couldn't be more clearly written with standard notation for a Euclidean inner/outer product, but I suppose everyone has their own preferences for notation.
scarmig 5 hours ago||
I like the math vs physics toggle.
_davide_ 4 hours ago||
Loved this incremental evolution, things gets way more understandable...usually xD
MrFiskarBengt 2 hours ago||
You could've also came up with Newton's laws. After all, they look trivial in retrospect. But, there's an important lesson in a story about balancing an egg here that can teach us something.

Filippo Brunelleschi said he could build the large dome for the church that had stood unfinished for a century. Skeptical, other's demanded he'd explain how. He refused. Instead he challenged everyone to balance an egg on its tip. Nobody could do it. He then demonstrated by lightly tapping the egg on the table, flattening the tip, making it stand. "Anyone could've done that! You never said we could break the egg!". And that's the point. Anyone could've done it. But nobody did. Nobody thought 'outside the box'. And likewise, his solution to building the dome is as simple, and as ingenious.

It's called Egg of Columbus. (there's a similar story about Columbus that's more famous, but apparently fictitious). It teaches us that hindsight is 20/20.

edflsafoiewq 8 minutes ago|
"You could have invented..." a just a genre of expository writing which builds a path from something you know to something you want to learn. The sequence of development is not necessarily even historically accurate, it can be completely invented as long as the learner finds it natural and it helps motivate them.

The fact that invention is hard and actually you probably couldn't have invented a bunch of hard stuff is not really that, um, relevant.

andai 4 hours ago||
>You Could Have Come Up With Kimi Delta Attention

What? Little old me! Well, then, let's have a look...

> (First paragraph)

> A note on notation: this article defaults to bra-ket notation because (in my quantum-inspired opinion) it makes the shapes in this derivation very clear. The Math notation switch above rewrites every equation using conventional bold vectors and explicit transposes instead. In bra-ket mode, ∣ q ⟩ ∣q⟩ is a column vector, ⟨ k ∣ ⟨k∣ is a row vector, ⟨ k ∣ q ⟩ ⟨k∣q⟩ is a number, and ∣ v ⟩ ⟨ k ∣ ∣v⟩⟨k∣ is a matrix. Vectors face right by default, while keys face left when written into the linear-attention state. We work with one causal attention head and real-valued vectors, assume DeltaNet’s keys are normalized, and let the state map from key space to value space.

Hmm... Guess not!

5555watch 4 hours ago|
I love that they let you switch to a more common q'k notation!
bee_rider 4 hours ago||
Where do linear algebra folks go to get started with ML stuff? It seems pretty easy but the hardware is expensive.
sva_ 3 hours ago||
I think Karpathys nn zero to hero is a good starting point. And you can experiment on small networks using pretty normal hardware.
thatjoeoverthr 2 hours ago|||
I’m having a great time with an NVIDIA 3090. 24 GB RAM will run a lot of neat models. But at zero you can for sure just do CPU until you build a project ambitious enough.
stuxnet79 3 hours ago|||
> It seems pretty easy but the hardware is expensive.

Huh?

If your aim is to truly 'get started' with ML then hardware is absolutely not a bottleneck (either local or cloud).

Remember that ML is much more than LLMs. Even modern day LLMs can be quantized to a point where they can run on local hardware although their capabilities won't be as impressive.

I would recommend looking into some of Andrej Karpathy's videos if you want a grasp of the basics.

nifets 3 hours ago||
what is a linear algebra folk?
sodapopcan 3 hours ago||
Ohhhhh Diag(αt), right. I was almost there but had left the placeholder "Diag(foo)" and never noticed. I now see is why I didn't come up with it first. So close!
luciana1u 2 hours ago||
love the toggle between math notation and physics notation. two flavors of confusion, nicely packaged.
anshumankmr 4 hours ago||
https://www.youtube.com/watch?v=p0CEOkoSsgA
joe_the_user 2 hours ago|
Overall, all the different linear attentions out there are approximations the original (quadratic) attention and this is important for the whole "AI" enterprise[2].

Original attention involves (very crudely) an approach of scanning how every token (roughly a word) relates every other token and training a classic neural network on related tokens - to get either language translation or next word prediction (and next word prediction is what "seems intelligent" in LLMs). [1]

The problem is that since original attention is "everything to everything else" it scales quadratically (O(n^2)) with the size of the train set (or train set window) and so basically even the largest data center can use that once a truly vast training set is accumulated. Which is to say that "dirty little secret" of LLMs following the "Attention Is All You Need" paper don't actually scale. That model (in my crude, amateur understanding) is elegant for allowing every word's connection to every other word to be weighed and still brute-force for not starting with or achieving "understanding" of the words [3 give only some background but also why "full" attention is powerful].

Linear attention is a way around the quadratic quality of original attention so everyone is naturally using clever approaches to make it work. Simplifying terribly - you're trying to determine the value of word before you see in context. But my intuition is that since (Everything X Everything) is inherently a quadratic relationship, none of these can capture their expanded data set in the way original LLMs did - not they are worse but all the models seem likely to hit diminishing returns in terms of blindly capturing meaning from all-the-world's text (and data).

Background and notes: [1] https://en.wikipedia.org/wiki/Transformer_(deep_learning_arc... [2] Linear Transformers Are Secretly Fast Weight Programmers: https://proceedings.mlr.press/v139/schlag21a/schlag21a.pdf [3] Transformers are Deep Infinite-Dimensional Non-Mercer Binary Kernel Machines: https://arxiv.org/pdf/2106.01506

More comments...