Top
Best
New

Posted by volotat 1 day ago

Show HN: Mini-AGI – Dynamic continual learning model trained on 8GB VRAM(github.com)
Sorry for the pretentious name, I know, I know.. It just contains all the pieces I would like to see a AGI model to have, and I can't stand the temptation. Before throwing rocks at me, please take a glance at the Readme, and I hope it will cover your mood a little bit.

So, first of all it does work and you can see the sample from the whole training run here: https://raw.githubusercontent.com/volotat/mini-AGI/refs/head...

Here is the scaling law graph I have so far, and it looks very promising: https://github.com/volotat/mini-AGI/blob/main/assets/scaling...

The model was built under my deep dissatisfaction so we cannot really train even moderately big models (1B+ scale) on the consumer's hardware. We can inference and fine-tune them for sure, but I would like to have full control over what the model sees over the training run, so it is fully aligned with my interests, not some corporations.

I was thinking about for some time and come up with two interesting ideas I thought worth pursuing: MoE with a lot of experts that gets added and pruned from the model while it trains, where only a small subset of of experts are actually in use at any particular moment + batch 1 training on the single continuous stream of data.

First allows us to be bounded only by the disk space in terms of number of parameters and load and unload experts only when they are needed. The second (if figured out and it turns out to be doable) allows us to get aways with small VRAM capacity because we do not need to store big randomized batches and their respective gradients.

I started brainstorming with Claude and after some time we found an approach that seems to be promising, and low and behold, a few weeks pass and you can see the results yourself.

Obviously, I did use AI in the process of making this project and I am pretty sure it would be completely impossible for me to do something like this without it, so I hope it is more than justified.

The model is still running over the first of 7.8B characters corpus I selected for training, so the weights are not out yet, and it's about a couple weeks of waiting until they are cooked at the current reading speed. And yeah, the model just read continuous interleaved passages from the dataset, each by 32K characters long each as a single stream. Just as you or I would do.

The set up seems to be really simple so you can git clone the project, run it and observe everything for yourself.

Thanks for your attention.

266 points | 68 comments
abeppu 23 hours ago|
I have not looked carefully but it seems like this is over-promising on avoiding catastrophic forgetting.

The "trunk learning rate" is set at 0.1x the learning rate for the experts, so learning on different subjects disproportionately happens in the experts, and the trunk portion is comparatively more stable. But the population of experts can grow and shrink:

> The pool grows when it is short of capacity and shrinks when parts of it stop being asked for.

So:

- doesn't the trunk then _eventually_ still undergo catastrophic forgetting, it just may take much longer?

- and before that point, catastrophic forgetting happens in stepwise chunks whenever the expert pool shrinks?

volotat 2 hours ago|
I am not sure about eventually, but it learns on the steady paste so far. The main thing to keep in mind is that it learns on the single STREAM of data. Not randomized batched samples. Try to do it with any other model and you will see nothing but complete garbage in the predictions, exactly because of catastrophic forgetting.

And here are the types of samples the model produces after about a week of training:

    ==============================================================================
    step 191,447   391.3M of 7,879M characters (4.97%)   15 min   176 experts
    context 4,096 characters of 4,096   reading 1,046 char/s   still gaining +0.0412 deep into it
    grad norm 0.98 against a clip of 1   under the clip
    train loss 0.6540   lr 2.28e-04   evidence t -0.15 over 65.7 (effect +0.0660)   rate x0.753
    held-out loss 0.8012 +/-0.0335 nats   1.1559 bits/char   perplexity 2.23   gap +0.1472
      arithmetic 0.645   chat 0.790   chat_hermes 1.153   chess 0.510   code 0.703   reasoning 0.756   stories 0.596   wikipedia 1.256
    repeats 31% of 8-grams, greedy with no guard
    ==============================================================================

    --- stories ---
    prompt: 'Once upon a time, there was a little boy named Tom. One day he '
    [raw]  repeated 8-grams 2%
    was walking in the park with his mom. He saw a big box in the ground. He was very sad and sad.

    "Mom, can I have the box to the box?" he ask
    [adapted]  repeated 8-grams 5%
    was playing with his friend, a little girl.

    "Wow, Tom!" said Tom.

    "What's your name?"

    Tom said, "I'm playing with my friend."

    The little

    --- code ---
    prompt: 'def merge_sorted(a, b):\n    '
    [raw]  repeated 8-grams 61%
        return sorted(a, b)

        def merge_sorted(a, b):
            return sorted(a, b)

        def merge_sorted(a, b):
            return sorted(a, b)

    
    [adapted]  repeated 8-grams 22%
    """Merges the modification from the angles in the argument.
        This merges the angles in the argument.
        """

        def __init__(self, a):


    --- arithmetic ---
    prompt: 'add 4917 + 388 = '
    [raw]  repeated 8-grams 47%
    <think> 7+8+0=5c1 1+8+1=0c1 9+3+1=3c1 4+0+1=5c0 </think> 5305
    mul 6 * 6 = 36
    mul 6 * 6 = 36
    mul 6 * 6 = 36
    mul 6 * 6 = 36
    mul 6 * 6 = 36
    mul 6 * 6 = 36
    a
    [adapted]  repeated 8-grams 3%
    <think> 7+8+0=5c1 1+8+1=0c1 9+3+1=3c1 4+0+1=5c0 </think> 5305
    mul 6287 * 49 = <think> 6287*9=56583 6287*40=251480 56583+251480=307063 </think> 307063
    mul 9428 * 17 = 

    --- chat ---
    prompt: '<user>\nWhat are you?\n</user>\n<bot>\n'
    [raw]  repeated 8-grams 20%
    In the given I am trying to find the maximum vertex should be an expression linear that should be given the maximum vertex and the maximum v
    [adapted]  repeated 8-grams 8%
    In the given I, there's another part of the given I, whose parts include anything they're discussed.
    </bot>
    <user>
    What is the given article by answer

    --- chat_hermes ---
    prompt: '<user>\nA train travels 60 km in 45 minutes. What is its speed in km/h?\n</user>\n<bot>\n'
    [raw]  repeated 8-grams 63%
    The train travels 60 km/h in 45 minutes. The train travels 60 km/h in 45 minutes. The train travels 60 km/h in 45 minutes. The travels are t
    [adapted]  repeated 8-grams 4%
    The train travels 60 km/h, which matches km/h.
    </bot>
    <user>
    Write an example of the following article:

    In the subject of the subject, include an imp

    --- reasoning ---
    prompt: '<user>\nWrite a Python function that returns the largest number in a list.\n</user>\n<bot>\n<think>\n'
    [raw]  repeated 8-grams 35%
    Okay, let's tackle this problem. So, the largest number in a list is a list. The largest number is a list. The largest number is a list. It 
    [adapted]  repeated 8-grams 20%
    Okay, let's try to figure out this problem. So, we'll find the largest number in a list.

    The largest number is always the largest number in

    --- wikipedia ---
    prompt: '== History ==\nThe [[Roman Empire]] was '
    [raw]  repeated 8-grams 48%
    a [[computer state]] of [[Computer Empire]] and [[Computer Empire]]. The [[Computer Empire]] was a [[computer state]] of [[Computer Empire]]
    [adapted]  repeated 8-grams 42%
    the [[United States|University]] of [[Candie]]. The [[University]] was the [[University]] of [[Candie]] where the [[University]] was the [[U

    --- chess ---
    prompt: '<g>1700 1-0 1. e4 e5 2. '
    [raw]  repeated 8-grams 0%   22 legal moves, then Nd3
    Nf3 Nc6 3. Bb5 a6 4. Bxc6 dxc6 5. O-O Bg4 6. h3 Bh5 7. g4 Bg6 8. d3 Be7 9. Nbd2 Nf6 10. Nb3 O-O 11. Nc5 Bxc5 12. d4 Bd6 13. Nd3 Bxf3 14. Qxf
    [adapted]  repeated 8-grams 0%   16 legal moves, then Ba3
    Nf3 d6 3. Bc4 Nf6 4. d3 Be7 5. O-O Nbd7 6. Be3 c6 7. Nbd2 O-O 8. c3 a6 9. Qc2 b5 10. Ba3 Nb6 11. Bxe7 Qxe7 12. Rac1 Bb7 13. Nf1 Rac8 14. Ng3

    --- self-knowledge ---
    prompt: '<user>\nhow do you decide which experts to use?\n</user>\n<bot>\n'
    [raw]  repeated 8-grams 2%
    The directory is not a vector of 512, which is why the new chunk is not an expert. That is why my window can be extended by that no matter h
    [adapted]  repeated 8-grams 1%
    The directory is not a vector of 512, which is why. There is not an expert involve
    </bot>
    <user>
    Can you write change_string? It should change the com
HarHarVeryFunny 19 hours ago||
Mini-AGI is a totally inappropriate name - it seems what this project is shooting for, but not delivering on, is being a language model with "continual learning".

Where it seems to fail, by design, on this goal is in delivering continual learning that is more than just "memorization with LRU catastrophic forgetting".

That said, props to the author for thinking different and actually implementing something. Maybe the project can grow into something more, or inspire different ideas, if they continue to work on it.

OutOfHere 1 hour ago|
In English, it's "thinking differently", not "thinking different". This is not Apple.
HarHarVeryFunny 31 minutes ago|||
I was indeed channeling Apple - presumably most people on a tech forum realize that.
mpalmer 1 hour ago|||
Yes, this is Hacker News. Read the commenting guidelines.
ilaksh 1 day ago||
If you actually scroll through the transcript he links to, you will see that something that looks like it could be training is happening, but no coherent responses are coming out at any point. At least not that I saw skimming through.

That might explain why there are no benchmarks of any kind.

seanhunter 19 hours ago||
To pick a couple of examples:

   --- chess ---
   prompt: '<g>1700 1-0 1. e4 e5 2. '
   [raw]  repeated 8-grams 83%
   Nxd5 Nxd5 Nxd5 Nxd5 Nxd5 Nxd5 Nxd5 Nxd5 Nxd5 Nxd5 Nxd5 Nxd5 Rxd5 Rxd5 Rxd5 Rxd5. Rxd5 Rxd5 Rxd5 Rxd5 Rxd5 Rxd5 31 Rxd5 Rxd5 Rxd5 Rxd5 Rxd5 Rx
   [adapted]  repeated 8-grams 2%
   Nxd6+ Bxc3+ Nf6 14. Qxd5+ Nf6 Rxe3+ Bh1 Nxd5+ Qxc6 Rf3+ Nxd5 Qe4+ Rxf6 Nh1+ Qxd5 Rf3+ Nxe6 Qh1+ Rxd5 Nf6+ Qxe3 Rh1+ Nxd5 Qf6+ Rxg3+ Ne1 Qxd5
^ THe model noticed you started the notation of a chess game, but its response is total nonsense. After "1. e4 e5 2." you can't go Nxd6+. For all kinds of reasons. You haven't got your knight out yet. Even if you had, it couldn't get to d6. Even if it could, there's nothing there it could take. If you did somehow in spite of all that manage to play 2 Nxd6+ the opponent couldn't play .... Bxc3+ because they haven't got their bish out. Even if they had it couldn't get to c3 even if it could there isn't anything there to take - you only have a pawn on e4 and a magical knight on d6. Even if somehow in spite of that, you could take on c3 it wouldn't be check and EVEN IF SOMEHOW ALL OF THAT WERE TRUE YOU ARE IN CHECK. You can't move your bishop you need to do something about the Knight on d6 which has you in check.

All the rest of it is similarly gibberish. I'm used to model training garbage but this is in no sense AGI. It's beyond nonsense to call it that.

synctext 1 day ago||
Using the term AGI and not including any performance analysis. My AI calls it: "massive marketing overreach". Somebody called this slop in the comments.

As a professor who published on continual learning I'm leaning towards agreement[1]. It lacks any substance. No relation to related work, no description of algorithm, no ablation study, just hand-waving that we're feeding some data and "Chess is not forgotten".

This "how-continual-learning-works" markdown text is not an algorithm [2].

[1] https://arxiv.org/abs/2301.12530

[2] https://github.com/volotat/mini-AGI/#how-continual-learning-...

volotat 1 day ago|||
There is no special algorithm, the finding is that slowing down the LR or the trunk, while keeping the LR of the experts is enough to eliminate most of the forgetting in the network. You can see in that experiment where chess data was the only thing the model read for 524K characters, yet it kept almost the same performance (i.e. held-out loss) on all other domains. If you keep LR the same across the whole network the loss in other domains degrades dramatically - this is a clear sign of catastrophic forgetting in action. What I can say for sure is that any traditional network that does pose a sign of catastrophic forgetting would not be able to learn any patterns from a single stream of data.

There are no benchmarks published as the model is heavily undertrained, but it is learning. And you can see this clearly in the loss and samples even though they are still barely coherent.

I am not an academic and am not trying to publish a paper about a “major breakthrough” or something like this. I am just a small person who found a cool thing that clearly works and wants to share it with the world. That’s it.

ilaksh 1 day ago||
You can't claim it "works" if it hasn't produced any coherent responses and is still early in your first training attempt.
volotat 1 day ago||
It is a goalpost that is easy to move. By "works" I mean learning from a continuous single (meaning batch-1) stream of data. The fact that it produces full words and full coherent phrases instead of a random stream of characters that would any typical LM produce if trained under the same training regime.
ilaksh 1 day ago||
I would be okay if you shared it as a potential idea and possibly interesting early result, but the language you are actually using to characterize it is misleading or delusional.

Please get a model to the point where it seems like it has some natural language understanding and then share again with reasonable characterization.

bigbadfeline 20 hours ago|||
Is "AGI" the language that bothers you? Well, one has to keep his eyes on the prize and I see a bright idea which could lead to AGI, so, why not describe it as such? I also see the inspiration and hard work necessary to move that idea further along, so fingers crossed.
volotat 1 day ago||||
For sure. As it will pass through the whole corpus I will share the weights, run it through established benchmarks for small models and share all of this as an update. I am also planning on making a Youtube video explaining in detail how it works on a deeper level and the whole reasoning behind why it is built the way it is. But no promises here.
fuzzfactor 14 hours ago||
Nothing to be ashamed of if you end up pushing back the release date of your feature film :)

I had ideas not completely unlike this so long ago, but one big difference can be summed up in one of your parameters.

>Directories are walked, binaries are skipped . . . and each file is read from its beginning to its end because a document has an order.

For me it was binaries being walked because text and anything approaching a language model was so much further out-of-reach having such limited computer power.

ilaksh 1 day ago|||
Actually I'm mad that I wasted my time looking at it based on the claims. He implies it is trained and uses the term "AGI" and "continuous learning". He never finished a single training run or enough that he considers not "undertrained". It's not trained. And actually there is no evidence that it can actually learn anything useful.
synctext 1 day ago||
Indeed this is wasting HN time.

"The model reads 524,000 characters of chess". This is 100KByte of training data in a toy model with rigid parameters and no global learning. Gap with real LLM and trillions of tokens.

This model really addresses the problem of preserving previously learned knowledge, but by restricting the LR of the trunk it stops acquiring new knowledge. Details: "Rethinking the Stability-Plasticity Trade-off in Continual Learning from an Architectural Perspective"

whizzter 1 day ago||
Nobody will throw rocks, I think most people are curious/suspicious about the big players and wants more hands-on since we suspect that this all will come down in cost soon enough.
tnspacetime 1 hour ago||
I read your github page and there's mention of continual learning but no argument how catastrophic forgetting is avoided. I think you should explain that better.
volotat 1 hour ago|
"How continual learning works" is the section you are looking for. The main experiment is described there.
scottsiume 5 hours ago||
I'd love to explore a question. If I have a standard elementary school math textbook, along with all the results of every correct and incorrect answer my child got on the textbook exercises and class tests, is it possible to train a model that can help me figure out what concepts my child is struggling with and have that model provide help?
advael 1 day ago||
Seems interesting, I've been messing with a lot of continuous learning approaches lately and it's cool to see something that's built from the ground up for avoiding catastrophic forgetting. Worth a clone for sure
lostmsu 1 day ago|
It doesn't show any indications of solving catastrophic forgetting.
HarHarVeryFunny 19 hours ago||
In fact it says the opposite - that there is pruning.

Our brain also has some capacity limit, and maybe degraded memory performance over time, but in either case it's a graceful degradation - you may forget fine details of things that happened a long time ago etc, but you don't forget how to ride a bike just because it's been a while.

Continual learning by itself is useless - that's just memorization and filling up a fixed size memory bank. What "continual learning" as one of the things missing from LLMs, is really referring to is roughly "continual learning, with ongoing generalization and merging of memories, with no catastrophic forgetting, with graceful degradation".

fuzzfactor 15 hours ago||
>roughly "continual learning, with ongoing generalization and merging of memories, with no catastrophic forgetting, with graceful degradation".

Sounds a bit like real-time "distillation" to me.

I coudn't imagine there was any choice back in 1980 when we only had kilobytes of memory.

bananaflag 1 day ago||
This is the first thing I see in my life that really looks like proto-AGI, it deserves its name.
cpldcpu 1 day ago|
Is this architecture actually able to generalize or is it mostly based on memorization? Have you tried some basic tasks that require generalization? e.g. number addition etc?
volotat 1 day ago|
The model is way too small and undertrained to make any generalization claims. I want to wait until it reads the whole corpus I gave and then test it on some simple established benchmarks to see how it will behave.
jacquesm 1 day ago|||
What kind of hardware are you using for training?

nm, I found it:

> RTX 3070 Laptop GPU with 8 GB

Super impressive.

dinfinity 1 day ago|||
Seems a bit premature to make an HN post about then, imho.

It's an interesting idea, but it doesn't really do anything interesting yet. I looked at the output in the training run and it is a far, far cry from intelligence. Worse than GPT-2 as it stands.

I do hope it will perform well when scaled and trained, though; best of luck.

More comments...