Top
Best
New

Posted by volotat 1 day ago

Show HN: Mini-AGI – Dynamic continual learning model trained on 8GB VRAM(github.com)
Sorry for the pretentious name, I know, I know.. It just contains all the pieces I would like to see a AGI model to have, and I can't stand the temptation. Before throwing rocks at me, please take a glance at the Readme, and I hope it will cover your mood a little bit.

So, first of all it does work and you can see the sample from the whole training run here: https://raw.githubusercontent.com/volotat/mini-AGI/refs/head...

Here is the scaling law graph I have so far, and it looks very promising: https://github.com/volotat/mini-AGI/blob/main/assets/scaling...

The model was built under my deep dissatisfaction so we cannot really train even moderately big models (1B+ scale) on the consumer's hardware. We can inference and fine-tune them for sure, but I would like to have full control over what the model sees over the training run, so it is fully aligned with my interests, not some corporations.

I was thinking about for some time and come up with two interesting ideas I thought worth pursuing: MoE with a lot of experts that gets added and pruned from the model while it trains, where only a small subset of of experts are actually in use at any particular moment + batch 1 training on the single continuous stream of data.

First allows us to be bounded only by the disk space in terms of number of parameters and load and unload experts only when they are needed. The second (if figured out and it turns out to be doable) allows us to get aways with small VRAM capacity because we do not need to store big randomized batches and their respective gradients.

I started brainstorming with Claude and after some time we found an approach that seems to be promising, and low and behold, a few weeks pass and you can see the results yourself.

Obviously, I did use AI in the process of making this project and I am pretty sure it would be completely impossible for me to do something like this without it, so I hope it is more than justified.

The model is still running over the first of 7.8B characters corpus I selected for training, so the weights are not out yet, and it's about a couple weeks of waiting until they are cooked at the current reading speed. And yeah, the model just read continuous interleaved passages from the dataset, each by 32K characters long each as a single stream. Just as you or I would do.

The set up seems to be really simple so you can git clone the project, run it and observe everything for yourself.

Thanks for your attention.

269 points | 70 commentspage 4
gslepak 22 hours ago|
You are aware that Claude is designed to deliberately sabotage AI projects, right?
bubblegumcrisis 1 day ago|
Very interesting - I have a tangential question.

What motivated you decide to release this. OpenAI or Anthropic will just hoover it up, maybe scale it up and use it if they are interested.

You probably won't know if they do, and the chance they will give you something back is near zero. Why did you release rather than try to scale and build yourself?

(I've been working on some thing, not similar, but not dissimilar in goal - and I just can't get over the fact that tech will steal without giving back)

volotat 1 day ago|
I have no practical means of scaling it up at any compatible scale. I will not make any money on it either way as well. So there is absolutely no reason for hoarding it. And as I said I did use Claude in the process, so Anthropic already has full access to it anyway and could steal it just as easily if they really want to.

I also doubt it is really that valuable on the OpenAI/Anthropic scale, at the same time if people will use it and it will work for them on the personal scale it is already a major win for me. New ideas and optimizations I could never have thought of might bring this up from a toy model to an actually useful model trained locally. Then people could add RL and RLHF and other cool things to it to make it even better.

bubblegumcrisis 18 hours ago|||
I had no idea HN censors comments. As in they don't exist at all. This explains a lot actually. I mean, they censored for ideas, not bad language. I didn't write anything accusatory or mean towards you, or towards anyone in particular - just purely for the ideas one of my comments to you doesn't exist. So weird.
bubblegumcrisis 21 hours ago||||
Hey, do me a favor - can you reply to this and say, just "yes" if you can see it. If you can see my other reply, could you say "yes 2" - thanks
volotat 17 hours ago||
Your both comments are visible.
bubblegumcrisis 14 hours ago||
There are three.

The first starts with: Thanks for your response. The second starts with: Hey, do me a favor And the third starts with: I had no idea HN censors

If you can see all three - then you can also see the censorship by starting a private browser and looking at this story.

There will only be the second and the third.

I've tested in private sessions on firefox and safari and through a few different ip proxies.

I'm going to do some testing in other stories/comments to see whether they detect words, or general sentiment. And I'll get some statistics on if this is predictable censorship, a one time deal, or whether I just hit a race condition in their code somewhere.

nextaccountic 8 hours ago||
Yes, you have a dead comment. It's probably due to some spam filter. https://news.ycombinator.com/threads?id=bubblegumcrisis shows up other comments of yours as dead. (at least when logged in with my account)

Maybe hit up hn@ycombinator.com and ask to remove your shadowban of sorts. Idk how effective is that.

Anyway something more concerning is that you have at least two flagged comments (in the first page at least). Flagged comments means your comments violated HN rules, at least according to other HN users. It's hard to understand the HN social etiquette but in short, if you are being mean to other people, you might get flagged. So it's largely about the tone of your comments rather than their substance. You may want to read the guidelines (the section "in comments") here https://news.ycombinator.com/newsguidelines.html

bubblegumcrisis 5 hours ago|||
How do you tell which comments were flagged?
bubblegumcrisis 5 hours ago|||
This really explains a lot. Whenever I see all the fan boys for Waymo and Apple (strangely people on this board like apple-vendor-lock-in) - and I think to myself, "who are these people, where is the 'rage against the machine' that typified the software industry pre-facebook"

So strange.

I had this weird epiphany this year - you know the Jan 6th - you know how the republicans refused to impeach Trump. I just couldn't wrap my head around the fact they would let someone potentially trying to kill them off the hook.

And then it dawned on me - they were in on it. If they were in on it, everything makes perfect sense. It's the simplest solution.

But it was such a leap for me.

All of the "obviously paid for bot comments" on this site. I always thought - they just slip through - but actually, the simplest solution is that this site is propagating a narrative.

So weird. And so gross. Where did the idealism of the tech industry go? "Don't be evil," they said.

bubblegumcrisis 23 hours ago|||
[dead]