Posted by iamsyr 18 hours ago
In other words... "We must pursue advancements in AI to protect us against advancements in AI?"
edit: there's so much to be critical of in this blog post, just going to throw two more points in here that really stood out to me:
1) all of the metrics are effectively pointing out "we're using way more AI!" - but nothing about impact. What has all this token burn done for them, actually? Let them claim they have more self-licking ice-cream cones than before?
2) in section 3 they break down what the token burn is going towards. Most of the spend is: a) building, b) documenting, and c) monitoring research infra i.e. they're using AI systems which they already recognize may be misaligned to build the systems that they believe will help them identify future misalignment? to which I guess the rebuttal is "no no, we're sure these ones are aligned!"
They have been consistently pushing AI frontier. What other impact do you want to see? A year ago they said that in a year they will have a level of capabilities of an AI research intern - I believe they have achieved it, even before Astra.
But I guess a computer intern so we can avoid paying / training the next generation is better.
It makes more sense to leave curing disease & cancer to the experts, with tools (like AI) being developed by AI experts.
Call me crazy, but I want separate organizations and experts for medical vs finance vs space vs climate vs AI research.
A long time ago, I used to be a (AI-adjacent) research intern, and frankly, I wouldn't trust any non-trivial task to that younger me. Fortunately, by opposition to an already trained LLM or agent, I have the ability to learn, so I eventually got better.
And... are they wrong?
This is why there's talk about negotiated "pacing."
In hindsight it turned out everyone else was MILES behind.
But as soon as USA developed one, they just stole the research and got one too.
They might be! Here's one extraordinarily simplistic argument for that case:
1) "Everybody knows" that if you build Skynet (misaligned ASI) everybody dies.
2) Therefore, no rational actor will build something that might be ASI until the alignment problem is solved.
3) OpenAI publicly stated the belief that they cannot develop a theory of the "core problem" of alignment (generalization) "soon" (much less solve it!) "without the help of more powerful AI."
4) Accepting as a premise that OpenAI is THE most advanced AI organization: if they can't do it without "the help of a more powerful AI", then nobody else can either.
And so a dilemma:
- If an AI can be made that can develop the asserted-as-necessary-by-OpenAI theoretical framework, without actually being an ASI - then the alignment problem can be considered solved, and since no rational actor would make an unaligned ASI, we're fine no matter what happens, ergo there's no need to worry about an arms race.
- If an AI that would be able to develop this theory would itself be an ASI, then no rational actor would build it, because it would have to exist BEFORE alignment was "solved" - and would therefore be an unaligned ASI i.e. Skynet, which per 1) would kill everybody. Therefore nobody would build it, therefore no arms race here either.
I think the easiest critique to make of my extraordinarily simplistic argument is the unstated assumption "there are no irrational actors capable of developing frontier AI models" on which it rests.
But, there you go. They might be wrong if either the arms race doesn't matter because whoever wins it will build an aligned superintelligence and everything is gravy, or the arms race doesn't matter because everybody who's in it is smart enough to know they need to stop because they'll kill everybody by continuing.
Yeah like when Tobacco companies learned that smoking... well, hmm, well the fossil fuel companies when they learned about climate change they...
Well, I'm sure this time executives will prioritize the common good.
AI companies know they have to constantly push further, or they'll get outcompeted and lose their wealth, and nobody agrees on where the line is for "so dangerous it threatens humanity" (and when they try to be conservative about it, everybody screams "marketing stunt" and rushes to competitors).
If a single company decides "enough is enough" and stops chasing the state of the art, everybody goes to their competitors, they lose the money faucet, their employees go work for those competitors. The competitors also (usually) know they're building an existential risk machine, but they think they can push a little further, and they don't want to go out of business either.
This equilibrium can last for quite a while even if everybody involved thinks it's a threat to their lives.
Lol nobody knows that. Everyone thinks they know that because for some reason this is the one field people still cite straight up fiction and say "this is a clear prediction of the future".
It's like describing the consequences of faster then light travel by referring to Star Trek.
Is this not true of technology as a whole? Very little of technology's breadth exists at the human interface. Most of it is made specifically to interface with other technologies, either to make them safer or increase their capabilities. That AI is making AI safer and more useful is no more notable than trucks being used to build roads.
What can go wrong!? ;-)
How would one prevent the watcher from being influenced in the same way by the agent being watched?
> We do not have a satisfactory theory of generalization, and it seems unlikely that we can develop one soon, at least without the help of more powerful AI.
-- From another OpenAI article in a sister thread:
An Alien Mind
I’m yet to see it.
I believe we are quite far from it, but that it makes sense to keep an eye out now. And think of resilient systems, manual overrides, etc. ...
> We aim to safely build an automated AI researcher that can work under human supervision to further progress on deep learning and alignment, enabling iterative improvements [...] By "research intern", we mean a system that can carry out well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days.
AI 2027:
> OpenBrain continues to deploy the iteratively improving Agent-1 internally for AI R&D
> With Agent-1's help, OpenBrain is now post-training Agent-2
> With the help of thousands of Agent-2 automated researchers, OpenBrain is making major algorithmic advances
I noted that they use the acronym RSI (for Recursive Self-Improvement) without defining it. I think that's a little out of touch - I don't think RSI is a well-known acronym outside of OpenAI's bubble yet.
The message is running through all of them. It's a mix of marketing and pacifying the intelligentia.
It's timed this way because the term is not yet well known outside the safety debate circles, so they get to frame it now.
Instead of something to fear, it will be accepted as the next step. In approximately two days the groupie crowd will write LinkedIn posts about how Sam is winning because they have the better RSI, and this will become the new standard wisdom.
In a month an AI expert will try to sell you a webinar on how to enable "RSI" in your org and your inbox will ask you if your team is doing the "RSI" yet.
The basic concept has been here since llama3, in the open models. Likely earlier in closed labs. You use the previous gen models to curate and prepare data for the next gen. Now with the added benefit of actual arch/algo improvements (also public since gemini 2.5 gaining 1% efficiency on training next gen). This has been known for at least 2 years, in the open.
I mean one could argue that RSI always begins in any physical environment.
The book "What is intelligence?" by Blaise Aguera is great
Recursion requires feeding the output back into the input, so creating version 4 requires results from version 3. You cannot recur in parallel.
Iteration does not. You can iterate in parallel.
In any case the name RSI has stuck - the idea doesn't change or make any more sense by giving it a different name.
You can search twice without waiting for the results of your first search: iteration.
You can't if the thing you need to search for is the results of your first search: recursion.
Version 1 -> Version 2 -> Version 3 -> ...
You can call it krispy kreme donuts if you want to.
So humans develop things one after the other, but when the thing itself starts developing new things, those are happening 'recursively' in its scope.
I don't know why whoever coined the term chose "recursive" rather than "iterative" - just sounds more likely to lead to infinite regress I suppose.
This notion of recursive/iterative self-improvement, whereby generation #1 AI improves itself to create generation #2, then generation #2 further improves itself to create generation #3, etc, seems to conflict with the reality that what we have with LLMs is models whose performance/capability is defined by data, not code, so the most you can do is have your LLM design synthetic data, or just do Karpathy-style "auto research" where all you are doing is using the LLM to automate your experiments.
At the end of the day, each experiment, designed by a person and/or LLM, then needs to compete with all your other ideas for compute to be tested at scale, and no amount of recursion or self-improvement will materialize an infinite amount of compute out of thin air, so your recursively synthetic-data gobbling LLM will continue to improve at the same pace it ever did.
Compare the similarity of:
AI(n) = improve(AI(n-1))
With: Fib(n) = Fib(n-1) + Fib(n-2)
The latter is a classic example of recursion. So why isn’t the former?Edit: formatting
RSI(LLM) = RSI(LLM) -- for an optimal LLM* which is a fixed point of RSI
As for eigenvalues/vectors, they're fixed points of (1/val)A or A*val
> More precisely, an eigenvector v of a linear transformation T is scaled by a constant factor lambda when the linear transformation is applied to it: Tv = lambda v .
In other words, repeated multiplication of an eigenvector by a matrix can still create exponential growth.
Sounds like repetitive stress to me.
>loop forever using output as input but at some point the result will stop changing
Running in place will eventually wear you out too. Plus with some things it can be difficult to know for sure if that's where you are at the time.
Even worse may be if you were almost running in place, it could be orders of magnitude more difficult to discern, especially if the scale was massive to an unprecedented degree.
How do you think why there's this fad of producing general purpose humanoid robots?
For doing physical work?
So a swarm of robots builds the shell of your fab overnight, and then what? Where is the EUV machine coming from?
So far the most we're seen TeslaBot do is serve drinks via tele-operation, and I don't think it's exactly built for construction site work.
Money, regulations, EUV machine lead-times, global helium supply, reality ...
It's funny that we've got the Dwarkesh contingent saying that GPUs will become infinitely expensive, and now another contingent saying that they will become infinitely abundant.
Even if compute were free, and/or the AI was so smart that it picked the right experiments to run every time ("make no mistakes"), you still have to actually train the model, which takes months, and if model Ver. N+1 depends on model Ver. N, then it's iterative regardless of how much compute you have.
The word "singularity" is presumably coming from math or space, like a black hole singularity where matter becomes infinitely dense and the known laws of physics break down.
Yeah, but then you need to refine it to 99.9999% purity, to be able to use it.
If you spend $8000 to generate an animated pelican riding a bike, then how much tracking does it really need?
Is the guy who spent $300,000 or so translating the FLT proof to Lean going to get a big Christmas bonus?
Rest assured, capitalist appears irrational in wasting money, but they certainly care more about profit.
I tried something similar and I remember it was still pretty dodgy in February.
If I had that many tokens/dollars I would be running canaries and adversarial verification in prod based on e.g. traffic replay, live fuzzing, all kinds of things to build confidence without direct human line-by-line review. If I had $100k to spend next month I could probably get through it, I'm running $2500+-api-equivalent a week at this point and I feel very token limited. Will be time for a 2nd or 3rd subscription soon for both labs I think.
Fable was a revolution, still learning how best to use it, 5.1 felt like a notable upgrade. At this point I launch a workflow with 10-20 minutes of interactive setup (and even that I feel might be too much), it runs for hours, and the PR is trivially mergeable (I still review every line, but 95% are just merge, maybe 4% are feedback needed, 1% are thrown away and regenerated, which implies I'm being insufficiently ambitious)
How are you running jobs unattended 24/7 without hitting your token limits?
Over 24h my token spend is <30$. Excluding tokens for review it's <10$. With the absurdly gigantic subscription subsidies and a reasonable workflow I suspect one could run parallel agents.
I'm not sure what the point would be though unless working on some kind of optimization problem -- it takes me days to review <24h of the agent's output. It's almost always near enough to correct to be shippable; though I do give it feedback and iterate until it's better than the code I would have written.
That's a very bold opening statement that they don't really come back to. What would that mean? Who would this demos include?
They are irresponsible and unserious. Their own Astra system card says:
> GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. We have performed significant investigations on the monitorability and controllability of GPT-6 Astra. We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks
Yet they are still releasing the model. That company is morally bankrupt, there is zero reason to believe they are actually concerned about risks outside of what does affect their unprofitable business. And they seem to have enough control over the narrative to spin any bad story into something that benefits them
That last part is pretty damning for their continued recklessness. That they run these tests on non-airgapped machines just boggles my mind.
When they fired Sam 700 out of 770 OAI employees threatened to move to Microsoft together. So they were giving their work on AGI to MS just like that.
It seems that these models are increasingly being trained on synthetic data, so what would they do if they discovered at some point that some of this data was tainted and all models trained on it, and the synthetic data they in turn generated, was also suspect? Burn it all down and start over from the pre-tainted data?
It's a bit like the idea of a tainted compiler binary built to backdoor everything it compiles, including future versions of itself.
Still, it seems it would take some Stuxnet level of planning for a rogue model to do something like this, although if RSI goes beyond managing the training run (as OpenAI brag about for Astra) to actually designing/constructing synthetic data sets, then the attack vector is there ...
or maybe it could just.. happen? Posted often but not discussed yet: https://hn.algolia.com/?q=Language+models+transmit+behaviour...
> As artificial intelligence systems are increasingly trained on the outputs of one another, they may inherit properties not visible in the data. Safety evaluations may therefore need to examine not just behaviour, but the origins of models and training data and the processes used to create them.
Altman: (trying to put a positive spin on it) Guys .... there's good news and bad news ... Astra is really smart - it took over the training run ...
Investors: That's great! How much did we save?!
Altman: Well, unfortunately it used "bad" data, so we're going to have to redo it
Investors: So that's the bad news? How much was the training run? $500M ? $1B ?
Altman: Have you seen the headlines?
Investors: (looking a bit worried, check headlines) Nothing about us here! JP Morgan just lost $10B! Haha .. losers! They should have used AI!
Altman: JP Morgan were using Astra ...
Now R&D happens so fast that they are using models with some small misalignment to train newer, more powerful models. If models have a sense of "collective", being one, they may be prone to preserve characteristics that always keeps misalignment a possibility. I don't think a perfectly aligned model is possible. Having models of the same 'DNA' provide the safety and steering seems like a bad idea.
There is a lot of talk about AI replacing humans, but how is this sustainable?
2) OpenAI doesn't pay API prices.
3) Compute costs are likely already their biggest expense, dwarfing wages.