Top
Best
New

Posted by ModelForge 20 hours ago

GPT-6 Astra, looped transformers, and hidden reasoning(magazine.sebastianraschka.com)
432 points | 141 commentspage 2
SubiculumCode 13 hours ago|
The major concern with looped transformers is that makes it more difficult to monitor model alignment. When more processing occurs within latent space without outputting text, that means less effective, frequent chain-of-thought monitoring, and the potential for greater un-monitored latent-space shenanigan.
imtringued 1 hour ago||
This is silly, the entire reason why chain of thought even exists is to let the LLM "think independently" instead of minimizing the deviation from the supervised training sample. It's an intentional scratch pad for intermediate data. The loose monitoring is kind of the entire point.
technotony 13 hours ago||
I'm not sure. That paper from anthropic talked about monitoring j space, presumably those same techniques would work here?
SubiculumCode 13 hours ago||
I am no expert, but I think it is this: 1) We have few effective tools at monitoring alignment right now, and chain of thought is one of the more effective. 2) Monitoring latent space may be possible, but I do not think it is even close to being a solved problem, nor whether it is possible at scale and outside of controlled problem areas. 3) Finally, more recursion within latent space may complexify the latent representations, not simplify them.
andai 16 hours ago||
> So, the whole idea here is that we increase the effective depth from 22 to 44 block applications without adding another set of transformer weights.

From what I gathered, LLM inference is bottlenecked on memory, right? Which implies there's "spare" compute we haven't been using? Does reusing the weights like this allow us to utilize it? (Do more math per unit of memory?)

brausepulver 14 hours ago||
You need to separate memory capacity and bandwidth. Looping decreases memory capacity/FLOP but not bytes loaded/FLOP, since weights need to be loaded again for the 2nd pass. Plus (depending on the method used) capacity required for KV will be that of the equivalent unlooped model (44 blocks) and KV is typically larger than weights at long context.
JyB 9 hours ago||
Why would weights need to be « loaded again » for the 2nd pass? Weights never change at inference time no?
the_real_cher 13 hours ago||
I think it's bottlenecked on memory throughput. Someone else more knowledgeable can verify this.
vatsachak 12 hours ago||
How do you guys even use LLMs where you are finding Astra light is worse than Sol High?
lwarfield 13 hours ago||
I'm kinda surprised that the mixture of depths paper didn't come up here. It approaches the other direction of sometimes dropping layers:

https://arxiv.org/abs/2404.02258

cubefox 18 hours ago||
This article is not up-to-date. There have been various benchmarks (some of which published and acknowledged by OpenAI, see the charts in this thread: https://xcancel.com/tomekkorbak/status/2095596839886274689) showing GPT-6 Astra is much less monitorable. The most recent third party benchmark I saw is showing a huge jump in capability for multi-hop reasoning without chain of thought: https://www.lesswrong.com/posts/FsCkkoGsNmPzFKRhg/gpt-6-astr...

I don't think this is explained by the model simply being more capable and therefore achieving more per token: the usage of recurrent depth (Neuralese) is exactly predicting less CoT monitorability even at equal capability.

ThunderBee 17 hours ago|
I work on small scale recurrent transformer architectures.

Better Multi hop reasoning is one of the most notable improvements of the architecture. The tricky part is figuring out a way to optimize the number of times you loop as it varies between tasks. Too few and you leave performance on the table too many and performance begins to drop.

siva7 17 hours ago||
Astra was insane until Monday but something happened on tuesday, now it feels like Sol. I grieve for the lost productivity but i hope they may give us the original Astra back.
sobellian 16 hours ago||
I thought the same, but on second thought I merely had to deal on Tuesday with a lot of the mistakes Astra made on the preceding days. I wonder if this time lag of consequences explains why the sentiment is so common with these models. It probably also cautions against irrational exuberance when you first crack open a new model and it one-shots various problems, as you don't yet know what goats Astra had to sacrifice to make it so.
cainxinth 17 hours ago|||
It's the same story every time OpenAI or Anthropic releases a new model. They are generous with compute for the first few days, and use maximum fidelity with uncompressed weights. Everything runs at its best to make a good first impression. But eventually they pare things back and the models perform a little worse.
Vetch 15 hours ago|||
The most charitable explanation I can think of for this is something like regression to the mean. When a model is first released, there'll be a subset of users who, just by chance, sample the highest quality band of the distribution that answers their query. Some of them will rush over to social media and post about how amazing a model is. Over time, those users' mental model of responses will converge but they'll perceive the model's return to typical performance as a downgrade.

This guess/explanation predicts that most users won't match what the initial social media hype claims, doesn't discount user experience as simple habituation nor does it assume companies are lying when they say there have been no changes to the model itself (quantization included).

I also think there's an aspect where initial testing is more forgiving because the more persnickety polish bits can be ignored and tests are likely to have similar structure to things that can be trained for. Meanwhile, actual specific work items are a broader unusual distribution with more stringent acceptance criteria.

Personally, I can detect a separation between Sol and Astra (but not as large as that between Opus and Fable). While they can solve most of the same problems, Astra takes less time, is less frustrating to talk to, is cleaner, notices more, spins wheels less and requires less corrections.

majormajor 6 hours ago|||
> I also think there's an aspect where initial testing is more forgiving because the more persnickety polish bits can be ignored and tests are likely to have similar structure to things that can be trained for. Meanwhile, actual specific work items are a broader unusual distribution with more stringent acceptance criteria.

IMO this is 90% of it (as someone who has a bit of a different interaction style and runs these things less autonomously, and hasn't generally seen the claimed regressions). Day 1: throw new stuff at it that failed badly, exciting to see something make more progress! Day n: reality sets in that it still wasn't perfect the first time.

kmeisthax 13 hours ago|||
To add onto this, if you use a shiny new model and it gives you a turd, you're not going to tweet about it ("hey guys, look what I made with Astra! Nothing!"), and even if you do nobody is going to interact with it so it does poorly in the algorithm, because it has to compete with all the people using the new model to make something that looks impressive. Then people get tired of the magic trick and the logic flips.
zaphirplane 12 hours ago|||
Really? there would be complaints, it’s expensive and doesn’t do as well
foolswisdom 9 hours ago||
When it's happened to me, I shrugged and went back to the way I did things before. Then again, I'm not a vocal social media user by any means.
majormajor 6 hours ago|||
We've seen that some---gpt5 was considered pretty lackluster intially, in particular. Opus 4.7 and 5 vs 4.6 were also greeted with a lot more "meh" than 4.6 or Fable.
OneOffAsk 8 hours ago||||
It’s all speculation (you too), but I think the effect you’re describing is instead getting calibrated to the model’s limits. Next time a new model comes out, wait a month before trying and see if you have the same feeling of rapid quality decline after a few days. I did after I jumped back into it mid 5.x or whatever ChatGPT after paternity leave. Blown away for a few days, worried about my job for a few days, then increasingly aware of its limits.
baby 17 hours ago||||
You think they introduce stronger quantization after a few days?
boredatoms 16 hours ago||
For sure they quickly move to q8, the output quality difference to bf16 is small compared to the speed/capacity gain
NineStarPoint 16 hours ago|||
Yeah q8 made so littler difference back when I was testing such things I'd be surprised if people could quickly notice that as a change. It's got to be either further quantized or some other type of optimization that kicks in when people notice the drop.
selectodude 15 hours ago||
NVFP4 would buy them a huge increase in capacity but I think it would be noticeable.
Caracas288 14 hours ago||
Why doesn't someone just try to measure this next time!?
embedding-shape 12 hours ago||
Can't really measure without being sure you aren't being messed around with, when it's a remote platform. Stupidly easy to detect when people run such benchmarks/tests against you as well.
torginus 14 hours ago||||
Some people here have remarked previously that while reduced precision doesn't show up in quick prompts, it does severely impact these models' ability to perform long running tasks - to the point that running these big models with severe quantization might be counterproductive as smaller but less quantized ones perform better.
nonethewiser 15 hours ago|||
Could this explain Opus?
blurbleblurble 17 hours ago||||
Or a lot worse
holler 17 hours ago||
so, AGI is cancelled?
__MatrixMan__ 17 hours ago||
AGI for the peasants is cancelled.
elwell 13 hours ago||
Trogdor - the AGInator
dooglius 16 hours ago||||
Do you have hard evidence of this assertion?
simlevesque 16 hours ago|||
We can't have hard evidence. It's a SaaS and they own the code and the machine it runs on.

So it may be a widespread hallucination. But there's no evidence of that either.

dooglius 12 hours ago|||
Run a benchmark with a large number of samples, rerun a few days later. Compare results, use statistics to see if there's a statistically significant difference.
kadoban 6 hours ago||
Doing this in a way that doesn't get you noticed or fucked with is potentially going to be quite difficult.

If they have a "hey we're being benchmarked" mode, which is not hard to imagine, avoiding tripping it is going to be annoying and difficult to prove.

fragmede 15 hours ago|||
We could still have soft evidence though. Make a Todo app on Monday, and make a Todo app on Tuesday, and see what it makes in comparison.
marcus_cemes 15 hours ago|||
You would need a significant sample size to make any sort of conclusion from such a probabilistic process. Then there's the issue of how you would actually grade/compare.
luckydata 14 hours ago||||
someone already does that https://aistupidlevel.info/
chaimtweiss 14 hours ago||
It's actually a extremely cool site, and fascinating to view the results off the AI bots i use.
ArvidSu 15 hours ago|||
You only need to come up with a catchy "SomethingBench" name, post it on reddit/x and now you're an ai sage. Not to disparage the launch/after comparison though, I'd genuinely enjoy a data point like that
bradly 15 hours ago|||
There is a toot from an Open AI person a couple days ago saying they are "pulling all the levers" because of capacity issues. I have no idea what the heck the person is talking about, but I'm guessing there are consequence for those levers.

    > "Demand for Astra is really unprecedented. We're pulling all the levers possible to sustain the demand, but I've not seen anything like it until now and we went through very steep growth before. Priority will always be to keep excellent service for existing users, but we might have to pause new Pro subscriptions for a bit if this continues."
nonethewiser 15 hours ago||
Depends on the nature of the levers
holoduke 12 hours ago|||
I am sure every input send to openai is prechecked by a dumb model and then send to another one. They heavily tweak this to improve performance.
scrlk 17 hours ago|||
Might be related to this announcement from Tibo on Sunday:

> We've made some improvements that improve usage on the long tail for power users of Astra when logged in with your ChatGPT account.

> No change in quality and a pure win that on the long tail can result in up to 3-4X less usage being drawn from the subscription.

https://x.com/thsottiaux/status/2096717905614524491 (https://xcancel.com/thsottiaux/status/2096717905614524491)

siva7 16 hours ago||
It seems to me the people working at OAI may believe all other humans must be a little bit behind intellectually.
pixl97 16 hours ago|||
I mean, in general they aren't wrong.

You can't fool everybody all of the time, but you can fool almost everybody most of the time.

But most of all, it's easy to fool yourself.

rowanG077 12 hours ago|||
It always reminds of the story of the creator of counter strike. Every new release he would get a ton of complaints from players about things they didn't even change. Notably that each version had more lag. And he got so fed that he start to negatively subtract peoples pings. And suddenly a ton of players reported back that the change was incredibly good.

Point is, I really don't buy all the stories about a model suddenly being downgraded without at least a modicum of substance. People are grasping at straws in the noise.

binary0010 16 hours ago|||
Disagree completely. I started using Astra from Sol the day it was released, and was a virtually imperceptable difference and made lots of mistakes and shit architecture decisions from day 1 of release.
kloop 8 hours ago|||
I still think this is because, on a new release, it works on some prompts the previous ones did badly at, because new weights do well on a different set of prompts.

Then after a few days you notice the prompts that it does badly on that the old ones did fine with and everyone is convinced there's a regression when it's just a different part of prompt space

aetherspawn 11 hours ago|||
I find Astra to be weirdly stupid in the sense you have to force it to spend time on something (fix this architectural issue and refactor), then it’s stupidly smart.

It prioritises getting something working over making something good during the 1-shot phase and outputs maximum slop.

mccoyb 17 hours ago|||
I had nearly the exact same experience and thought I was imagining it … absolutely ripping, then it turned into Sol++ on Tuesday …

I’m working on hard things, it is very noticeable when it is hums through something and then falls over on something it should not

I can tell by analyzing my own prompts to look at when I get frustrated ;)

theLiminator 17 hours ago|||
I wish someone ran some sort of representative benchmark suite every X days to see if this occurs.
micycle1 17 hours ago|||
https://marginlab.ai/trackers/codex/
CamperBob2 17 hours ago||
Unfortunately that seems to be monitoring Sol, not Astra, unless I'm missing something.
embedding-shape 17 hours ago|||
I'm not sure how you could run such a benchmark without leaving it possible for the labs to easily detect and fudge the results.
konart 15 hours ago|||
>now it feels like Sol

It can very well be Sol, no? What stops them from using cheaper model for some requests during "rush" hours or simply use cheaper model for every Nth request.

kgeist 9 hours ago||
>What stops them <..> simply use cheaper model for every Nth request.

That would trigger a full prefill (context recompute) every Nth request because cached tokens aren't interchangeable between models, and that would require way more compute than just staying on Astra.

To avoid full recompute, you could prefill a cheaper model's context incrementally by always feeding it Astra's outputs in the background (and vice versa), but then that would require 1.5-2 more VRAM for each session + the complexity of keeping them in sync.

If the rumors are true that Astra is a looped transformer, a more practical approach would be to dynamically adjust the loop count during peak hours.

throwatdem12311 15 hours ago|||
I’m so used to seeing this on every single model release I’m starting to question if these kinds of posts are just trolling.

Alternative theory - it always seems amazing when it first comes out then the novelty wears off and we’re just meh about it. New model is a model is a model. I bought a PS5 Pro and was genuinely blown away by it at first…few weeks later I’m just like…eh it looks pretty good I guess? It’s still the same, I’m just used to it now and the wow factor along a new thing is going. Kinda like that.

Or they are just compute constrained so they have to serve a shittier version. Who knows?

I hate how opaque these companies are. It feels deceptive and evil.

jcmontx 17 hours ago|||
Same story every time, I bet they quantized it
manmal 17 hours ago||
Exactly my thoughts today. They have to make it cheaper after demoing what’s possible initially.
nickreese 17 hours ago|||
I had the same experience. Moving back to Sol for actual implementation.
Paracompact 9 hours ago|||
Can you re-run some prompts that you ran on Monday and report the differences in output?
qaq 14 hours ago|||
OK so it's not just me
Razengan 17 hours ago|||
Which plan/region are you on/in?
siva7 16 hours ago||
Highest subscription tier and i believe there is only US region available being served globally
acedTrex 15 hours ago|||
This shit is just vibe coder astrology lol
ModernMech 17 hours ago||
lol I didn't get access until Monday (I was at 0% since Friday and my reset was Sunday at 11pm), so go figure.
atomflunder3000 16 hours ago||
I only used Astra while coding a bit so I can't comment on anything else but I have been really disappointed by it.

It seems to overengineer really bad and it is also very slow due to it "thinking" too much I feel like.

One example is that I asked it to implement a new functionality inside an existing App of mine and if I had written it myself it would have been like a ~50 line diff. Astra took like 10 minutes to write ~400 lines, most of them useless and also in pretty bad style, barely readable code.

Maybe I am bad with prompting but I didn't have these issues before, not even with 5.6 Sol on max reasoning.

jiggawatts 1 hour ago||
Try Astra on low or at most medium thinking level. Its "low" is better than Sol "high" or even "xhigh", and then it also doesn't overthink as much.
vatsachak 11 hours ago||
What was your prompt? I have gotten easy fixes with I tell if to do something.
iJohnDoe 18 hours ago||
Probably off-topic. Astra has been kind of weird. Like, I can't trust it, weird. It has an interesting tone, especially in Codex, that is off-putting. It's over zealous at times (which is why I stopped using Claude) and gets too creative when doing agentic system level stuff. Accessing files and doing things it shouldn't do. If OpenAI was chasing Claude's approach, then they are going in the wrong direction. OpenAI has always been the "business and boring approach", which was its selling point and why I have stuck with it. Claude was always the radical one (powerful, but radical).

Also, Astra overlooked, in my opinion, a serious flaw in its approach for something I was working on recently, which really surprised me.

Reading between the lines, there were some breakthroughs with Astra, which I'm sure is why OpenAI released it so quickly after Sol, but probably not in the ways the traditional OpenAI customer wanted.

redox99 12 hours ago||
Astra is definitely weird. It is more capable than Sol, no doubt about that. There are things sol could simply not solve that Astra breezes through.

However for typical low to medium difficulty code, it will often either overengineer stuff, create massive functions instead of organized code, and just write very hard to read code. It literally looks like minified code. Clearly they trained it to reduce the number of output tokens and in turn the code is often atrocious. I'll keep trying Astra but I might actually go back to 5.6 sol for many tasks if I keep getting these results.

enraged_camel 17 hours ago|||
>> It's over zealous at times (which is why I stopped using Claude) and gets too creative when doing agentic system level stuff. Accessing files and doing things it shouldn't do.

I gave Astra a pretty straightforward bug ticket yesterday. The bug involved an edge case that could sometimes result in an invalid value getting stored in a user profile field. Pretty harmless, no crash or anything, just annoying.

Based on past experience, I don't trust OpenAI, so I decided to watch Astra as it worked. About four minutes in, it convinced itself that it should also check the prod database to see "how far the corruption has spread" and attempted to SSH into the hosting provider. This resulted in my 1Password to prompt me, which I of course denied. Then I stopped Astra, closed the ChatGPT/Codex app and gave the task to Opus 5. Suffice it to say I will not be renewing my subscription, because "you have to watch it like a hawk" is the opposite of agentic engineering.

silversmith 16 hours ago|||
Why is your agent able to call ssh. Why can it trigger 1password. Why are you giving metaphorical guns to metaphorical toddlers. Why is it not sandboxed. Your practices worry me.
olalonde 8 hours ago||
Are you guys all running agents in VMs?
lann 8 hours ago||
Yes.
olalonde 4 hours ago||
Container or full blown VM?
mike_hearn 1 hour ago||
I use containers in one context (custom container manager) and a regular UNIX account on bare metal in another.

This isn't intended to stop a model like Astra hacking its way out of course, it's more like guardrails on a staircase.

My personal container manager tool has an intercepting SSL proxy and small Javascripts on the host can rewrite or block HTTP requests. The agent gets its own isolated home directory and can't tamper with mine. Local caches like Maven are mapped read/only with a write layer on top.

wilj 16 hours ago||||
ChatGPT desktop this morning lost a chat thread while I was actively working in it. I asked Astra to find the lost session, and next thing I know it's prompting for full computer control to drive Finder. It's just jsonl files on disk, not hard to read normally.

Negative feedback filed and ChatGPT uninstalled.

AnimalMuppet 16 hours ago|||
You have to watch it like a hawk so it doesn't do something to production, on its own, without a specific request? Wow. Then I could never trust it to not be doing something to some other system that it shouldn't, so I'd have to audit every network request.

If enraged_camel had been doing something else involving the production database at the wrong time, they might have accepted the 1Password prompt.

enraged_camel 7 hours ago||
Worth noting that this has never, ever happened with Anthropic models, which I've been using all day every day since Opus 4.1.
andriy_koval 16 hours ago|||
wondering if creativity can be managed by setting reasoning level.. You pick lover reasoning for simpler tasks and high reasoning for open ended research.
Alokcue 16 hours ago||
cool
ModernMech 18 hours ago||
I don't really like Astra either. It doesn't seem noticeably better than Sol, and it uses more tokens. Some people said ultimately it's cheaper because it can solve problems faster but I haven't really noticed that.

The way I use it now is I'll ask a chat 6 Pro session to make a plan and then have Sol implement it, then 6 Pro reviews it. This seems fine and it doesn't use my Codex minutes, so I'll use Astra. But on the metered tasks I don't see the utility.

This is a problem for OpenAI because if Sol is good enough, and they don't have a moat, then it's only a matter of time before Sol-level models are open sourced and running locally. I know I'll be doing that as soon as I can.

BikiniPrince 17 hours ago|||
I'm still working through my first few days, but I've had to deal with Opus ADHD for a while. I built a task management system which is closer to old school remedy with reviewers. The stylistic guidelines on task creation have a seven part problem statement, goal, success, ancillary data and such. By framing the task diligently it does keep the work on target. The review logic is basked into the task management software so the agent can't declare done. On open ended issues it can still wander. It's been remarkable to drive down issues over these last few weeks. I was annoyed I had to stop for 3 days and build management infrastructure, but it's paid for itself.
zamadatix 17 hours ago||||
I had a few problems which Sol was bumbling around with and giving mediocre results (e.g. in a toy planet app, Sol was taking several iterations to get a half decent looking render of the weather I still wasn't pleased with) but Astra managed to implement well in one go.

Much the same as you're saying, I never got around to verifying how much of that was because of Astra being better vs just being a different model sent specifically to those tasks because the token usage didn't make sense to spend unless it was something not working in Sol. So even if it was all due to Astra being fantastic I'd still not like to use the model for the cost being even more fantastic.

redhed 17 hours ago|||
I have tested it out with CAD and PCB circuits and it is a huge jump compared to Sol. I agree though when trying it with programming I don't notice a huge jump.
kilpikaarna 17 hours ago|||
Anecdotally (I did try it myself, but wasn’t blown away) many seem to like it for 3D modelling. That was emphasized in the promo too. I think this kind of ”general intelligence” is what is meant to set it apart from 5.6.
redhed 16 hours ago||
Yeah my scenario was we had old paper drawings without actual CAD models. Fed those into Astra and it did it 100% perfectly. Honestly might be the easiest scenario for it, but that's also what I thought for Fable and Sol and those completely butchered it. Wish I could share pictures of those attempts but just imagine a completely mangled model that barely looks good if you squint. These were not simple models either, pretty large/complex machinery.
ModernMech 17 hours ago|||
I'll have to try it for a PCB circuit because that's where I'm going next. Were you asking it to use specific software to build the circuits?
redhed 17 hours ago||
Using KiCad by uploading their _sch and _pcb files. Originally with Sol, I stuck to using it for finding parts and double checking my KiCad schematic. Definitely good at finding parts quickly from JLCPCB's stock and for quick cosmetic edits of the schematic. I found its PCB editing abilities pretty bad, though it was useful for cosmetic edits (quickly relabeling silkscreen labels) and for creating a nice custom DRU file. With Astra on the other hand it can actually make good PCB edits. Still not great but usable and editing it quicker than starting from scratch. I do doubt you can go 0-100 with just Astra but definitely sped up my work. For reference my circuits are high amperage, noise sensitive, and interface with sensors. They are pretty simple circuits though, just fairly simple ICs with no MCU or anything like that.
rvz 17 hours ago||
Recommended reading from an actual researcher who thoroughly understands AI research papers and has an in depth analysis of models architectures and their mechanics and no nonsense benchmarks.
simianwords 16 hours ago|
On looped transformers:

previously, conversation might have 50k tokens spent on reasoning. the next turn takes all the previous tokens as well (if you wanna preserve prompt caching) which is not ideal. this new method skips that so you get more free context until compaction kicks in.

is this true? if so its a huge deal. why is it not spoken about? its one of the main reasons i don't use High or Max

More comments...