Top
Best
New

Posted by Philpax 16 hours ago

GLM-5.3-Flash(z.ai)
https://news.ycombinator.com/item?id=49450353
958 points | 486 commentspage 3
yipinwong 16 hours ago|
When reading this type of announcements, always have keen eyes on graphs.

e.g. "Agent Coding Performance by Effort Level" cuts Y-axis from 0~20.

- This makes it as if GLM-5.3-Flash made a bigger jump than it claimed as the Y-axis does not increase much (stupid trick used in biz reports)

I did mention that ox was working ok for me, and having an open-weight comparable to close to SOTA makes it very compelling for me to try it out locally (well, only if I got more VRAM)

nchmy 16 hours ago|
they also conspicuously omitted GPT 5.6 Luna from comparison. It scores lower, but is also cheaper. MiMo 2.5 is not a valid comp at this point

edit: nevermind. it is there in the artifical analysis scatter plot, but is greyed-out.

MUCH more interesting is that in that chart, their cost is WAY off. The actual chart shows GLM 5.3 Flash at $0.09, but their chart shows $0.045...

mrtesthah 15 hours ago||
The web page says 5.3 flash is discounted right now.
drob518 15 hours ago||
Seems disingenuous to draw frontier graphs with starter pricing.
seaal 15 hours ago||
Well, Luna debuted with 5x higher pricing than is currently available. With the pace of recent development these models might not be relevant by Thanksgiving.
drob518 14 hours ago||
Of course. Pricing is always changing, but typically it goes down over time, not up. So, if you're showing artificially low pricing from the start based on a teaser rate, IMO, you shouldn't be using that to show where you appear on a frontier graph. Place yourself on the graph based on your expected long-term pricing. Then, over time, adjust your position based on your standard rate, whatever that might be. Games are always being played for things like this, but this seems excessive.
yipinwong 14 hours ago||
I don't know if that's the standard pricing for US models to go down overtime, while Chinese ones go up (start cheap but pay more).

I don't have enough metrics to compare those costs but still Chinese models have been cheaper except against Luna for me.

FWIW, Luna does everything so well, I just keep using it for all my agents by default.

drob518 13 hours ago||
I haven't noticed the Chinese models going up in price for the same model. They do release new versions of the models with different prices that are higher. But everybody is doing that. One fine point is that deepseek-v4-flash-0731 is really a different model than deepseek-v4-flash and it's priced higher.
BrucecarlL 2 hours ago||
It is bench maxed during the stealth testing. And it can’t beat DS flash on speed
singularity2001 15 hours ago||
At the current 50%-off GLM-5.3-Flash price ($0.075/M input, $0.25/M output; cached input $0.015/M), surprisingly, roughly $400–900/month would buy token throughput comparable to fully exhausting Claude Max 20×
arizen 10 hours ago|
Probably apples to apples would be to compare z.ai subscription plans vs API pricing
iamsyr 16 hours ago||
Standard API Pricing for GLM-5.3-Flash (per 1M tokens)

- Input: $0.15 - Output: $0.50 - Cached input: $0.03

Xunjin 16 hours ago|
Is that cheaper than DS4 flash?
nateb2022 15 hours ago|||
Slightly more expensive than the (post-price hike) DS4 flash pricing, but in the ballpark.

https://openrouter.ai/compare/deepseek/deepseek-v4-flash-073...

drob518 15 hours ago|||
Hm. GLM is more expensive in all dimensions than DS but it has a lower weighted average input? How is that?? Something seems off.

EDIT: Looks like they are swizzling around the pricing dynamically on that page, on both the GLM and the DS sides, so who knows.

walrus01 15 hours ago|||
Comparison should be to 0731
nateb2022 15 hours ago||
updated thanks
javier123454321 16 hours ago||||
All I can say is that even if it is, I was almost glad to go back to using DS4 Flash. Because 0XAlpha was just so friggin slow to complete a task because of the level of circular reasoning that it would go over and over into, sometimes even returning no output. If I just wanted something done I would switch from a free model to a paid one which is crazy.
denysvitali 16 hours ago||
Tbh it was also slow because it was being hammered by everyone making use of the free tokens
javier123454321 16 hours ago||
Possibly influenced by that, but I believe that is a different issue. I meant the way it processed a request. It went into so many more loops of thinking.
swiftcoder 16 hours ago|||
It's even cheaper than DS4's off-peak pricing. Seems like DeepSeek have some stiff competition now
arizen 15 hours ago||
Few weeks ago, I wouldn't expect this statement to be true. Accelerate!
OldGreenYodaGPT 12 hours ago||
Tested this last week and couldn't get it to finish any task that took more then an hour with /goal keep getting errors
syntaxing 14 hours ago||
Ironically, our administration pushing for ban of the AI chips to China is forcing them to make smaller and more efficient models which seems like a requirement for running on Chinese chips. I wouldn’t be surprised this model was tailored to run purely on Chinese chips. Same thing with Deepseek MLA, the drastically lower KV cache memory requirement was born out of necessity so it runs on the Huawei chips.
adroitboss 13 hours ago|
This is exactly what Jenson said in all of his interviews. Banning it in the short term would have long term consequences.
simonw 11 hours ago||
Good bicycle, good pelican: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
bigyabai 10 hours ago||
It's feeling good on non-pelican workloads too. Less verbose than 5.3, cheaper/higher usage limits, vision capability and some good web design one-shots even with vision disabled.

With GLM 5.1 and 5.2, the big problem was tool calling and long-horizon coherency. 5.3 was more trustworthy at the cost of longer thinking traces, and now Flash seems to improve on it once again with a more concise, smaller model. As long as there aren't any noticeable regressions, I could see myself defaulting to this for >90% of my day-to-day coding work.

joquarky 11 hours ago||
I only see a mostly blank page with a "Paste" button, a "URL" button, and and a "Preview" label.
simonw 11 hours ago|||
By chance you have a browser extension that might block fetching data from raw.githubusercontent.com ?
kelvinjps10 10 hours ago|||
I just see the raw svg code
mariopt 16 hours ago||
It's only 320B, local frontier AI is getting closer, sooner than expected.
saberience 13 hours ago||
It's not possible to keep shrinking down parameters and keep "frontier" performance, it's like saying it's possible to take a 3 hour movie and compress it down to 3 megabytes, there are information theoretic limits on the amount of bits of information that can be compressed.

What I'm saying is, if you're expecting a model that can be run on a 16GB or 32GB machine with the intelligence/knowledge of Mythos or Sol, it will never happen. It cannot happen, just like you cannot watch the Odyssey saved as a 16MB file.

Smaller models can get faster and smarter, but by definition they can never compress all of the knowledge of a frontier model and they will approach a limit by which they cannot get better.

hypfer 13 hours ago|||
You can fit Shrek 1 into a 13mb gif tho

https://www.deviantart.com/sssfjknfvdknj/art/the-ENTIRE-shre...

twobitshifter 12 hours ago||||
The current models are not close to approaching the limit of compression for intelligence. They aren’t even focused on it like Chinese labs are. The training of Qwen’s 27B parameter model showed that by structuring model training from fundamentals to more difficult topics they were able to drastically reduce the number of parameters needed.

The ‘frontier’ models rely on scale to achieve their results but that’s not the only approach. Eventually we will hit up against the fundamental limits but we are not close with Sol and Mythos.

saberience 7 hours ago||
Yes they are approaching the limits, try asking smaller models niche questions about almost anything, they hallucinate massively because you cannot simply pack in all the raw knowledge from a massive frontier model into something that’s quantified down to 20GB etc.

It breaks fundamental laws of information theory. It’s like saying you can extract 100 joules of energy from 10 joules of energy source. Not possible.

fy20 5 hours ago|||
It doesn't really matter though. Hardware performance is still growing. The new Mac Studio could just about run this model locally (rather slowly) - something that sits on your desk, that you as a consumer can buy.

Imagine prosumer desktop hardware 10 years from now. The 2036 DGX Spark. For a few thousand dollars you will be able to buy something with hundreds of GB (maybe TB if manufacturers step up) of unified RAM, memory bandwidth in the 10-20TB/s range. Overall AI "compute" will increase 10-20x, while at the same time AI model capability per byte will increase 5-10x.

The hardware would fit today's models, something like Kimi K3, quite comfortably and give performance of maybe 100 tokens/second. So what needs data center hardware today will run on your desk.

But if we also assume the models become more efficient, a 2036 Fable-class model (in terms of intelligence/capabilities, not size) will easily run on this thing at hundreds of tokens per second.

Unfortunately it'll still slow to a crawl with 5 Chrome tabs open, and every Electron app will need at least 200GB of RAM.

twobitshifter 5 hours ago|||
Sounds like you are describing a quantized model which is a naive form of compression, not a model that is trained more efficiently.

Additionally the information theory angle is for information storage, but a model can access resources and tools to gain information and what we are really seeking to train is reasoning not information retrieval. We reduce the needs to the right capabilities and we don’t get upset if it does not know the lyrics to every song ever written.

computerex 10 hours ago||||
You heard of JEPA? LLM's have all sorts of garbage they have memorized. Reasoning in latent space instead of in text significantly reduces the number of needed parameters.
saberience 7 hours ago||
JEPA is a joke, let me know when those models do anything useful.
oceansky 16 hours ago||
Can't come soon enough!
pranav_tech26 5 hours ago||
Benchmarking is cool, but for production I care about real inference latency, self-hosting VRAM costs, and how cleanly it handles structured JSON output.
Aboutplants 11 hours ago|
When do Chinese models surpass US models? I thought there was at least be a 2 year runway but now I think they surpass it within 12 months, if not sooner.
More comments...