Top
Best
New

Posted by Philpax 19 hours ago

GLM-5.3-Flash(z.ai)
https://news.ycombinator.com/item?id=49450353
1013 points | 508 commentspage 6
Destiner 19 hours ago|
from the article, pareto frontier for open source models is completely dominated by GLM now.
montroser 18 hours ago||
Well, it will be interesting to see where Qwen3.8-Flash-Next ends up landing, also released today. These are exciting times!
Lalabadie 17 hours ago|||
I find GLM's idea of fast/flash is not really competitive with the speed DS4 Flash has, and it's hard to see them as being in the same segment for that reason.
knollimar 16 hours ago||
Even vision? Thought k3 might have an edge there
halyconWays 12 hours ago||
Between Gemma 31/26/12/4/2, Deepseek-v4-flash-0731, Qwen 3.8 27B, Qwen 3.8 Flash Next (which I haven't even gotten to run yet!), and now GLM 5.3 Flash, I can't keep up. I love all these open weight models and am continually stunned that it's largely the West fighting for closed, restrictive, anti-user bullshit and China absolutely mogging the likes of OpenAI and Anthropic, with some notable exceptions like Gemma. Still, I shudder to think what the world would look like if we only had closed models. In many ways the stagnation of open source diffusion seems like that: LLMs are just a few months behind frontier, but image gen is like 1.5 years behind.
beannt 15 hours ago||
Is it good compare to Opus 5 ?
freakynit 3 hours ago|
Personal testing results: it's on gpt-sol-low level.
jdw64 17 hours ago||
This was the ox-alpha model, right? I remember it performed really well for a model that had 'flash' in its name.
tokai 18 hours ago||
Why is their own coding plan always the last place z.ai release their models? Its even online, you just have to guess the model settings.
woadwarrior01 15 hours ago|
Captive audience.
Imustaskforhelp 18 hours ago||
> To overcome the relatively limited compute and memory capacity of individual chips, we built a dedicated inference engine for this architecture on top of SGLang. Notably, this effort was accelerated by our GLM-5.3-powered infrastructure agent, which assisted engineers in developing and optimizing kernels, diagnosing performance bottlenecks, and improving the serving stack — creating a feedback loop in which the model helped optimize the system serving the model itself.

> (...) Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale.

It might be one of the most actually practical tasks that AI might've done because the compounding effects of it and also its implications are/feels so immense. It feels as if Nvidia might be in a slight turbulence from it.

freakynit 3 hours ago|
Now translate this to physical world, robots building and optimizing other robots... getting iRobot (2004) vibes
scottfits 17 hours ago||
so is it confirmed if this is the mysterious OxAlpha model?
Gander5739 16 hours ago|
Yes; if you try to use Ox Alpha it will give an error saying it waa trial period, and that it is GLM 5.3 flash.
scottfits 10 hours ago||
interesting, thanks!
nkjvhb 11 hours ago||
I heard that Dario Amodei is not having a great day today.

2 really strong open models on the same day is a amazing.

kayleykiwi 18 hours ago||
This looks like it goes hard, can't wait to try it
toppy 18 hours ago|
By clicking this link you download some PDF in the background
krystofee 18 hours ago|
Its displayed in the html...
More comments...