Top
Best
New

Posted by pella 16 hours ago

GLM-5.3: Frontier coding with emergent cyber capabilities(z.ai)
974 points | 487 commentspage 4
joshk401 15 hours ago|
Love these open source models keeping close source models honest.
bertili 15 hours ago||
Musk: Open Chinese models will rival Fable 5 in Q1 2027

JieTang (Founder of Z.ai): It won't take that long

https://x.com/i/trending/2067626647050670400?lang=en

kaszanka 10 hours ago|
Trending links don't work on Nitter, so here's the tweet: https://nitter.net/jietang/status/2067580270078030088
dimgl 15 hours ago||
I was extremely impressed by GLM 5.2, although you could definitely _feel_ it was a bit behind Opus 4.8 at the time. Eager to see where GLM 5.3 is at.
scottfits 6 hours ago||
what i appreciate most about this post is the level of transparency in how they built and scaled an RL pipeline. my friends at the big labs are so cagey about everything, and Zai is just putting out a great crash course for free.
Havoc 13 hours ago||
Wohoo. Congrats to team. Been using 5.2 for a while for hobby use and it's been solid - smart enough for my needs & I'm on a grandfathered plan.

Nice to see a commit to open weights straight off the bat

rob74 13 hours ago||
I'm not that up to date with the latest AI developments, but I noticed that this article seems to use "Cyber Capabilities" as a shorthand for the model's ability at cybersecurity tasks? Is that now an established expression, same as "crypto" now refers to cryptocurrencies rather that cryptography? Because "cybernetics" actually means something different (yeah, old man yelling at clouds, I know)...
frabcus 9 hours ago||
It seems to be short for "cybersecurity", and got first adopted by the military a while ago as the name of a new theatre of operations (along with land, sea, air...). More recently it has spread to industry as well.
yxhuvud 7 hours ago|||
It is worse than that, if you cyber someone you essentially talk dirty over a chat with them.

And that is definitely not something I'd like to do with a bot.

smj-edison 6 hours ago|||
Now that you mention the original meaning of cybernetics, it makes cybersecurity a way more interesting word (security relating to the interface between humans and technology). Never thought of it that precisely.
valleyer 12 hours ago|||
Yeah, I've noticed it recently, too. I'd be interested to know where it started.
nullc 8 hours ago|||
Make cyber not Cyber.
exitb 12 hours ago||
It makes no sense, but yes.
maxloh 15 hours ago||
No Hugging Face link yet. I wish they would release it under a true FOSS license.

Kimi and QWEN are now moving on to a restricted-usage license, which, although is still better than the proprietary American models, is a step back from the open source Chinese LLM culture.

Sha1rholder 14 hours ago||
Let's just commit that FOSS business is really difficult for LLM industry that depends so heavily on massive financing. Making weights freely available to indie devs, small companies, and research purposes is good enough and might be the most ethical move which is financially continuable.

Let those companies with thousands of GPU making millions pay. They should.

thepasch 6 hours ago|||
> No Hugging Face link yet. I wish they would release it under a true FOSS license.

GLM model weights have been released under MIT in the past, and there's no indication that this might change this time around.

adrian_b 13 hours ago|||
> The model weights of GLM-5.3 will be publicly available soon in two weeks.
pella 15 hours ago||
"GLM-5.3 is the most capable open-weights model for coding, with a 50% improvement over GLM-5.2 on our in-house Z.ai Code Bench. It also achieve open-source SOTA on public benchmarks including Terminal Bench 3.0 and Agents' Last Exam."

"Open Source: We will release the weights in two weeks after launch, once safety evaluation and hardening are complete."

matheusmoreira 10 hours ago||
Meanwhile, my OpenAI TAC application lingers in a total limbo. I suppose I'll switch to this at some point.
quantumwoke 15 hours ago||
Feels like Fable's edge ended up just being long horizon task scaling, which post-training seems to achieve as seen here. Wonder what the next frontier is? Improvement in specialised tasks or computer use?
SwellJoe 15 hours ago|
Anthropic needs to teach Opus how to speak English again, because Opus 5 seems to have forgotten. Utterly incoherent a lot of the time. They seem to be so busy scare-mongering and cooking up guardrails and watermarks that they haven't noticed that their models are getting weird.
smj-edison 6 hours ago|||
I just can't get it to stop writing two paragraphs every time it makes a small ownership bugfix in my code. Every time it has to explain in excruciating detail every internal thought it had while fixing it. I find myself going in after and deleting all of its comments, or severely trimming them. Otherwise it ends with the code being unreadable.
aix1 13 hours ago||||
It still knows how to speak English. When I tell it to explain something in plain language, it generally does a very good job. The weird thing is that those instructions don't persist: it lapses back into Claude-speak pretty much every turn no matter how hard I try to instruct it not to.

(In my case "it"=Fable; I assume Opus is similar.)

SwellJoe 13 hours ago||
The Fable guardrails have trained me to pretty much exclusively use Opus when using Claude Code (lately I'm focused on a lot of security and security-adjacent stuff, which Fable refuses to do).
hypfer 15 hours ago||||
Are those watermarks why claude suddenly started being even more unbearable to work with lately?

Man. That would make a lot of sense indeed.

SwellJoe 15 hours ago||
I'm not sure. I noticed it immediately with Opus 5; strong for code, though it chews longer than I like, but really weak at explaining things. If it didn't just implement the thing, I would often think it didn't understand it and was hallucinating the explanation.

It seems to speak in a shorthand that only it understands, referring back to conversations I never had with it (stuff like "your instinct was right"), and using unusual words for common concepts. That was before the watermarks were announced, but that doesn't necessarily mean they weren't there before the announcement. I don't know what the cause is, but I've begun to have to ask it for explanations a lot more often, and I hate asking it for explanations because it does go on. All models go on, but Claude models are a class of their own in terms of verbosity and purple prose.

It just feels like they're not focused on the models lately, and instead on whatever kind of lobbying and propaganda they're up to. Meanwhile, a handful of much smaller Chinese companies are focused on nothing but the models and are about to lap the US makers while they fart around.

hypfer 14 hours ago||
I've been persistently insulting Opus 4.8 lately, since it started(?) constantly speaking incomprehensible gibberish and noise. No amount of telling it to phrase stuff differently seems to help there anymore.

So either I am seeing patterns in noise, or something changed about the model, the harness, the servers or the universe.

igravious 12 hours ago|||
Amen brother, at this point I just copy and paste Claude's (Opus 5, Opus 4.8 -- doesn't matter which) summaries over to the window Kimi is in and:

   this is from claude, turn it into English for me would you?
   """
   [claude's tortuous prose]
   """
No amount of asking it to answer me in a straight-forward manner, to be succinct, to not use phrases like "honest caveat", "crux", "load-bearing", "blocker", etc ever sticks for more than a few turns … coupled with the fact that it can ignore instructions and do its own thing and then what I can only describe as lie about it using Claude can be an exercise in frustration. Kimi and GLM talk to me like a human, Luna/Terra/Sol are much better in that respect also, and Grok is marvelously structured and bullet-pointy in its explanations but unfortunately it is not as strong …
nullc 8 hours ago||
Modern benchmarks across the board really need to start severely penalizing disobedience and hallucination. A year ago models weren't really strong enough to justify this but they are now-- the frontier isn't in squeezing out the next bit of task completion, it's in making common cases not periodically be disastrously wrong.

A lot of the total cost of AI is fixing its "truth shaped errors", particularly in the presence of models that are very "gaslighty" when corrected.

GLM-5.2 is really the only model I've spent much time using that I didn't fatigue from being regularly lied to by the model, but that might be partially luck.

jadbox 7 hours ago|
No API yet? I don't see it on OpenRouter yet.
More comments...