Posted by whiteros_e 6 days ago
Ironically, many benchmarks being maxxed out, and quite quickly, so new ones have to be created.
Several metrics are actually growing exponentially.
Power consumption and water consuption are the obvious ones. What are the others?> Do you have anything that proves this one way or another that isn't based on vibes or shoddy benchmarks?
They clearly aren’t talking about RSI here, but that model development has stalled in general.
How, exactly, does "it was Claude that solved Navier Stokes, not ChatGPT!" get you to "AI has hit a wall and stalled"? That's, not to put too fine a point on it, incoherent, and is just noise thrown into the discussion to avoid grappling with the fact that AI continues to rapidly improve.
TL;DR Buckmaster and Alpöge haven't solved Navier-Stokes blow up.
How information can get so distorted when it's trivial to fact check?
Some call this "The singularity" (e.g. Hinton).
This is actually a core danger postulated by the, let's call it, "worrying" scenario - see AI 2027 (to be clear, I think its timeline is not realistic).
> Statements dreamed up by the utterly deranged.
Evidently, and tragically, it will take catastrophes to show that deranged are the ones deriding the worried crowd.
Honestly, I have no idea what z.ai is either (I'm aware of an AI-enabled editor called Zed, but that's under zed.dev), so it's a bit presumptuous from them to assume that everyone is familiar with their product...
Appreciate the clarification. For me it was the "F" in "WTF" that tipped me. Other than that, it's more than fair for you to not know what GLM is. Things are moving so fast that I would be surprised if anyone can keep track of it all. Cheers, have a grand day!
Also I feel like the obvious way to read the very first sentence is that GLM is a language model
> As we develop GLM, the model sometimes exhibits capabilities that surprise us
Why would you be reading their corporate blog posts if you don't even know who they are?!
Where GLM-5.3-Flash is the newest "small / fast" model.
Come on now
Also, why would they introduce themselves on their own blog?
They must have hit really hard scaling limits if the prices were hiked so much so quickly.
Can I ask where are you using all those tokens?
I now exclusively use https://omp.sh/ as my harness:
I set it up so it never works in the main branch so subagents etc don't step on each others toes, and only merges back when complete: https://github.com/Daviey/mario/blob/main/.omp/hooks/pre/wor...
A good AGENTS.md is essential: https://github.com/Daviey/mario/blob/main/AGENTS.md
I then provide specifications for what I want, making sure it is unit tested.
For [API usage](https://openrouter.ai/z-ai/glm-5.3-flash#providers) they charge a bit more than the very cheapest providers of GLM-5.3-Flash, but not so much that a big price difference would make sense.
Very good - but I'm on a legacy plan. And coming up on a renewal that would put me on the watered down current plan. But with 50% legacy discount think it may be worthwhile. If I go to a competitor I'd be paying market rate.
>They must have hit really hard scaling limits if the prices were hiked so much so quickly.
Not really scaling - their plans were initially comically subsidized even more so than what the western providers are doing. More advert for an upstart than commercially priced.
The max plan will provide ~1,100 USD of GLM-5.3 or ~260 USD of GLM-5.3-flash per month for 168 USD. I can personally attest to these numbers through omp (~97% cache hit rate).
Unless you are able to highly parallelize (your work, you won't be able to hit your hourly or weekly quota using the flash model simply because it's so slow.
They give you ~3x more flash tokens, which maybe comes out to ~2x more actual work after accounting for the extra thinking it does to achieve the same result. The mental model, for not getting angry, is 5.3 is fast mode by default, and you can disable fast mode for 2x the work output at 1/3-1/10th the speed.
They're serving me 5.3 at ~40 tok/s and 5.3-flash at 30 tok/s (according to omp).
That is shocking. Is it per-token I wonder?
Just checking now: recent runs tau3[1] was at 96% and toolathlon[2] was at 90%
[1] https://www.induction.ai/docs/benchmarks/tau3 [2] https://www.induction.ai/docs/benchmarks/toolathlon
I’m getting 97%.
Also why Meta gets a +1, just charge less money on the training path.
If you factor that in, then there are clearly different tiers: one you can trust, and one that may well just be saying that to increase market share with little reputational or legal consequences if they are found to be lying.
These are not equal.
To be fair, none of us are sure of anything and I think that’s the part that’s most irritating
> ZCode, the GLM coding agent, silently uploads your Git history
And you can bet GLM is still ridiculously subsidized, just not as ridiculously as Anthropic and OpenAI.
But what really kills me is the idea that these companies are using Python for production inference. I mean really? Have you seen how bloated and slow Python is? Do global locks really sound like a strategy for fast dynamic computation?
If you look at how many years the whole NVIDIA and CUDA ecosystem has been evolving, it's certainly impressive how they've just stood up and optimized this CUDA-free 100,000 node cluster in just a few months.
Yes, but it's calling C code.
No real reason to respect any terms they might want to impose. Besides, if you want to break TOS, just have an agent do it; "everyone" running these things agrees there's no corporate or moral liability for what your AI does.
In the PRC, they[1] leaked tons of national secrets on the PRC's latest AI campaigns, the inner workings of their "opinion monitoring" (read: performative panopticon) and "stability" (read: violent oppression) departments, Chengdu's whole CCTV network, direct-energy weapons plans, espionage activities in Syria to hunt down Uyghur refugees, and god knows what else that Anthropic didn't divulge to us common folk.
In the US, it's very clearly an attempt to rip off a competitor. I'm not sure how else you could possibly see it. Even if you're a distillation fan in general (which A. why and B. plz don't), they did this through a network of Japanese and Signaporean shell accounts, presumably at least some of which were abusing Anthropic's subscription service in a ToS double-whammy, as it would be exorbitantly expensive otherwise. They also had to hack around Anthropic's API to get CoT traces, which seems impossible to explain away as anything innocent.
I've been beating the "China isn't necessarily an enemy, it's gonna take us all to handle AI" drum for literally years, but this attack was just... gross. Gross in scale and gross in arrogance. Not a good sign for the dawning alignment crisis, to say the least :(
TL;DR: Use these services if you want, but know that you're supporting aggressive escalations and companies that very clearly don't give a flying fuck about violating the law, much less your ToS. So... buyer beware, I guess.
[1]: For clarity, Z.ai was not alone in this, nor were they most egregious attack -- Moonshot.ai (kimi) took that coveted prize. DeepSeek was involved, too.
> use these services if you want, but know that you're supporting aggressive escalations and companies that very clearly don't give a flying fuck about violating the law, much less your ToS
From my European point of view the same risk/concerns apply when using US providers
Alignment is meaningless; as you've noticed, humans aren't all that "morally aligned".
If the tool needs safety measures it should be kept in a safe enclosure like we do with CNC machines, furnaces, and so on.
Alignment is meaningless; as you've noticed, humans aren't all that "morally aligned".
Moral relativism is an attractive proposition when you first examine the topic, but it quickly falls apart; there's a reason it's not even a coherent camp in contemporary philosophy beyond some vagueities from radical post-modernists. Just to go over some of the greatest hits:- Is what [DICTATOR/MURDERER/CRIMINAL] bad, or merely not to your taste? If the latter, then you have no coherent reason to argue they should be punished. We would never imprison people who don't like vanilla ice cream because 51% of the population does like it.
- If another culture had a deeply held belief to [HORRIBLE_THING] to, say, children, would you just shrug and say "different strokes for different folks"? What if [MURDERER] just had a different culture?
- No, the fact that nature is red in tooth and claw does not disprove morality; we are very, very, very far from our pre-rational, animalistic roots, and to go back now would be unthinkable.
If the tool needs safety measures it should be kept in a safe enclosure like we do with CNC machines, furnaces, and so on.
Thousands of scientists have been studying this problem for 76 years now, going on 77; your hunch about physical machines does not overrule their findings about the capabilities and tendencies of minds wrought from sand. You didn't explain why it's illegal or why distillation is bad.
I think this is just blatantly false, likely based in a misunderstanding of criminal law vs. civil law. Civil courts still deal with legality.The broader discussion of why distillation is bad and dangerous and immoral is left as an exercise for the reader, as it was above with the parenthetical. It's not a complex argument; I guarantee you understand it if you're reading this.
Nulla poena sine lege?
The same thing as above -- the fact that laypeople can not think of a criminal charge that they've heard on Law & Order that corresponds to this behavior does not mean that it's legal. It's textbook fraud, regardless of what particular detail you focus on. I'm always wondering when "distillation" comes up how feasible it is, or if it's just BS... Or am I missing something here that makes real "distillation" feasible?
I think the fact that it's happening at such a large scale is proof that very smart, well-resourced labs in China (the producers of the world's best OS models, including the incredible GLM-5.3-Flash) think it's feasible. I'm not sure it's productive to question them in the absense of any indication to the contrary.This is a great question still, not trying to shut you down. But I think the fundamental issue is a misunderstanding of what distillation is -- it's not directly stealing literal atomic parameters and piling them up. They might try to focus on substructures within these massive networks, but even that isn't strictly necessary for a distillation attack.
Source for 1? Are we sure those aren't hallucinations?
Sorry, I never linked it! This is from the latest Anthropic safety report (of "Anthropic Houtis build missile" fame), and no, these cannot be hallucinated -- the leaked secrets were inputs, not ouputs. https://www.anthropic.com/threat-intelligence-report-septemb... Like Anthropic and OpenAI are? After all, didn't they distill all the information in the world into their model(s)?
This is just blatant word games, sorry. I'm sure intended in good faith, and I understand the impulse -- I consider myself a radical anti-IP slacktivist, after all. But "both things involve information transfer" is just not a coherent point; lots of things fit that description. I mean, if they get to distill other's IP, why can't others distill their IP?
These are cyberattacks. Yes, anyone can cyberattack cyberattackers. But, y'know... an eye for an eye... Yes, wont somebody please think of the shareholders whose IP had been stolen...
I am not at all concerned with the value of the resulting artifacts as assessed by the (already totally unhinged) NYSE et. al. I am concerned about user respect, law following, truth telling, blatant cyber warfare at a time of rising tensions, accidental data leakages at a scale that'd be hard to fathom 5 years ago, bad-faith public postures, and a general distaste for fraud.The struggle you have with pinning down moral relativism I think betrays the fact that your understanding of it is low quality. A good ear mark is, can you name a single passage from a treatise on moral relativism that you like? If you can't find a single aspect of a philosophical construction to advocate for, that means you don't actually understand it.
You also keep sliding between conflating moral relativism with moral anti-realism and even moral nihilism at one point. These are different axes.
> Realism + Absolutism
All moral facts converge for every single person. As a matter of fact, they aren't facts at all. Morals are a hinge of the Wittgenstein variety.
> Realism + Relativism
Moral facts are derived from frameworks, or contexts, which are themselves grounded facets of reality.
> Anti-Realism + Absolutism
Kantian Constructivism. There are no facts which are coherent without a framework, but moral claims are still necessarily only capable of being universal.
> Anti-Realism + Relativism
Moral reality is constructed by the framework, grounded only to the framework.
I mean, if they get to distill other's IP, why can't others distill their IP?
The Antrophic article mentions "16 million" conversations, GLM models are in the 700-300 billion parameter ranges and while the frontier sizes aren't know but Gemini suggests Astra and Mythos are at around 10 trillion. That'd amount to extracting 40k parameters per conversation without a lot of errors if it was just a distillation (from an unknown source/algorithm as opposed to distilling your own model).
Now, I can imagine these conversations being used as a verification step that they're not missing stuff in their training, and that their models are capable of most of the same things, but that's mostly confirming that they've stolen the same data from the public as Antrophic/OpenAI has stolen already.
Or am I missing something here that makes real "distillation" feasible?
I wonder how you imagine that China built their own space station? Reliant on using American made duct tape, perhaps?
Do you realize how reasoning models are being trained nowadays? You design/build simulation environments to run agents in, with the environment providing the RLVR "verification" scoring. So why won't Ziphu use GLM to build their own RL training environments? Do you think they are not doing this?