Top
Best
New

Posted by whiteros_e 6 days ago

How GLM built its own inference infrastructure(z.ai)
411 points | 285 commentspage 3
OhNoNotAgain_99 6 days ago|
[dead]
_aavaa_ 5 days ago||
[dead]
tefkah 6 days ago||
[flagged]
binsquare 6 days ago||
Given the rate of improvement, why is this deranged?
sixeyes 6 days ago|||
because the rate of improvement is fairly stalled?
pizza234 6 days ago|||
You're tragically misinformed; it isn't. Several metrics are actually growing exponentially. But if you want emprical information, you can just have al look at the nature of the late AI incidents.

Ironically, many benchmarks being maxxed out, and quite quickly, so new ones have to be created.

ttoinou 5 days ago|||
The AI "incidents" are pure marketing ploys to get free word of mouth. Like what you're doing.
imp0cat 5 days ago||||

    Several metrics are actually growing exponentially.
Power consumption and water consuption are the obvious ones. What are the others?
glimshe 5 days ago||||
If you knew what "exponentially" means, you probably wouldn't be saying that.
heresiarch39 5 days ago||||
Which metrics?
georgefloydd 6 days ago||||
[dead]
simonw_simonw_ 6 days ago|||
[dead]
owebmaster 5 days ago||||
Kind of but there's a lot of improvements available to the competitors catching up
xyzsparetimexyz 6 days ago||||
Do you have anything that proves this one way or another that isn't based on vibes or shoddy benchmarks?
koe123 6 days ago||
You prove your own point no? You are asking for a benchmark to prove AGAINST ASI. Surely the burden of proof for such a scientific fiction concept should be the other way around.
phoghed 6 days ago|||
No, they asked for a reliable measurement to prove that model development has stalled. Nobody is talking about ASI except you.
jdiff 6 days ago|||
Nah, they're fine making that small leap. RSI is the new marketing term for the sci fi singularity.
phoghed 6 days ago||
>> because the rate of improvement is fairly stalled?

> Do you have anything that proves this one way or another that isn't based on vibes or shoddy benchmarks?

They clearly aren’t talking about RSI here, but that model development has stalled in general.

koe123 5 days ago|||
Sorry I meant RSI
xyzsparetimexyz 5 days ago|||
I don't think these benchmarks like humanity's last exam or whatever have much value.
scarmig 6 days ago|||
Yeah, it's been over a week since a Millennium Problem was solved. AI has hit a wall.
etamponi 6 days ago||
It was not solved. ~OpenAI~ Buckmaster and Alpöge found one (or a few) singularities in the forced version of the Navier-Stokes equations. Then magically 2 weeks later OpenAI found them too. Again, I am not saying this is not a great feat. I am just saying that everyone should be a bit more careful when making statements about RSI.
scarmig 5 days ago|||
As the sibling comment points out, that is incorrect: they solved a smaller, simpler problem, and OAI solved the actual Millennium Problem. But even if they had solved the actual problem and OAI stole it by digging through chats, that hardly supports the "AI has stalled" thesis; Claude played the major role in creating their blowup to a different problem.

How, exactly, does "it was Claude that solved Navier Stokes, not ChatGPT!" get you to "AI has hit a wall and stalled"? That's, not to put too fine a point on it, incoherent, and is just noise thrown into the discussion to avoid grappling with the fact that AI continues to rapidly improve.

red75prime 5 days ago|||
Buckmaster and Alpöge has found a forced finite-time singularity for the 3D incompressible Euler equations (and two other types) building on the work by Diego Córdoba and Luis Martínez-Zoroa with the assistance of Anthropic and OpenAI models. Then OpenAI found a forced finite-time singularity for the Navier-Stokes equations.

TL;DR Buckmaster and Alpöge haven't solved Navier-Stokes blow up.

How information can get so distorted when it's trivial to fact check?

tefkah 4 days ago||||
Because pursuing your own replacement is deranged.
ramon156 6 days ago||||
if you can be replaced by an algorithm, how useful were you really?
azan_ 6 days ago|||
Very useful. That’s a weird question.
fc417fc802 5 days ago||
I aspire to be at least as useful as bogosort.
simonw_simonw_ 6 days ago|||
[dead]
meyer3423 6 days ago||||
[dead]
simonw_simonw_ 6 days ago|||
[dead]
pizza234 6 days ago||
> This is known as Recursive Self-Improvement, or RSI.

Some call this "The singularity" (e.g. Hinton).

This is actually a core danger postulated by the, let's call it, "worrying" scenario - see AI 2027 (to be clear, I think its timeline is not realistic).

> Statements dreamed up by the utterly deranged.

Evidently, and tragically, it will take catastrophes to show that deranged are the ones deriding the worried crowd.

almaight 5 days ago||
[flagged]
rob74 6 days ago||
This article left me with one immediate question: "WTF is GLM?".

Honestly, I have no idea what z.ai is either (I'm aware of an AI-enabled editor called Zed, but that's under zed.dev), so it's a bit presumptuous from them to assume that everyone is familiar with their product...

jbonatakis 6 days ago||
z.ai is a fairly well known AI lab out of China and their GLM models are probably the most popular outside of Anthropic or OpenAI’s. I don’t think it’s presumptuous for them to not introduce themselves in a post on their own blog, I think you’re just a bit out of the loop here.
ma2kx 5 days ago||
And honestly that's for a reason. GLM5.3 on max has in my experience far less hallucinations than any other open weights model and it feels it has some intuition to bring in the right information when it's in principle out of context but relevant to the topic. Like its goal is more to bring value and assist you than just solving the given task with the least token spent.
bogdan 5 days ago|||
I don't get the outrage. Do you post this kind of stuff on every topic on hackernews that you are not knowledgeable about?
rob74 5 days ago||
Maybe my post sounded harsher than I intended, and yeah, it's probably on me that I'm not familiar with GLM. Actually the other major Chinese LLM Kimi does ring a bell, maybe it's because three-letter acronyms are a dime a dozen and annoy me because I'm confronted with them regularly at work too (people at my company seem to love acronyms), but that's obviously on me too...
tokai 5 days ago|||
It didn't read as harsh. Only unaware and you broadcasted that you don't have the decency to do basic searches.
bogdan 5 days ago|||
> Maybe my post sounded harsher than I intended

Appreciate the clarification. For me it was the "F" in "WTF" that tipped me. Other than that, it's more than fair for you to not know what GLM is. Things are moving so fast that I would be surprised if anyone can keep track of it all. Cheers, have a grand day!

fxwin 6 days ago|||
It's presumptuous for them to assume that a reader of their blog is familiar with their product?

Also I feel like the obvious way to read the very first sentence is that GLM is a language model

> As we develop GLM, the model sometimes exhibits capabilities that surprise us

peri-cl 6 days ago|||
It's only the top open-weights LLM in the world,

https://artificialanalysis.ai/#intelligence-category-tabs

HarHarVeryFunny 5 days ago|||
Ziphu, aka Z.ai, is the company that makes GLM (a very competitive Chinese LLM).

Why would you be reading their corporate blog posts if you don't even know who they are?!

Mashimo 6 days ago|||
A ai model family similar to Codex, Gemini or Claude.

Where GLM-5.3-Flash is the newest "small / fast" model.

0x457 5 days ago||
Except there is no ai model family called Codex or Claude. There is Gemini, I give you that.
drbscl 6 days ago||
>As we develop GLM, the model sometimes exhibits capabilities that surprise us, and even unsettle us.

Come on now

Also, why would they introduce themselves on their own blog?

embedding-shape 6 days ago||
I was gonna ask how people found their coding plans, and realized, have they massively ramped up the prices? Seems the middle plan is ~$80/month now, didn't that used to be like $20/month? Cheapest plan is ~$20/month currently.

They must have hit really hard scaling limits if the prices were hiked so much so quickly.

Daviey 6 days ago||
I paid $360 annual for Max plan and currently averaging about 1BN tokens a day with their frontier GLM-5.3 model. This was clearly unsustainable for them and they've dropped this package.
world2vec 6 days ago|||
1 billion tokens a day?!! I've done a lot of work these past 2 weeks with GLM-5.3. Like, a lot. And I've just passed 300 million tokens in total.

Can I ask where are you using all those tokens?

_0ffh 6 days ago|||
Well, there's essentially two major ways to use these models: Pair programming or fully autonomous fire-and-forget code generation. The second strategy needs essentially zero input, so the number of tokens you can blow is practically only limited by API speed.
rubslopes 5 days ago||
There's also a third way that can spend the most tokens: if the AI is used as part of the product, and not just a tool to build the product.
wartywhoa23 6 days ago||||
Something like this I guess: https://youtu.be/U-Rqv9dOB1U
absqueued 5 days ago|||
This was such a gem of a video.
p2detar 5 days ago|||
This is such a good video. Instant sub. Next to tech bros, we should also put AI-cringe bros.
Daviey 5 days ago||||
I have 3-5 agent harnesses with large context windows working on different applications concurrently.
embedding-shape 5 days ago||
Share the resulting code from any one of those please? I've tried so many times to find a setup that facilitates parallel work + high quality results, but it's just impossible regardless of harness or model. Leave the agents alone for too long, and the entire thing just balloons out of control, and next you know you're sitting there with half a million LOC where 80% isn't even needed.
Daviey 5 days ago||
Most of them are not public, but a fun thing I did was a mario cli game - https://github.com/Daviey/mario/ (or `ssh mario.baby`).

I now exclusively use https://omp.sh/ as my harness:

I set it up so it never works in the main branch so subagents etc don't step on each others toes, and only merges back when complete: https://github.com/Daviey/mario/blob/main/.omp/hooks/pre/wor...

A good AGENTS.md is essential: https://github.com/Daviey/mario/blob/main/AGENTS.md

I then provide specifications for what I want, making sure it is unit tested.

buckle8017 6 days ago||||
That's easy to do with many agents independently told to find bugs in a large codebase.
tokai 6 days ago|||
300M for two weeks is surprisingly low. What are you doing that need so few tokens?
world2vec 6 days ago||
It's not my main model (that would be Fable 5.1 Extra) but it's been doing agent-driven search and optimisation of a cross-trading ranking model (it's for work).
disiplus 5 days ago||
I would suggest you to hook fable or 5.6 to check it regularly and its work because it gets lost easily on stuff it was not trained on. I'm doing some custom inference engine optimization and it's a workhorse but it can easily lose its way and if you don't recheck it you will get wrong answers in the end.
embedding-shape 5 days ago|||
Kind of feels like this applies to every single model, from Astra to Qwen, they all eventually lose track of the plot unless you feed it some human's input that can steer them right every now and then. The only difference is how often you need to do so, and also how often you want to do so heavily influences how good quality the results will be.
disiplus 5 days ago||
you are not false, but there is still difference. its just that the better models are correct more of the time and will better validate its own steps. glm sometimes will understand the plan start implementing and then forget part of it and then say it finished. or then take a wrong turn somewhere and not correct. but they will all happily proclaim they are correct till you question it.
world2vec 5 days ago|||
Yeah that's what I already do. Fable writes the plan and checks things at certain milestones. Otherwise it does get lost indeed.
disiplus 6 days ago|||
I also have a legacy pro plan and the only limitation is if you are trying to work in the morning from Europe because you are in the 3x usage overlapping China time but after 12 or so you basically can run it at least for me at least 3 parallel sessions all the time.
bbor 6 days ago|||
It's hard to know, since no one advertises the actual token limits (partially cause they're prolly complex / adaptive). So it seems much more likely that they just offer different pricing tiers than you're used to. Like, the $80 plan is still ~$80 of subscription quota, regardless of what else is offered.

For [API usage](https://openrouter.ai/z-ai/glm-5.3-flash#providers) they charge a bit more than the very cheapest providers of GLM-5.3-Flash, but not so much that a big price difference would make sense.

Havoc 6 days ago|||
>I was gonna ask how people found their coding plans

Very good - but I'm on a legacy plan. And coming up on a renewal that would put me on the watered down current plan. But with 50% legacy discount think it may be worthwhile. If I go to a competitor I'd be paying market rate.

>They must have hit really hard scaling limits if the prices were hiked so much so quickly.

Not really scaling - their plans were initially comically subsidized even more so than what the western providers are doing. More advert for an upstart than commercially priced.

_aavaa_ 5 days ago|||
Their plans are still worth it if you use their models. You can see how many tokens you can except to get based on plan here: https://docs.z.ai/devpack/overview#estimated-token-allowance

The max plan will provide ~1,100 USD of GLM-5.3 or ~260 USD of GLM-5.3-flash per month for 168 USD. I can personally attest to these numbers through omp (~97% cache hit rate).

Unless you are able to highly parallelize (your work, you won't be able to hit your hourly or weekly quota using the flash model simply because it's so slow.

They give you ~3x more flash tokens, which maybe comes out to ~2x more actual work after accounting for the extra thinking it does to achieve the same result. The mental model, for not getting angry, is 5.3 is fast mode by default, and you can disable fast mode for 2x the work output at 1/3-1/10th the speed.

They're serving me 5.3 at ~40 tok/s and 5.3-flash at 30 tok/s (according to omp).

Schlagbohrer 5 days ago||
That table assumes cache hit rate of 95% or better. Am I understanding this correctly that people really are doing such repetitive prompts (compared to each other, across the concurrent user base at that time) that only 5% or less need actually be computed by the intended LLM?

That is shocking. Is it per-token I wonder?

workbreak 5 days ago|||
Every tool call is essentially entire prompt so far sent again with the response and that's why cache rates are so high for agentic workloads. This really bites when using expensive models since most models are 1/10 for cached input.
bleonard 5 days ago||||
We have been running a lot of agentic benchmarks with the various loops and tool calls on longer threads - we routinely see 90%+

Just checking now: recent runs tau3[1] was at 96% and toolathlon[2] was at 90%

[1] https://www.induction.ai/docs/benchmarks/tau3 [2] https://www.induction.ai/docs/benchmarks/toolathlon

_aavaa_ 5 days ago|||
If you are using their coding plan for coding, then yes you can easily hit such cache rates, with a good harness.

I’m getting 97%.

asp_hornet 6 days ago|||
The way I look at it, their coding plan doesn’t retain data or use it for training making it one of the cheaper plans for me.

https://docs.z.ai/legal-agreement/privacy-policy

andy_ppp 6 days ago||
You believe any of these companies care about the law? They care about winning and building the self improving AI as quickly as possible.
asp_hornet 6 days ago|||
I too am sceptical but I’ll take my chances. At least it’s helping the open weights.
criley2 6 days ago|||
I believe the that the companies who claim to not train on my data are more likely to not train on my data than the companies who refuse to even claim they won't.

Also why Meta gets a +1, just charge less money on the training path.

orf 6 days ago||
I’m not sure that follows. You’re assuming that all those claims have the same weight, without considering the size, jurisdiction, reputation or even the general vibe of the company making that claim.

If you factor that in, then there are clearly different tiers: one you can trust, and one that may well just be saying that to increase market share with little reputational or legal consequences if they are found to be lying.

These are not equal.

andy_ppp 5 days ago|||
Yes I sometimes think the "don't train on my data" is actually a good signal for "this data/person is probably better to train on because they want to keep something private". The whole copyright system should have stopped these guys from training on everyone's data and it did not, if you think they care about the privacy checkbox I think you're dreaming personally, based on their past behavior.
asp_hornet 5 days ago|||
> I’m not sure that follows

To be fair, none of us are sure of anything and I think that’s the part that’s most irritating

orf 5 days ago||
It’s more a polite way of saying “that’s crap”
asp_hornet 5 days ago||
And mine a polite way to say “you are equally uninformed”. We’re not getting anywhere. All the best.
orf 4 days ago|||
Absolutely galactic-level ooof :)

> ZCode, the GLM coding agent, silently uploads your Git history

https://news.ycombinator.com/item?id=49752422

orf 5 days ago|||
FYI it’s helpful to actually say your point during a discussion. And if you don’t want a discussion then why did you comment?
probst 5 days ago|||
Way to restrictive in terms of tokens provided. I am on their largest plan, and quickly run into their limits. And that is using it selectively in addition to codex.
broodbucket 6 days ago||
Yeah it went from a great deal to unviable compared to other providers imo. They really need to find a healthy middle ground
lompad 6 days ago|||
It just gives a taste of what we are all going to have to pay soon, once the model providers actually have to make money. And the era of "let's charge a dollar for every 10 dollars running the infra actually costs" is rapidly coming to an end.

And you can bet GLM is still ridiculously subsidized, just not as ridiculously as Anthropic and OpenAI.

chobbledotcom 6 days ago||
This isn't true, you can pay for GLM 5.3 from a provider like Neuralwatt or Friendli who have no incentive to subsidize or loss-lead their inference APIs
breakingcups 5 days ago|||
They didn't pay for training
jdiff 6 days ago|||
This introduces other incentives to cut corners and over-quantize.
pyrophane 6 days ago|||
What provider are you using currently?
0xbadcafebee 5 days ago||
This is a really funny sounding post. They sound like they just found out that increasing your automation gives you increased capabilities at faster speeds. They also sound like they just realized AI makes hard things easier.

But what really kills me is the idea that these companies are using Python for production inference. I mean really? Have you seen how bloated and slow Python is? Do global locks really sound like a strategy for fast dynamic computation?

HarHarVeryFunny 5 days ago||
It's not that they "just found out" - what they are saying is that while they were previously dogfooding because it's good practice, now that their models are so much stronger they are using them because it helps accelerate.

If you look at how many years the whole NVIDIA and CUDA ecosystem has been evolving, it's certainly impressive how they've just stood up and optimized this CUDA-free 100,000 node cluster in just a few months.

saagarjha 5 days ago|||
Most of the fastest inference and training code in production today is written in Python. There are no global locks on the GPU except the ones you put there
wolttam 5 days ago|||
Python acts as an orchestrator of accelerator libraries and does none of the inference math directly
esseph 5 days ago|||
> Have you seen how bloated and slow Python is?

Yes, but it's calling C code.

kamranjon 5 days ago||
Someone tell this man about vLLM!
dilyevsky 5 days ago||
they did rebuild their serving layer from fastapi to some rust thing so...
kgeist 5 days ago||
I have a similar approach where I optimize kernels and find numerical differences between the CPU oracle and CUDA kernels using an automated AI agent in a feedback loop. Usually it solves numerical problems easily (it compares outputs of every layer and finds where they diverge), but so far no matter how many different SOTA models I throw at it, and even show it reference code from other inference engines, they aren't able to much the speed (my engine has a modification which is not found in reference code, although a lot of stuff is similar). Either I'm doing something wrong, or z.ai's Infra Agent is actually an agent swarm, i.e. a bruteforce with heuristics. My project is 2 weeks old so maybe I just need more time.
gpugreg 5 days ago|
For me, DeepSeek-V4.1-Flash works very well for CUDA kernel optimization. Access to ncu (NVIDIA Nsight Compute CLI) also helps.
bbor 6 days ago|
Well, other than the infrastructure they got from illegally routing millions of paying customers' requests through Anthropic's Opus 4.8 in a distillation attack...
pjc50 5 days ago||
Anthropic infringed the copyright of basically every author on the planet: https://www.anthropiccopyrightsettlement.com/

No real reason to respect any terms they might want to impose. Besides, if you want to break TOS, just have an agent do it; "everyone" running these things agrees there's no corporate or moral liability for what your AI does.

_aavaa_ 5 days ago||
I'm not defending their actions, but we should be clear about where the law currently stands: Anthropic was found to infringe because of the torrenting, not because of the training.
woadwarrior01 6 days ago|||
That is such a canard, IMO. FWIW, Anthropic and OpenAI encrypt "thinking" token outputs in their models, while Chinese labs don't. If anything, it's more likely that everyone is using open-weight models in their synthetic training data generation pipelines. It's way easier to distill from logits than it is to distill from hard tokens.

https://x.com/EricSimons/status/2099252922098061714

phoghed 6 days ago|||
We weep for Dario, that he had to suffer such a devastating attack against his Terms of Service.
jensb1 6 days ago|||
What is "illegal" about it?
bbor 6 days ago|||
Are you joking...? Sorry if so! Just in case: It's illegal in both the PRC and the USA.

In the PRC, they[1] leaked tons of national secrets on the PRC's latest AI campaigns, the inner workings of their "opinion monitoring" (read: performative panopticon) and "stability" (read: violent oppression) departments, Chengdu's whole CCTV network, direct-energy weapons plans, espionage activities in Syria to hunt down Uyghur refugees, and god knows what else that Anthropic didn't divulge to us common folk.

In the US, it's very clearly an attempt to rip off a competitor. I'm not sure how else you could possibly see it. Even if you're a distillation fan in general (which A. why and B. plz don't), they did this through a network of Japanese and Signaporean shell accounts, presumably at least some of which were abusing Anthropic's subscription service in a ToS double-whammy, as it would be exorbitantly expensive otherwise. They also had to hack around Anthropic's API to get CoT traces, which seems impossible to explain away as anything innocent.

I've been beating the "China isn't necessarily an enemy, it's gonna take us all to handle AI" drum for literally years, but this attack was just... gross. Gross in scale and gross in arrogance. Not a good sign for the dawning alignment crisis, to say the least :(

TL;DR: Use these services if you want, but know that you're supporting aggressive escalations and companies that very clearly don't give a flying fuck about violating the law, much less your ToS. So... buyer beware, I guess.

[1]: For clarity, Z.ai was not alone in this, nor were they most egregious attack -- Moonshot.ai (kimi) took that coveted prize. DeepSeek was involved, too.

dgellow 5 days ago|||
What does any of this have to do with the legality of distilling Claude?

> use these services if you want, but know that you're supporting aggressive escalations and companies that very clearly don't give a flying fuck about violating the law, much less your ToS

From my European point of view the same risk/concerns apply when using US providers

cmrdporcupine 5 days ago||
It comforts Americans to believe their exceptionalism is both persistent/eternal and fully justified.
bbor 5 days ago||
Replying to a massive, blatant cyberattack with "lol America would prolly do the same" is not helpful nor rational. I am under no illusions about the exceptionalism of my state, especially considering the ongoing fascist self-coup. That's not the end of this discussion, not by a long shot.
pjc50 5 days ago||||
> alignment crisis

Alignment is meaningless; as you've noticed, humans aren't all that "morally aligned".

If the tool needs safety measures it should be kept in a safe enclosure like we do with CNC machines, furnaces, and so on.

bbor 5 days ago||
Replying to everyone to hack HackerNews' Gish Gallop feature (nulla poena sine lege!):

  Alignment is meaningless; as you've noticed, humans aren't all that "morally aligned".
Moral relativism is an attractive proposition when you first examine the topic, but it quickly falls apart; there's a reason it's not even a coherent camp in contemporary philosophy beyond some vagueities from radical post-modernists. Just to go over some of the greatest hits:

- Is what [DICTATOR/MURDERER/CRIMINAL] bad, or merely not to your taste? If the latter, then you have no coherent reason to argue they should be punished. We would never imprison people who don't like vanilla ice cream because 51% of the population does like it.

- If another culture had a deeply held belief to [HORRIBLE_THING] to, say, children, would you just shrug and say "different strokes for different folks"? What if [MURDERER] just had a different culture?

- No, the fact that nature is red in tooth and claw does not disprove morality; we are very, very, very far from our pre-rational, animalistic roots, and to go back now would be unthinkable.

  If the tool needs safety measures it should be kept in a safe enclosure like we do with CNC machines, furnaces, and so on.
Thousands of scientists have been studying this problem for 76 years now, going on 77; your hunch about physical machines does not overrule their findings about the capabilities and tendencies of minds wrought from sand.

  You didn't explain why it's illegal or why distillation is bad.
I think this is just blatantly false, likely based in a misunderstanding of criminal law vs. civil law. Civil courts still deal with legality.

The broader discussion of why distillation is bad and dangerous and immoral is left as an exercise for the reader, as it was above with the parenthetical. It's not a complex argument; I guarantee you understand it if you're reading this.

  Nulla poena sine lege?
The same thing as above -- the fact that laypeople can not think of a criminal charge that they've heard on Law & Order that corresponds to this behavior does not mean that it's legal. It's textbook fraud, regardless of what particular detail you focus on.

  I'm always wondering when "distillation" comes up how feasible it is, or if it's just BS... Or am I missing something here that makes real "distillation" feasible?
I think the fact that it's happening at such a large scale is proof that very smart, well-resourced labs in China (the producers of the world's best OS models, including the incredible GLM-5.3-Flash) think it's feasible. I'm not sure it's productive to question them in the absense of any indication to the contrary.

This is a great question still, not trying to shut you down. But I think the fundamental issue is a misunderstanding of what distillation is -- it's not directly stealing literal atomic parameters and piling them up. They might try to focus on substructures within these massive networks, but even that isn't strictly necessary for a distillation attack.

  Source for 1? Are we sure those aren't hallucinations?
Sorry, I never linked it! This is from the latest Anthropic safety report (of "Anthropic Houtis build missile" fame), and no, these cannot be hallucinated -- the leaked secrets were inputs, not ouputs. https://www.anthropic.com/threat-intelligence-report-septemb...

  Like Anthropic and OpenAI are? After all, didn't they distill all the information in the world into their model(s)?
This is just blatant word games, sorry. I'm sure intended in good faith, and I understand the impulse -- I consider myself a radical anti-IP slacktivist, after all. But "both things involve information transfer" is just not a coherent point; lots of things fit that description.

  I mean, if they get to distill other's IP, why can't others distill their IP?
These are cyberattacks. Yes, anyone can cyberattack cyberattackers. But, y'know... an eye for an eye...

  Yes, wont somebody please think of the shareholders whose IP had been stolen...
I am not at all concerned with the value of the resulting artifacts as assessed by the (already totally unhinged) NYSE et. al. I am concerned about user respect, law following, truth telling, blatant cyber warfare at a time of rising tensions, accidental data leakages at a scale that'd be hard to fathom 5 years ago, bad-faith public postures, and a general distaste for fraud.
ux266478 5 days ago||
Moral relativism is about handling the disparity of moral frameworks in the world, which you acknowledge exists at several points. The inverse of moral relativism is moral absolutism, which is that there exists precisely one set of moral facts. The space between these two is a spectrum, by the way, and usually we choose our position on the spectrum depending on what our purposes are. If we're talking about a disparate cultural collective, moral relativism is structurally necessary. Relativism just means we have a contextual differentiator, just generally in philosophy, not just in moral studies.

The struggle you have with pinning down moral relativism I think betrays the fact that your understanding of it is low quality. A good ear mark is, can you name a single passage from a treatise on moral relativism that you like? If you can't find a single aspect of a philosophical construction to advocate for, that means you don't actually understand it.

You also keep sliding between conflating moral relativism with moral anti-realism and even moral nihilism at one point. These are different axes.

> Realism + Absolutism

All moral facts converge for every single person. As a matter of fact, they aren't facts at all. Morals are a hinge of the Wittgenstein variety.

> Realism + Relativism

Moral facts are derived from frameworks, or contexts, which are themselves grounded facets of reality.

> Anti-Realism + Absolutism

Kantian Constructivism. There are no facts which are coherent without a framework, but moral claims are still necessarily only capable of being universal.

> Anti-Realism + Relativism

Moral reality is constructed by the framework, grounded only to the framework.

tuesdaynight 5 days ago||||
You didn't explain why it's illegal or why distillation is bad.
Bluestein 6 days ago||||
Nulla poena sine lege?
mitxela 5 days ago||||
Interesting. Where can I see this leaked data about the inner workings of Chinese opinion monitoring?
lelanthran 5 days ago||||
Like Anthropic and OpenAI are? After all, didn't they distill all the information in the world into their model(s)?

I mean, if they get to distill other's IP, why can't others distill their IP?

jLaForest 5 days ago||||
Yes, wont somebody please think of the shareholders whose IP had been stolen...
podocarp 5 days ago||||
Source for 1? Are we sure those aren't hallucinations?
whizzter 5 days ago||||
I'm always wondering when "distillation" comes up how feasible it is, or if it's just BS.

The Antrophic article mentions "16 million" conversations, GLM models are in the 700-300 billion parameter ranges and while the frontier sizes aren't know but Gemini suggests Astra and Mythos are at around 10 trillion. That'd amount to extracting 40k parameters per conversation without a lot of errors if it was just a distillation (from an unknown source/algorithm as opposed to distilling your own model).

Now, I can imagine these conversations being used as a verification step that they're not missing stuff in their training, and that their models are capable of most of the same things, but that's mostly confirming that they've stolen the same data from the public as Antrophic/OpenAI has stolen already.

Or am I missing something here that makes real "distillation" feasible?

jensb1 5 days ago|||
[dead]
bingud 6 days ago|||
breaking Anthropic TOS and misleading users
drbscl 6 days ago||
Breaking TOS isn't illegal per se. It just allows for denial of services, and may define terms by which the provider can reclaim costs.
lelanthran 5 days ago|||
I have very little sympathy for thieves who get robbed of the goods they have stolen.
butterNaN 5 days ago|||
Eh, even if this was true, then they're merely stealing from thieves. Anthropic did break a ToS or two to get training data themselves.
HarHarVeryFunny 5 days ago|||
If you understand what they have achieved here, then the notion that they are bottle-necked on training data is absurd.

I wonder how you imagine that China built their own space station? Reliant on using American made duct tape, perhaps?

Do you realize how reasoning models are being trained nowadays? You design/build simulation environments to run agents in, with the environment providing the RLVR "verification" scoring. So why won't Ziphu use GLM to build their own RL training environments? Do you think they are not doing this?

mitxela 5 days ago|||
That's not illegal.
Laurel1234 6 days ago||
[dead]