Top
Best
New

Posted by pella 13 hours ago

GLM-5.3: Frontier coding with emergent cyber capabilities(z.ai)
922 points | 466 commentspage 2
virgildotcodes 13 hours ago|
OpenAI and Anthropic need to just go ahead and give people access to the cyber models.

Otherwise we have a world of attackers using open and closed source models against a much smaller group of maintainers that are likely heavily dependent on Anthropic and OpenAI and for whom it may not be a simple matter to just get approval to start using the open model flavor of the month.

LeonidBugaev 13 hours ago||
Not only attackers. I have to switch to Kimi or GLM even in cases of basic issue triage on my own projects! Current guardrails are ridiculous.
SwellJoe 13 hours ago||
I've been building a harness for security work, and had to switch to GPT 5.5 when even Opus started refusing security work. Then 5.6 Sol arrived, and it refuses security work, too. So, I switched to Kimi K3 and DeepSeek for API testing just because it's so much cheaper. But, if GLM is better, I'm here for it, as I think GLM is also cheaper than K3.
Synthetic7346 7 hours ago||
Mind sharing a link?
mindwok 12 hours ago|||
At least OpenAI seems to want to do that, but the US is now forcing them to go through approvals. Anthropic seems much more hesitant.
bryceneal 10 hours ago||
OpenAI seems to understand that these guardrails hurt the good guys. This is why they released Daybreak Blue, which is a step in the right direction (but the model itself is weak as it's just Sol with fewer guardrails). Anthropic seems to believe that harming defenders is worth it if it means they can achieve regulatory capture. They do a lot of mental gymnastics to try to pretend that this is not actually what they are doing. As a result they have lost a lot of customer goodwill, which hasn't yet caught up with them yet, but absolutely will IMO.
35129ab 5 hours ago|||
They won't. People will notice that the models are overhyped once they can test them.
surgical_fire 9 hours ago|||
Can't the maintainers use the same models as the attackers?

The maintainers don't need approval to use GLM.

virgildotcodes 8 hours ago||
They may need approval from their employers.
matheusmoreira 7 hours ago|||
[dead]
worldsavior 13 hours ago||
[flagged]
fcanesin 4 hours ago||
GLM-5.3 is further proof that all >1T models are currently undertrained. I was looking at inteligence density ( https://www.pasteboard.co/6q2-5f92mtj9.png ) from recent open models (where parameters sizes are known) and taking DS-v4-flash as upper limit GLM-5.x can 3x its performance.
vmware508 11 hours ago||
Apple will release M7 MacBook Pros / Mac Minis next year, and they will be able to run free LLMs locally at native speed. All software developer notebooks will be replaced to run local models, saving a lot by cancelling Claude Code subscriptions. Developers win. Apple stocks will be rocketing. Everything else will go down. You're welcome.
schleck8 11 hours ago||
You'd need the 256 gb memory model which will be expensive because apple has trouble getting capacity (got turned down by cxmt). And even then you can only run a 2 bit quant which is noticeably worse than 8 bit
Havoc 7 hours ago|||
That’s not how that works. The hosted models don’t stay still in size and capability while Apple advances. Both will advance their frontier and there will still be a gap and developers will still prefer the stronger option.
gehsty 9 hours ago|||
Local vs remote compute is a constant thread in tech history - mainframes and desktops then local and cloud compute (think Google Photos bs Apple photos - one indexes on device the other indexes in cloud). Now we have the next chapter local vs cloud LLM models.

There will always be a market for frontier labs in the cloud based models - these models will always be able to be bigger, and that will likely translate to doing things local models can’t.

Logically also we’ll likely get to a point where RAM drops in price as production ramps up, and local LLM is both capable and cost effective. This feels like it is coming for Siri / Gemini / Alexa personal assistant type use cases.

So I think the local LLM will become a thing in laptops and phones in a year or two, offering PA type use cases. Professional LLM services will likely remain at the frontier (and in the cloud) for the foreseeable.

Gecko4072 11 hours ago|||
They will cost an insane amount as well. Maybe less than subscriptions or tokens. But running massive models on laptops with batteries and poor cooling doesn’t make much sense.
LeBit 10 hours ago|||
Until hiding PII from the cloud LLM is a resolved issue, running local LLMs will remain a necessity.

There are workplaces that refuse to use LLMs because they fear the devs will expose sensitive data without care.

ignoramous 48 minutes ago|||
> will cost an insane amount as well

We will get to a point where prosumer laptops that etch SoTA LLMs in removable silicon will be as expensive as cars.

kube-system 4 hours ago|||
Every single MacBook built in the past half-decade already has an LLM built into the latest version of their OS.

But there's a significant difference in hardware required between running a 3B parameter model and a 700B-1T+ parameter model.

toasty228 10 hours ago|||
Sure buddy, all you'll end up with is a $10k machine that run gimped models at like 30tok/s for about 5m before the fan kicks in and it starts to sound like a turboprop, while offering maybe 30% of the context size of hosted models.
layer8 8 hours ago|||
The RAM shortage situation won’t be sorted out within the next year.
rammler 5 hours ago||
Bold to believe it will be sorted at all
fearmerchant 3 hours ago||
If the margins are there it will get sorted.
WarmWash 5 hours ago|||
I'm still waiting for Linux to topple Windows
Flavius 10 hours ago|||
> run free LLMs locally at native speed

This reads like a hallucination. What does native speed even mean?

kyxsc 10 hours ago|||
for example, models running at like 100-150 tokens/second (or faster!) vs 15 t/s

(fable/sol are ~60 t/s, and OpenAI just announced their Cerebras partnership(?) for "ultrafast" mode of 750 t/s)

models aren't able to run that fast right now on our consumer/prosumer hardware. M5 Max for example has a memory bandwidth of 600 GB/s. a 5090 has 3x that, so running the same model on a 5090 is that much faster (provided the model is within 30GB).

running a bigger model on an M5 Ultra is still much slower than running it on a Blackwell chip with sufficient vram, CUDA being a major difference. if apple can bridge this gap, interesting things will happen... and just imagine if M7 Ultra has comparable speeds to Blackwell (or even Rubin)!

toasty228 10 hours ago||
Meanwhile the GB300 used by hosted llms:

GPU Memory Bandwidth: 7.1 TB/s Interconnect Bandwidth: 900 GB/s bidirectional

https://pi3g.com/nvidia-gb300-specifications-including-memor...

If you think M7 will hit even 15% of these speeds you're very optimistic.

andsoitis 8 hours ago||
A hosted instance serves multiple customers at a time. A local model only one.
toasty228 7 hours ago||
How many though? At 1m context you quickly fill a full gb300's 280gb of memory
lmpdev 10 hours ago||||
I assume they mean same t/sec as a SOTA cloud model
csomar 8 hours ago|||
There should be some kind of moratorium on new accounts. HN's always had waves of newcomers, but their impact was always limited. The wave passes and people either get filtered out or adapt. That doesn't seem to be happening anymore, since bots can churn out endless gibberish.

He did answer you though. Native is x10 the non-native speed. 50/50 that's not a bot; though it could be a meat-proxy

flexagoon 8 hours ago||
> There should be some kind of moratorium on new accounts.

OC was registered in 2016 though? What do new accounts have to do with this?

csomar 5 hours ago||
I am talking about my general impression not this particular occurence. I have a suspicion that someone/some entity is buying old accounts to bypass the new accounts penalty. I even created (https://chromewebstore.google.com/detail/hn-users-filter/ine...) to filter these accounts/comments (disclaimer: vibe-coded)
scotty79 10 hours ago|||
I don't know why you'd want to burden your laptop with a large model. But I can totally see a new "developer workstation" product that's just a semi-large box that's optimized for running frontier open weights models for one to few users.
nater5000 4 hours ago|||
I'll give you credit for at least offering a specific, somewhat unique take. But this is a pretty dumb take lol
re-thc 10 hours ago||
> Apple will release M7 MacBook Pros / Mac Minis next year

The latest on Apple is TSMC is stuck on the next iPhone due to lack of RAM. Good luck getting any Macs. Memory shortage is getting worse.

zmmmmm 12 hours ago||
Missing multimodal again?

It is so valuable in practise to be able to have the models see screenshots - I guess if they aren't in the benchmarks then nobody will focus on it. But it completely nixes these for some of my main use cases.

xscott 11 hours ago||
Probably not what you're after, but I've considered having a separate small mm-model act as a seeing-eye dog for the bigger more capable one.
pllbnk 11 hours ago|||
I can’t come up with a use case where I couldn’t extract the image details using another, multimodal model and pass it into the GLM’s context with as many details as I need.
zmmmmm 10 hours ago||
I think you lose a lot by not having the vision capability shared with the text. It is the joint reasoning across them where the power lies (the same model that sees the code and made the changes to produce the visual presentation, sees the image of it and reasons about it).
cmrdporcupine 7 hours ago||
I mean this is assuming the thing you're working on has a UI? Not all of us work in that space.
arcanemachiner 12 hours ago||
I would assume that GLM 6 will be multimodal, but 5.x will be text-only.
lazarus01 3 hours ago||
I’m using deepseek v4 flash to build a complex full stack production ai app and it’s a total beast.

I break out Claude when I hit some serious roadblocks, but that doesn’t seem to be happening much after the last deepseek flash release.

Deepseek prices just went up, but are still low.

I will def try GLM on my next project

KronisLV 11 hours ago||
Their coding plan switched to credits, didn’t it? What are the rate limits like, compared to Anthropic or Kimi K3?

I remember trying their Coding Plan out before the change and the 5 hour limits felt too restrictive then even for light/medium work, especially cause of the whole peak and off-peak thing: https://blog.kronis.dev/blog/z-ai-s-glm-5-2-is-a-great-model...

Nowadays, I’d probably go with their Max plan if the rate limits are okay? Anyone using them now?

Oh also unrelated but ZCode was surprisingly good, which is surprising for a tool that came out of nowhere - even some of the critiques in my blog post have been patched out. Sadly they don’t support using Claude Code as an agent so can’t use it like Paseo or Kepler or Agent Orchestrator.

Havoc 7 hours ago||
> Anyone using them now?

You're gonna have a had time getting straight answer to that out of the internet. There are now 4 different flavours of the Max plan floating around (Legacy V1, Legacy V2, New plans, and the current credit ones). And on top of that they have peak times. So ~8 scenarios, 24 in total across all feedback for their coding plans.

So when someone tells you they're having a good time on a GLM coding plan it's damn near unusable as a datapoint unless both parties are very clear about what precisely is being discussed

[It's been good for me though...V1 Max off peak...which is basically the best of the 24]

ipsod 6 hours ago||
I have V1 Max, and I think they throttled me for using it too much. I was maybe abusing it, by sending out 8 or 16 review agents at a time.

I haven't tried it in a few months, but it went from amazing to unusable really fast.

andai 4 hours ago||
That could also just be random fluctuations in quality of service. Some days it's super fast, some days super slow or I get constant errors.
KronisLV 2 hours ago|||
Update: tested it out myself on their Max plan, on some parallel agentic sessions.

Currently 20% of my 5 hour limit and 4% of my weekly limit.

  Total: 58.46M
  GLM-5.3 Cached: 56.91M
  GLM-5.3 Uncached: 1.23M
  GLM-5.3 Output: 315.18K
  Cache hit rate: 97.9%
Extrapolating from that (inaccurate for now but oh well):

            Full 5-hour  Full weekly
  Total     292.3M       1.461B
  Cached    284.6M       1.423B
  Uncached  6.15M        30.75M
  Output    1.576M       7.88M
All of the work was off-peak I think, using OpenCode not ZCode in these examples.

Their own estimates are quite different, probably due to their conservative caching estimates vs what I normally get on longer form work: https://docs.z.ai/devpack/overview#estimated-token-allowance

ljosifov 9 hours ago|||
Wdym "sadly they don’t support using Claude Code"? For the longest time that's all Zai supported - Claude code. I'd run it via

  export ZAI_ANTHROPIC_BASE_URL="https://api.z.ai/api/anthropic"
  export ZAI_ANTHROPIC_AUTH_TOKEN="$ZAI_API_KEY"
  claude-zai() {
      { local -; set -x; } 2>/dev/null
      ANTHROPIC_BASE_URL="$ZAI_ANTHROPIC_BASE_URL" ANTHROPIC_AUTH_TOKEN="$ZAI_ANTHROPIC_AUTH_TOKEN" claude "$@"
  }
  $ claude-zai
I liked Claude Code to start with. But over time between 'CC cache thrashing undo' seetings (I see now accumulated in ~/.claude/settings.json) and Anthropic-anything becoming a liability - have not used it in while. ZCode is ok and use it to take advantage of the discount tokens on offer from time to time. But really glad to see that in omp (oh-my-pi) Zai is a 1st class provider, can be selected on it's own no configs shananigans needed. And fits in the overall picture. E.g. can select GLM-5.2 (now 5.3) assign role [plan] or glm-5-turbo [advisor].

Got reminded now of glm-5v-turbo - that 'v' was for vision - will try assign it role [vision] now in omp. See what happens. :-) Often times it's handy when describing gui problems if the harness/model 'can see'.

KronisLV 7 hours ago||
I am not talking about GLM models being served through an Anthropic compatible API, that part is perfectly fine and I'm glad they support it!

I am talking about ZCode, the program, being unable to delegate to other harnesses, like using Claude Code (or even OpenCode) within their UI, so that an Anthropic subscription can be used, because Anthropic don't let you use 3rd party harnesses directly.

It's basically what Paseo: https://paseo.sh/ and Kepler https://www.gitkraken.com/kepler and Zed https://zed.dev/ support doing.

ZCode doesn't seem to work at that level, it instead feels comparable to OpenCode or Codex or Claude Code directly, while also being desktop oriented - you just make API calls directly within it.

It's okay if it's not a goal of theirs, it's just that their UI is really really nice and that would be a cool direction for them to also go in some day.

ljosifov 4 hours ago||
Ah sorry - I misunderstood. Thanks for explaining it. Have not heard of Paseo nor Kepler, and have never tried Zed. Yeah I too assumed if I'm to try use OpenAI subscription outside Codex, or Anthropic subscription outside Claude Code - I'd get my account banned it's agains their rules. So I have never looked how using the whole harness from outside looks like either (except for 'claude -p'). Interesting. BTW I see now https://docs.z.ai/devpack/tool/codex Zai added OpenAI compatible end point.
KronisLV 4 hours ago||
My current view on things:

Paseo had a really nice UI/UX, except sometimes sub-agents within OpenCode sessions would hang. Still, quite pleasant if you want something like the Codex or Claude Code desktop apps, but across various providers.

Kepler integrates with issue trackers like GitHub, you can just create a worktree from a ticket and let it churn, seemed like the second most polished option I tried, but there are obvious gaps - like moving cards manually, some missing UI options etc., which I'd chalk up to either the software just being that new or maybe being a little bit vibe-codey. Either way, one of the more promising options if you want something like Kanban board for agents.

Zed is mostly just a (really nice) text editor with some AI integrations, though it seems like they're also building a more agentic product as well - https://delta.dev/ haven't used that one much and am not in circumstances where I'd collaborate with people that closely, but there was a pretty cool podcast episode with the creators recently and it seems like it works pretty nicely for them! As an editor though, it succeeded where Fleet failed and has mostly replaced Visual Studio Code for me. Nothing against VSC, Zed just does most of the stuff I actually need out of the box.

Some of those tools interacting with Claude Code instead of trying to replace it is more or less the way to get Anthropic's models in other tools while still on a subscription (at least for now). How it works under the hood, go figure, there's ACP https://agentcommunicationprotocol.dev/introduction/welcome but also any number of hacky approaches.

To be fair, you can use Anthropic's models in many other harnesses directly, it's just that it then counts against API billing instead of your subscription, which ends up being way more expensive for individuals, but is kinda what you're supposed to do as a company.

cmrdporcupine 7 hours ago|||
I found with GLM I was better off using plans from either Neuralwatt or Ollama.

But Neuralwatt significantly raised their rates since then.

scotty79 10 hours ago||
I feel like quota on their subs is extremely generous. I pay 3-4 times less for larger quota than gpt-5.6-sol.
Gecko4072 13 hours ago||
People familiar with the topic, how will models continue to get better? Post training it seems? Labs have already used up internet-scale data, so are there any limits to architecture improvements and post training or can we expect this trend to continue? ByteDance is training a 10T-parameter model. Here, GLM 5.3 outperforms models 3-4x its size of roughly 700B, so parameter count doesn’t seem to be a direct correlation anymore.
npn 12 hours ago||
> used up internet-scale data

yet but it is still contain a lot of trash. you need better models to process those trash and create a curate dataset. this will happen again and again until there is no more juice to squeeze. and I'm sure we are still not done with it.

> post training

yeah this will be crucial. the big models are already too capable, they are just not that aligned with current agent tasks.

> parameter count doesn’t seem to be a direct correlation anymore

I don't think so, remember that chinese labs do not have as much compute power compare to US frontier labs. that's why deepseek v4 flash had that huge jump and deepseek v4 pro is kinda a disappointment, they just do not have the compute power to proper posttrain the pro model like they wanted. glm is also a relative small model so you also can see the huge jump with just post training. so it does not mean the size does not matter, it is just mean that the chinese labs currently only capable of training smaller models effectively.

alightsoul 11 hours ago|||
GitHub dumps are about 115 terabytes. The common crawl is in the petabyte range uncompressed for every year. Apparently there are dumps of Reddit too in spite of their efforts to ban bots and it's not solely due to the use of residential proxies. For a 1:20 parameter to token ratio, you can still train up to 10 trillion parameters so 10T parameters times 20 is about 200 trillion tokens. Then each token is 4 bytes so 200 times 4 is about 800 terabytes, which is not inconceivable, the common crawl alone has more data than that. So does the internet archive if you donate to them, Anna's archive is 2 petabytes including images, etc etc not all of it is text, but training on multimodal data increases model intelligence by virtue of being multimodal
alightsoul 3 hours ago|||
also reddit has eliminated their api entirely, but dumps of it can still be made. every website can be seen as its DOM with html, css, javascript, which can be seen as source code especially if you only look at its javascript, and its dom with css, html, javascript or only javascript can be added to a source code dump together with github and can be duplicated as plain text with no html markup, no css, no javascript, as an information source. if you pay youtube, instagram, tiktok, bilibili to crawl their data, you can probably get data into the exabyte range.
miohtama 10 hours ago|||
Maybe Reddit dumps explain why Opus 5 is talking like a retarded.
gr_norm 13 hours ago|||
Yeah, the comparison here between GLM 5.3 and Sol + Fable is impressive on its own, but incredibly more so when you consider it's a fraction of the (rumored) size. The miniaturization trend is as strong as ever.
justapassenger 13 hours ago|||
You basically need both. Parameters and good post training. If you keep on growing both, you’ll have good models.

LLMs are still surprisingly “easy”. You need maybe a couple dozens of right people, a lot of good quality data and a lot of GPU that you know how to operate. There’s relatively little “secret sauce” needed.

CuriouslyC 4 hours ago|||
How to structure experiments/scaling and hyperparameter tuning regimes are most of the secret sauce (besides massive compute). If you don't create an experimental ladder to verify scaling and optimize your hyperparameters well, you'll waste a ton of money.

The data is mostly coming from places like Scale/Mercor/etc and net dumps with some filtering and batch prioritization, and RL on verifiable domains like code/math/games.

FergusArgyll 12 hours ago|||
I think there's still a ton of secret sauce needed for serving them economically
justapassenger 11 hours ago||
Sure, same for building a model in an economically sustainable way. But barier to entry is surprisingly low (expect for the huge amount of cash, of course). That’s fairly surprising, given how extremely powerful that tech is.

10 years ago it was super hard to have usable “frontier” ML. You needed very complex data warehouse, feature engineers, feature stores, multi level ranking, calibrations, tons of different model architectures, etc, etc. Each by itself was extremely hard engineering problem and really only handful of companies could deal with that complexity.

With LLMs, 95% of that is gone, infra to support them is greatly simplified. Of course, to make really reliable, performant, user friendly, etc - you still need to a lot of engineering. But it’s very different challenge.

NitpickLawyer 12 hours ago|||
> Labs have already used up internet-scale data

Despite this being the topic du jour of 2025, it was never true. Most of the "we've hit a wall with data" came from communicators / media and not researchers. It got popular because negativity sells. It's a false premise for a number of reasons:

a) Data curation is as important, if not more important than bulk data. Models becoming better at classification leads to better curation leads to cleaner data. Throwing common crawl and pray is so 2023. We've known this since llama3 days, it worked then, there's no reason to think this will not continue to work as the models imrpove.

b) Models are today good enough that you can augment / multiply your data easily with enough compute. You can now have a model take "authoritative content" and create more data from that + scenarios. Say you take a book on computer architecture. You ask models to break it down. Then you ask models to find examples for each topic. Then you ask models to ask questions and offer answers from several viewpoints. Then you take each of those and ask other models to flag inconsistencies. And so on. But you can whateverX your data from one authoritative source + bulk data into 5x - 10x "scenarios".

c) RL is really really really powerful. It's hard to do right (reward hacking, instabilities, etc) but once it works it "keeps" on working. Again, we knew this to be true a few years ago, ever since models really started to do well on math (highly verifiable). It only follows they're getting better on cybersec and other verifiable tasks. But now, with models improving, you get the same data augmentation pipelines as above, just better because they're also verifiable. For example, the way cursor augments their data: take a repo, ask an agent to identify a feature (it can be a large multi-file feature). Remove all code relating to that feature, but keep the original tests in the repo. While training, that becomes a RL scenario: implement this feature in this repo. Verify it with the original (hidden for training) tests. Reward appropriately. Now you can get 1 repo -> 20-50-100 scenarios. Instead of "feed everything into the pretraining", you're now creating scenarios, verify them w/ existing tools, and get your scoring function for the rewards. And, importantly, as the models become better in general, they also become better at this pipeline building exercise. So the next iteration gets trained on more scenarios, better scenarios, and so on.

> how will models continue to get better?

Probably the same. No one can know for sure, but at the moment, despite all the "walls this, slowdown that, plateauing" and so on, there are no signs of slowing down. And, as you noted, this works across the field of model sizes. There are, of course, theoretical information-based limits on size, but smaller models also improve, once "bigger" models can be used as training data generators, oracles for verification, rubric verifiers for open ended questions, and so on.

And smaller models (i.e. cheaper to serve) get to generate more traces during RL, and more rollouts give you better training, and so on. Next up - hardware optimised inferencing (ASICs basically). Once you have that, we can expect another wave of improvements. And so on.

CuriouslyC 4 hours ago|||
Model output is pretty mid at augmenting, it can lead to distribution collapse. It's useful for smaller models because nobody wants manually to curate a specialized corpus and those models can't represent the diversity anyhow, but if the plan for infinite scaling was just to keep feeding the biggest model more of its predecessor's slop, that's not going to work out so well. It might work as a supplement for "thin" areas that have outsize importance for the amount of training data available for them though.

Big models are going to "tap out" on non verifiable fields within ~2 years, just because the pool of experts able to reinforce the models is going to get very small, and as the nuances get finer, the signal from reinforcement is going to get progressively less aligned with the intent. Math and code will be mostly tapped out in that time frame as well, even though we can technically scale them "infinitely," just because the cost benefit won't line up. At that point, most RL will be "gyms" with games that are designed to model designated valuable economic activity.

In the next few years, we'll get small domain specific distillates that are ridiculously smart in their domain (imagine if Qwen 3.X 27B went super saiyan), and even frontier labs will be routing to experts/orchestrating because the cost to serve/TPS difference is huge. They'll still train the god models for PR/marketing, c-suite use and distillation, but using them for day to day work would be like making houseware out of solid gold.

WarmWash 4 hours ago||||
Refreshing to see someone actually understand training rather than treat it like dragging and dropping "internet.zip" into the LLM "knowledge" folder.
Gecko4072 11 hours ago|||
Thank you for your response. Part c was especially insightful. Quite a smart way to do it and makes the possibilities of post training seem almost endless. Makes sense that you just need more time and compute.

A positive feedback loop then. RL->better model->better RL pipeline -> better model…

And we’ve only recently started getting into the much better RL pipelines

CuriouslyC 4 hours ago|||
Small models can be super smart. Big models mostly give you baked in world knowledge, domain flexibility and long context stability/coherence. I wouldn't be surprised if we see Fable level smarts in a coding model that fits in 24GB by next year, but it'll be a savant style coder that needs in context learning, and it'll get very wonky after >100-200k tokens consumed.
andai 4 hours ago|||
Roughly in order: data from simulated environments, data from robotics, data from brain waves.
nullc 5 hours ago||
> Labs have already used up internet-scale data

Not really, but a lot of what isn't used isn't very good.

More important is synthetic data. Use a teacher model with RAG with a huge reference library to write synthetic transcripts of idealized behavior for the model. Use models to judge and correct these transcripts. Train on the good ones. Use bad traces to train the model to correct its own errors (e.g. don't train it to produce a bad transcript but if it finds itself in the middle of one train it to self correct).

Similarly, for tasks that can be closed loop evaluated -- e.g. running computer software and programming, unlimited amounts of novel training data can be generated... including for highly original tasks: e.g. run publications in any domain through a model prompted to look for programming problems suggested by the material. Then write/judge/improve transcripts of solving those novel problems.

I expect in the future smaller models won't be directly trained on any internet data at all-- but entirely on simulations of idealized expected behavior from the model under construction. Raw internet data in that case would show up in prompts, but never in the target output (except of course for prompts that are asking it to copy the input).

anana_ 13 hours ago||
What a week for AI model releases
_ache_ 12 hours ago|
No yet finished! Still waiting for tonight Qwen3.8-27B and the unsloth Q5_K_M/S quantification.

Hopping for an AgentWorld variant from Qwen but I guess, I have too high expectations.

mraza007 13 hours ago||
Such an interesting times we are in,

We just had amazing releases this past two months

kimi k3, glm5.3 qwen3.8 and now glm5.3

These open models are getting really good

w4yai 10 hours ago|
You wrote GLM5.3 two times :)
InsideOutSanta 4 hours ago|||
An LLM so nice, they named it twice.
czottmann 7 hours ago||||
Because it's doubly good.
mraza007 4 hours ago|||
Sorry , It was 5.2 :)
jamesponddotco 3 hours ago|
Is there a plan somewhere that gives access to Kimi K3 and GLM-5.3? I was thinking of testing both to run security reviews of my code.

I know OpenCode Go has both, but their limits seem kinda low, so I'm not sure how feasible it is to run such a task with them.

More comments...