Top
Best
New

Posted by altertable 13 hours ago

Qwen 3.8 27B available on Cerebras at 1500 tokens/s(inference-docs.cerebras.ai)
520 points | 167 commentspage 4
drchaim 12 hours ago|
The idea of custom software on the fly is coming
Marciplan 13 hours ago||
used their Code product with GLM4.7. its fun but if the model is bad it just doesn’t do much useful.

Hope they add such models to Code too :)

altertable 13 hours ago|
Yeah GLM 4.7 is from another decade at the speed we're going
trvz 13 hours ago||
Normal people: tok/s or t/s

Psychopaths: tok/SEC

scotty79 13 hours ago||
I like tps
verdverm 12 hours ago||
do you get reports on them?
altertable 13 hours ago||
ok fair, caps lock kept ON /o\
jing09928 6 hours ago||
[flagged]
byako 13 hours ago|
[flagged]
miohtama 12 hours ago||
Your brain can wash laundry and cook pasta, so there is still a long way to go
qiine 12 hours ago|||
(requires additional fleshy bits sold separately)
davrosthedalek 11 hours ago|||
regarding my brain, my mother might disagree on the laundry part.
dgellow 12 hours ago|||
Your brain updates itself constantly and maintains your whole body, LLMs are static.

Still, 1500tokens/s is indeed wild

eli 12 hours ago|||
If you read the reasoning trace for Qwen 3.8, it does a whole lot of "uh" and "But, wait..." too
howunfortunate 12 hours ago|||
You're absolutely right - filler words are genuinely load-bearing
Zambyte 12 hours ago|||
At 1500 tps, "uh" is about 0.7 ms, instead of 200-300 ms for a human.
ripbozo 12 hours ago||
fyi this is an AI bot account