Posted by altertable 11 hours ago
The next time I hear about them I am laughing, because when I could enjoy these powers? How many years I should be sitting in a waitlist...
This model had its knowledge replaced with reasoning ability. The chain of thought what makes this reasoning effective.
So this is why you need to let it think and don’t quantize the kv cache.
Yesterday I did have success with Gemma-4-12b with 128k context. It fits in my RAM and it's relatively fast on my hardware.
I had to give it prompts that are quite a bit different from the way I use foundation models, but I did get it to work quite well. I feel like I could learn it's differences and get good at using it for real work.
(update: I got my answer. support@ replied and said my email domain is on their blacklist. It was just me (and I've resolved it)).
EDIT: Or, maybe it's just token pricing, but $10 is the minimum? Maybe it's that.
There is a separate subscription based plan, which is sold out now.
The CEO was on Gradient Dissent a couple years ago: https://www.youtube.com/watch?v=qNXebAQ6igs
[1] https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultraf...
It would look bad for cerebras if other people are hosting the 27b version and show a higher TPS than cerebras.
Sure, 4x input , but cheaper output. Though Cerebras doesn't have prompt caching, so not great for agentic workloads. (they do, but it doesn't affect the price.
Most of the cost for agentic coding is input tokens, you pay for the whole context at each tool call or message. Output tokens is just a small rate