Top
Best
New

Posted by itvision 15 hours ago

AMD acquires Taalas to boost inference performance by etching models in silicon(www.theregister.com)
https://ir.amd.com/news-events/press-releases/detail/1296/am...

https://chatjimmy.ai/

681 points | 513 commentspage 7
tech234a 9 hours ago|
See also: Twitter statement from Taalas https://x.com/taalas_inc/status/2085458427757937097
tonyhart7 4 hours ago||
so in the future I can buy KIMI, GLM or whatever model that get "soldered" directly into GPU ????

so instead of RTX xx70 series, I can buy xxTA that have kimi integrated ??? is that right ??

fellowniusmonk 14 hours ago||
Token quantity will have a quality all its own.
ycui7 14 hours ago||
so qwen3.x-27b on hardware? or better deepseek-v4-flash on hardware .
ilaksh 14 hours ago|
I wrote them an email asking for PrismML Bonsai 27b Ternary which is like 6b or something crazy small and would be a lot easier for them to do initially.
mdp2021 13 hours ago||
They were specializing their forthcoming system on 4-bit FP - which I understand is a structural decision.

Bonsai Ternary (1.7bits/weight) is a compromise, compromise that has to make sense in the context - efficient when translated into transistors.

galaxyLogic 10 hours ago||
"... the chip serve Meta’s Llama 3.1 8B at a blistering 16,960 tokens a second — when announced last February, that was 48x faster than Nvidia's GPUs and 8.5x faster than Cerebras' accelerators. "
andrewvl 13 hours ago||
It must be a “super model”. What will be if new model released? New chips?
downrightmike 10 hours ago|
Chip pops out like a gameboy cartridge. AI not working? Blow on it and jam it back in
api 12 hours ago||
I've had an endgame idea in mind for a while.

Models, probably first open weight ones like Kimi K3 class, are etched into silicon like this and sold as cartridges almost like old school game cartridges.

You buy a USB-C dongle that the cartridge goes into, or for data centers you have PCI cards that take these in slots.

wmf 12 hours ago|
Each cartridge costs $1,000. Do you still want it?
anigbrowl 10 hours ago|||
For fast Kimi K3? You're damn right I do
wmf 9 hours ago||
$1,000 only gets you the Qwen 27B cartridge. For Kimi K3 it would be more like $100,000 (and the "cartridge" is the size of a refrigerator).
trollbridge 4 hours ago||
I would gladly pay $100,000 for local K3 running at 18,000 tok/sec.
singingtoday 9 hours ago||||
Yeah. I have 3 max20 plans.
api 12 hours ago|||
Me? Probably not. A business or a hoster, sure. There'd probably end up being an aftermarket in used cartridges with slightly older but still good models on them.
ur-whale 12 hours ago||
Yeah, so https://chatjimmy.ai/ ... the model is crap, but the speed is amazing. Worth checking out.
empiricus 1 hour ago|
Worth wondering why they used a crap model.
jijji 9 hours ago||
taalas is great for llama 3.x 8B models, really bad for one board serving Kimi K3, it seems like you would bottleneck at a few hundred tokens no matter what you do.... spreading the big model against multiple cards seems the only way to get into the 1k+ tok/sec range. Another thing taalas is doing is masking the model weights into the silicon itself, not a flashable firmware, which would increase latency....
walrus01 14 hours ago|
Imagine the size of chip needed to 'etch' something like Qwen 3.6 27B in size.
golem14 13 hours ago||
Interesting thought, because it's a yield question. How tolerant are models today to a few broken weights.

If tolerant, they could churn out many cheaper chips, some perhaps with slight abnormal tendencies ;)

thepasch 13 hours ago|||
> How tolerant are models today to a few broken weights.

Extremely! You can remove entire layers and the model will still work just fine, with barely perceptible capability losses.

I've cut/bypassed ~15% of total parameters out of Gemma 4 31B on a pod once. Still got perfectly coherent responses out of it. Certain layers are a lot more important than others, particularly early and late ones; but it's honestly astonishing how much can be cut out from the middle without destroying the model's coherence.

I didn't run any meaningful benchmarks, so I have no idea what the capability loss looks like exactly. But "produce coherent and sensible English in response to a wide variety of prompts" was definitely not among the things the model unlearned.

walrus01 13 hours ago||
Brings to mind the scene in '2001' where Bowman is pulling out individual pieces of hardware that represent the mind of HAL, and it becomes increasingly incoherent as more physical hardware is detached.

https://www.youtube.com/watch?v=UwCFY6pmaYY

walrus01 13 hours ago|||
I wonder if you had a few percent of problems in the yield, if it would be functionally equivalent to the difference between a unsloth-published Q6 standard size GGUF vs. the nearly perfect precision of an unsloth Q8-K-XL. Or more like Q4 vs Q8 where a lot is lost.
mdp2021 13 hours ago|||
Not too dissimilar to the first HC1 (6nm 815mm² 53B Transistors embedding an 8b LLM):

> Our second model, still based on Taalas’ first-generation silicon platform (HC1), will be a mid-sized reasoning LLM

flog 14 hours ago||
If someone has that sort of knowledge; how big a chip would be required? Is it possible?
mdp2021 13 hours ago||
Well, given the data above, roughly a 220b transistors chip for the HC1 tech.
More comments...