Top
Best
New

Posted by itvision 16 hours ago

AMD acquires Taalas to boost inference performance by etching models in silicon(www.theregister.com)
https://ir.amd.com/news-events/press-releases/detail/1296/am...

https://chatjimmy.ai/

714 points | 538 commentspage 8
walrus01 15 hours ago|
Imagine the size of chip needed to 'etch' something like Qwen 3.6 27B in size.
golem14 15 hours ago||
Interesting thought, because it's a yield question. How tolerant are models today to a few broken weights.

If tolerant, they could churn out many cheaper chips, some perhaps with slight abnormal tendencies ;)

thepasch 14 hours ago|||
> How tolerant are models today to a few broken weights.

Extremely! You can remove entire layers and the model will still work just fine, with barely perceptible capability losses.

I've cut/bypassed ~15% of total parameters out of Gemma 4 31B on a pod once. Still got perfectly coherent responses out of it. Certain layers are a lot more important than others, particularly early and late ones; but it's honestly astonishing how much can be cut out from the middle without destroying the model's coherence.

I didn't run any meaningful benchmarks, so I have no idea what the capability loss looks like exactly. But "produce coherent and sensible English in response to a wide variety of prompts" was definitely not among the things the model unlearned.

walrus01 14 hours ago||
Brings to mind the scene in '2001' where Bowman is pulling out individual pieces of hardware that represent the mind of HAL, and it becomes increasingly incoherent as more physical hardware is detached.

https://www.youtube.com/watch?v=UwCFY6pmaYY

walrus01 14 hours ago|||
I wonder if you had a few percent of problems in the yield, if it would be functionally equivalent to the difference between a unsloth-published Q6 standard size GGUF vs. the nearly perfect precision of an unsloth Q8-K-XL. Or more like Q4 vs Q8 where a lot is lost.
mdp2021 15 hours ago|||
Not too dissimilar to the first HC1 (6nm 815mm² 53B Transistors embedding an 8b LLM):

> Our second model, still based on Taalas’ first-generation silicon platform (HC1), will be a mid-sized reasoning LLM

flog 15 hours ago||
If someone has that sort of knowledge; how big a chip would be required? Is it possible?
mdp2021 15 hours ago||
Well, given the data above, roughly a 220b transistors chip for the HC1 tech.
cubefox 14 hours ago||
> At 20 billion parameters per chip, you’d need just 50 accelerators to support a trillion-parameter model

I don't see any evidence that this is possible. From my understanding, the whole model needs to be on a single chip. Which rules out any popular frontier models with several trillions of parameters. Even smaller sub-frontier models have hundreds of millions of parameters, so these would be ruled out as well.

wmf 14 hours ago||
The methods for splitting weights across multiple chips are well established. Groq/Cerebras can't hold a model on one chip either.
pyrolistical 11 hours ago||
Umm I have an extra 35, do you have layer 6?
IsTom 14 hours ago|||
I think it's enough that a single layer fits on each chip if you can daisy-chain them with good interconnects.
octoberfranklin 9 hours ago||
They pipeline-parallelize across multiple chips. DeepSeek v4 Pro will be 30 chips.
moralestapia 13 hours ago||
Taalas is just a phenomenal startup from Toronto. My dearest congratulations to the founders.

Edit: Lol, downvotes? Stay jelly, meanwhile Talas goes brrr.

runtime_lens 3 hours ago||
[dead]
khanhnguyen8386 8 hours ago||
[flagged]
gavinbuilds 11 hours ago|
[flagged]