Top
Best
New

Posted by itvision 14 hours ago

AMD acquires Taalas to boost inference performance by etching models in silicon(www.theregister.com)
https://ir.amd.com/news-events/press-releases/detail/1296/am...

https://chatjimmy.ai/

650 points | 498 commentspage 6
christkv 3 hours ago|
There is a big risk in etching a model into silicon like this. We are still evolving what small models look like and improving their performance. When do you decide to etch one into silicon knowing that right now an improved one can be 3 months away.
yousif_123123 10 hours ago||
If things like this get traction, will we need all the datacenters?
downrightmike 8 hours ago|
You are mistaken about what the datacenters are for
tripledry 4 hours ago||
What are they for?
jauntywundrkind 11 hours ago||
Core rope memory is back baby!

Enjoying the Ian Cutress / TechTechPotato video on Taalas. Some ok good technical details on the tech, and some good insider baseball, whose who stuff. (What a treasure having tech discussions like this about.) https://youtu.be/3MKRjt59hh4

OddMerlin 9 hours ago||
Congrats to the Taalas gang.
tonyhart7 3 hours ago||
so in the future I can buy KIMI, GLM or whatever model that get "soldered" directly into GPU ????

so instead of RTX xx70 series, I can buy xxTA that have kimi integrated ??? is that right ??

bob1029 13 hours ago||
I feel like NAND process tech could become useful at solving some of these problems. A GPU where you can update the weights a few thousand times may be sufficient.
mdp2021 12 hours ago||
The basis of Taalas is "compute in memory" electronics - past Von Neumann's separation of processor and memory.

You need to be able to add|mul where the data (the weights) are stored.

addaon 12 hours ago|||
NAND hasn't been scaling great lately. It seems like PCM or MRAM would both be better fits.
kridsdale1 13 hours ago||
FPGA model storage?
concraper 9 hours ago||
A massive L for Canada
tech234a 8 hours ago||
See also: Twitter statement from Taalas https://x.com/taalas_inc/status/2085458427757937097
fellowniusmonk 12 hours ago||
Token quantity will have a quality all its own.
ycui7 13 hours ago|
so qwen3.x-27b on hardware? or better deepseek-v4-flash on hardware .
ilaksh 13 hours ago|
I wrote them an email asking for PrismML Bonsai 27b Ternary which is like 6b or something crazy small and would be a lot easier for them to do initially.
mdp2021 12 hours ago||
They were specializing their forthcoming system on 4-bit FP - which I understand is a structural decision.

Bonsai Ternary (1.7bits/weight) is a compromise, compromise that has to make sense in the context - efficient when translated into transistors.

More comments...