Top
Best
New

Posted by itvision 13 hours ago

AMD acquires Taalas to boost inference performance by etching models in silicon(www.theregister.com)
https://ir.amd.com/news-events/press-releases/detail/1296/am...

https://chatjimmy.ai/

630 points | 478 commentspage 5
yigalirani 3 hours ago|
what prevents amd to just do what they do without acquiring them?
roschdal 2 hours ago||
Is this the singularity?
christkv 2 hours ago||
There is a big risk in etching a model into silicon like this. We are still evolving what small models look like and improving their performance. When do you decide to etch one into silicon knowing that right now an improved one can be 3 months away.
jackdoe 10 hours ago||
Can you imagine in few years getting Fable level intelligence at 20k tokens per second?

"You are not prepared" --Illidan Stormrage

drob518 10 hours ago||
So, Kimi K3 in silicon sometime soon?
roughly 10 hours ago||
How's that jive with the fact that they're introducing a new model every other week?
drchickensalad 9 hours ago||
The new model every week is not necessary at this point really. What if you could run opus 5 for the next couple years at 1/20 the cost?
roughly 9 hours ago||
What's interesting about this is that I as a user would find this useful, but I think the AI industry as a whole would find it an absolute goddamn disaster. Opus 5 is a very good tool, but it is not a human-replacement-level intelligence, which means the entire revenue stream the industry's built on - labor replacement - is not met by this, and the only slightly charitable read of the industry's finances is that they're gonna bootstrap their way to creating the labor replacement hypothesis by getting people to spend money on Opus/etc, whereas if the actual product is a 1/20th the cost Opus-on-a-chip, the entire business and financing model that's tying up $N Trillion dollars of investment money goes out the window.

Great for us, looks like a recession as far as the Market is concerned.

jaggederest 9 hours ago||
Pipeline the burn into silicon, lower the latency as much as you can, for the 10-100x operation cost it's worth it. Imagine if frontier models cost $5/mtok and the 2nd or 3rd tier models cost $5/billion tokens for 3-month-old models.
tecoholic 11 hours ago||
With web search and tool call a decent current generation model at the speed of the chatjimmy could do a lot. People saying it would be out of date are missing the point. It’s not going to make much sense for frontier companies that’s chasing the SOTA. But for a lot of business use cases if someone can put GLM 5.2 and sell it as a box, it would make so much sense.

My partner has been asking for a “completely private” model for doing research and shifting through volumes of data that can’t leave the office and $$$ for the current hardware makes no sense. It would be an easy sell if someone walks in with a black box that contains “ChatGPT”.

5555watch 10 hours ago||
In my understanding the first Deep Think / Pro models were already very good as they were doing some kind of parallel repeated reasoning, thus were slow and expensive. So if chatjimmy speeds enables a fast deep think level performance, I think that would be great.
equinumerous 11 hours ago|||
100% agree - you don't need the most up-to-date model to have something that's useful in agentic contexts. They could even produce chips with weights that make all the decision making/logical reasoning and have it delegate to other specialized agents. If it becomes cheap enough to print a run of custom chips, releasing a batch for each major advancement does not seem unreasonable for SOTA companies.
cephei 11 hours ago||
There are so many use cases for supremely fast offline models. The first thing that comes to my mind is for real-time video processing or other non-textual content in real time.
anigbrowl 8 hours ago||
I wouldn't call it supremely fast but zippy and versatile, yes: https://shop.m5stack.com/products/ai-pyramid-computing-box-p...
tonyhart7 2 hours ago||
so in the future I can buy KIMI, GLM or whatever model that get "soldered" directly into GPU ????

so instead of RTX xx70 series, I can buy xxTA that have kimi integrated ??? is that right ??

yousif_123123 9 hours ago||
If things like this get traction, will we need all the datacenters?
downrightmike 7 hours ago|
You are mistaken about what the datacenters are for
tripledry 3 hours ago||
What are they for?
jauntywundrkind 10 hours ago|
Core rope memory is back baby!

Enjoying the Ian Cutress / TechTechPotato video on Taalas. Some ok good technical details on the tech, and some good insider baseball, whose who stuff. (What a treasure having tech discussions like this about.) https://youtu.be/3MKRjt59hh4

More comments...