Top
Best
New

Posted by itvision 11 hours ago

AMD acquires Taalas to boost inference performance by etching models in silicon(www.theregister.com)
https://ir.amd.com/news-events/press-releases/detail/1296/am...

https://chatjimmy.ai/

538 points | 413 commentspage 3
bhouston 10 hours ago|
Toronto Canada startup btw.
cmrdporcupine 10 hours ago||
Seems to be somehow some kind of offshoot from or connected to Tenstorrent, which is just down the road. Founder looks like he was/is maybe at Tenstorrent and previously associated with Keller?

Always fantasize about applying at Tenstorrent, but wrong side of Toronto. 2 hour commute.

kridsdale1 10 hours ago||
Works well, I remember driving by the ATI building as a kid.
MarkWayneNewton 10 hours ago||
While this design is self-limiting I think its a good approach. It doesn't take an entirely new architecture or infinite memory to produce significant performance improvement.
Legend2440 8 hours ago|
This is a new architecture. It's a non-vonn neumann device.
zkmon 4 hours ago||
I guess the idea is, gains from inference speed could offset the cost of upgrading the chips to a new model when really required. I think general purpose models would consolidate and release frequency might flatten out, favoring this strategy.
num42 4 hours ago||
I have used chatjimmy before, it is incredibly fast, waiting for latest SOTA model on the chips in future. Great!
nojs 9 hours ago||
Can anyone comment on the economics and likely turnaround times of this process, when it’s more mature?

Would it be realistic for a frontier lab to deploy this or would the turnaround time mean the model is always too out of date?

Assuming the weights and architecture are eventually stable, how much cheaper would this end up being?

2001zhaozhao 9 hours ago||
There are always uses for outdated models.

Claude Code is still using haiku 4.5 from ages ago for explore subagents for instance. Not to mention production uses like customer service that only need to be "good enough"

edot 9 hours ago|||
Just looked this up, no longer true. Explore subagents inherit whatever model the parent is. And you can of course make other subagent configs.
samtheprogram 9 hours ago|||
That's solely so that you burn more money. It's totally unnecessary to assume the parent model. Sure, it could be upgraded from Haiku if there was a solid reason to, but...
AussieWog93 9 hours ago|||
I mean, if you could get Opus or even Sonnet 4.5 at 1000+ tok/s exploring the codebase, they would probably change that setting back.

But either way, I think GP's overall sentiment of "delegating intelligence-saturated tasks to an outdated but fast subagent" makes a lot of sense.

alightsoul 9 hours ago|||
Customer service has really degraded huh. 4 years ago they expected opus performance out of human call center agents

I guess losing some customers due to poor customer service is ok if the price of customer service is right.

cogman10 9 hours ago|||
2 to 3 months optimistically assuming everything goes smoothly and is fully automated.

6 months or even a year if something goes wrong in the fabrication process and you need to update things.

If they do more standard asic design, it could be a lot longer as the design needs to be validated on an FPGA cluster, which would necessarily need to be very big for something like a LLM. Easily up to 2 years.

There's a reason chatjimmy isn't demonstrating newer models and why they only show of an 8B model.

shangofox 9 hours ago||
I mean even if it take a few months, it'll still be out of date. But there was a hypothetical when it came up in Feb, would you want Qwen 3.5 at like 10k tokens per second.

At the time people were no doubt saying yes but now 3.8 is out, is that still desirable?

xienze 9 hours ago||
There's soooo much stuff that such a model is still capable of doing in the pursuit of getting a better overall answer. Imagine a powerful research agent that blasts out dozens of the small, cheap models to fetch and summarize one page each. Then the beefy researcher model performs the final analysis.
matheusmoreira 4 hours ago||
> Once the chips are deployed you’re stuck with that model.

At least we can be sure that's the model we wanted. Service providers could be serving modified versions and nobody would ever know.

ratsbane 3 hours ago||
Smart move by AMD. Chatjimmy is very fast and not very good, but I think it might become very fast AND very good.
sgc 7 hours ago||
What does it take to go from here to a model on a pcie card or an m.2 card, so I can plug one into my workstation / laptop? Will 'intelligence' become much like a gpu, where most people just live with the performance of whatever they have installed, outside large companies that must have cutting edge, or prosumers that have a incrementally better version than the masses?

Are we a couple years away, a decade away, or something else?

mdp2021 7 hours ago|
> What does it take to go from here to a model on a pcie card or an m.2 card

It is already that.

> Will "intelligence" become much like a gpu

As an option among the implementations.

> Are we a couple years away

They could mass produce now, but it makes no sense at this rate of improvements in the models.

yunnpp 5 hours ago||
I would've hoped the company stayed independent instead of being engulfed into a behemoth. I'd like to see more diversity in the hardware ecosystem, but I guess the economics of hardware manufacturing aren't there.
syntaxing 10 hours ago|
Honestly, this is starting to make more and more sense. SOTA models are starting to converge to certain architecture and capabilities. I wouldn’t be surprised we end up with a base model ASIC + “fine tune” card where it’s a physical LoRA style adapter.
encyclopedism 10 hours ago||
Imagine a multi-modal model with 1000's of tokens per second. Realtime inference for a host of applications. This is a BIG deal and will change the landscape in unfathomable ways.

The https://chatjimmy.ai demo was impressive.

Once models settle down this makes sense. Imagine a cartridge with a physical model on it. You purchase a cartridge and stick it in your computer/phone/server. Want to upgrade? By a new 'cartridge'.

This should bring inference cost down dramatically, I wonder how OpenAI/Anthropic feel about that.

2001zhaozhao 9 hours ago|||
i'm looking forward to Qwen3.8 27B launch to see how much models have peaked at a given size.

it might already be time to start burning the best small models onto hardware since it's possible they can't get much better at many tasks like knowledge recall due to the inherent information density limits for models at a given size.

Grosvenor 9 hours ago||||
> Imagine a cartridge with a physical model on it.

I can finally have my own Dixie flatline. Cool.

mdp2021 9 hours ago||
> Dixie Flatline

In case some did not know: also the movie (actually TV series) is finally happening.

# Neuromancer - Official Teaser ( https://news.ycombinator.com/item?id=49055037 )

pstuart 2 hours ago||||
The cartridge could be a small mac-mini type unit connected and powered over thunderbolt. If it included like an m5 or m7 with 64GB of memory and a PCIe5/6 4TB Nvme it would be amazeballs. Hopefully when the bubble corrects and hardware advances and prices reset something like that will become available.

Just even comparing compute from 10 years ago (Apple silicon vs Intel) and it's significant. 20 years it gets crazy. My first computer was an 8 bit 6502 with 64K RAM and a 128K floppy drive (I think, it's fuzzy). Everything amazing now will look quaint in due time.

yassa9 2 hours ago||
It is not linear anymore, take in consideration the Moore's law, the curve is nearly saturated now and gains in performance and memroy are not accelerating any more, BUT there is some hope with new different technologies, like the PHOTONIC chips , doing GEMMs through light particles instead of electrons
anthonypasq 9 hours ago|||
very interesting idea. i didnt think of that. i was just assuming youd have an additional one of these in your phone for actual lightning fast local inference
VladVladikoff 10 hours ago|||
Wouldn't this mean someone with sufficient hardware could lift the SOTA model weights off the chip? Or are you saying that these chips would only be used internally by these companies and not sold to the public?
dumberquestions 9 hours ago|||
I wouldn't expect companies not sharing their weights today to be any more likely to share them if they're on hardware, this doesn't sufficiently hide weights from a local user.
snek_case 10 hours ago||||
The weights are very unlikely to be on the chip itself. That wouldn't work for SOTA models that are terabyte scale, even quantized. This is probably an accelerator for specific kernels in the model, but the weights are likely loaded from memory. The chip may have SRAM to store some of the weights temporarily during inference.
foltik 9 hours ago||
At least in the case of Taalas the weights are physically encoded directly on the chip.

It’s composed of 4-bit multiplier cells that compute all 16 possible results in parallel. The top metal wiring layer physically selects the one that corresponds to a multiplication with that cell’s constant weight, and routes it to the next layer.

mdp2021 7 hours ago||
Are you sure? Source? (does not seem to be https://taalas.com/the-path-to-ubiquitous-ai/ , for example)
syntaxing 10 hours ago||||
I don’t get why this is an issue? You can run Claude/OpenAI SOTA models through Amazon bedrock. These weights have to live somewhere to run on Bedrock.
wmf 8 hours ago||
somewhere = an AWS data center with multiple layers of security and NDAs

They won't sell/rent/license the weights to an end user at any price because they don't trust your security.

syntaxing 3 hours ago||
I work in embedded space. Just because it’s in hardware doesn’t mean you can’t “protect” it. Most modern software (regardless if it’s hardware or not) can be cryptophically signed.
bluezly 2 hours ago||
[dead]
amazingamazing 10 hours ago|||
One idea would be to use an open model.
kevin_thibedeau 9 hours ago|||
Then we can have machine psychologists pull cards when they run amok.
all2 8 hours ago||
You have a robot. You need it to be smarter. You buy a new model cartridge (probably a PCIE 9.x). Now you need some domain specific skills. You'd like it to be able to cook, and you'd like it to not dent your walls anymore. You buy 'improved spatial reasoning LORA' card and 'Gordon Ramsey's Chef ULTRA9000' card.

Now your robot can respond sarcastically when you ask for chicken nuggets. Again. It also doesn't dent your walls anymore.

walrus01 9 hours ago|||
Having a base model ASIC as a physical piece of hardware makes me think of the early days of microcomputer desktop stuff where having a socketed ROM or PROM was a key piece of hardware, and people actually knew/cared what ROM was on their system's motherboard.

Imagine if like instead of having a specific Mac Plus ROM, you had a thing that looks like a fat ASIC that can hold models sitting on a slotted daughtercard directly next to the CPU and RAM.

breadislove 9 hours ago|||
we have not converged at all, if you look at how different the chinese models in terms of architecture you can guess that the labs are experimenting a lot as well. we are seeing all different types of hybrid architectures, different attention methods and so on. Of course on a high level its still a transformer but if you take a proper look we are seeing more divergence then a convergence.
smokel 10 hours ago|||
The technical aspects of SOTA models are not publicly documented. How do you know if something is converging?
syntaxing 10 hours ago|||
SOTA American models are not. SOTA Chinese models are. From a physics aspect, closed source models cannot be too far from open source ones in terms of size. There’s only so much you can squeeze out a B100 style cluster even with fancy Dflash style diffusion model for the speculative model.
_aavaa_ 10 hours ago||||
If we had deepseek v4 flash 0731 etched on a chip it would be more than capable enough and fast enough for so many people's needs, even hardcore engineer.
nurumaik 10 hours ago||
Will be capable and fast enough for 2-3 weeks until new sota drops
amazingamazing 10 hours ago||
If it is capable today why would a new model change this?
thombles 9 hours ago|||
I think it’s tongue in cheek. When I first got access to Sonnet 4.5 I remember thinking to myself “y’know if they never got any better and I just had access to this forever then that would be pretty okay”. Turns out my expectations have changed since then and I would like a higher baseline now.
singingtoday 5 hours ago||
Interesting. I've yet to find a model I consider sufficiently intelligent.

Fable is nice, but still requires a lot of guidance for large scope tasks.

catchnear4321 9 hours ago||||
if capability is a commodity then the differentiator becomes taste.
FridgeSeal 9 hours ago|||
Because new stuff instantly makes anything prior bad and incapable and garbage of course! Did you forget the hype-machine speaking notes??? /s
cyanydeez 10 hours ago|||
if they were still exponentially increasing, they wouldn't be preparing for an IPO. IPO is where companies go to die and founders escape.
cyanydeez 10 hours ago||
I don't think there'll be a fine tune card; you'll have the base model vintage whatever year, and then your GPU will do whatever LoRA layers you want it to do; the LoRA will wrangle older dated models into the current of whatever your looking at.

But yeah, for things like programming, if it can do linux and python and some go and sql and javascript, larger domains can be threaded with LORA

More comments...