Top
Best
New

Posted by jacquesm 12 hours ago

On A.I. regulation and messaging(twitter.com)
https://xcancel.com/DarioAmodei/status/2088758816376807762
144 points | 270 commentspage 3
toasty228 4 hours ago|
The tech industry cooking up some new way to screw people over for 25 years

The tech execs waking up on a random monday:

> I think that ordinary people don’t trust companies, governments, or the tech industry and always suspect that we are cooking up some new way to screw them over.

No shit Dario, no shit...

Havoc 4 hours ago||
Yeah and we’re not even done with dealing with the societal consequences of social media and dopamine algorithm driven everything
knollimar 2 hours ago||
That sounds like a "dismissal attack" and we're giving you less now to combat it
saidnooneever 3 hours ago||
he notes how to critisze frontier labs for failing to deliver, but says also not to blame their marketing. That is next level dim, because the marketing sets expectations for what they will deliver.

this guy need to stop vibeposting nonsense.

onion2k 6 hours ago||
When people talk about Qwen 3.8 being on a par with Fable, they're really talking about Qwen 3.8 Max aka Qwen3.8-2.4T-A95B. That's a 2.4 trillion parameter Mixture of Experts model with 95B active parameters. You need about 400GB of RAM to run it. No one is running that locally.

The distillations of Qwen 3.8 down to a 27B model are good, but they're not on a par with frontier models.

When Dario talks about open weights not being a solution this is what he means - if you don't have 400GB of VRAM lying around the fact that there's an open model like Qwen3.8-2.4T-A95B doesn't really help much. If we're not regulating how models are available, or making sure access is open, then RAM prices will mean everything concentrates on a few very rich companies.

xscott 4 hours ago||
I don't really understand the argument you're making, but just to add a data point:

DeepSeek V4 Flash 0731 is 167 gigabytes from the developer and as a GGUF with no additional quantization. It limps along on my 192GB M2 Mac from several years ago [0]. This model tests better[1] than Claude Opus 4.6 released in February. That's six months ago - what will be available 6 months from now?

https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/tr...

https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF

So yeah, enthusiasts aren't going to run frontier models on their gaming machines, but a small office could easily justify the $30k - $100k cost to run something like this at high speed. The small company I worked for routinely spent that kind of money on Dec Alphas twenty five years ago, and that's not accounting for inflation adjustment.

And this is completely discounting the advances smaller models are making. You're right that Qwen 3.8 comes in different sizes. However, Qwen 3.8 27B and Qwen 3.6 27B do run on gaming cards, and they're better than the frontier models from twelve months ago.

I have no idea what will happen in the future, but I wouldn't base my guesses solely on the largest open weight models.

[0] Yes, it's unpleasantly slow (5-8 tok/sec)

[1] Yes, benchmarks should be taken with a lot of salt.

woadwarrior01 5 hours ago|||
> The distillations of Qwen 3.8 down to a 27B model are good, but they're not on a par with frontier models.

Does it have to be? There are plenty of coding tasks, where it's good enough.

onion2k 5 hours ago|||
Practically, no, the distill is great. It's fine to use it.

However, if you're having a discussion about access to frontier models, and using Qwen 3.8 as an example of how open weights is a solution, then you should be honest and accurate about what you're talking about. Making an argument like "People can run Qwen 3.8 at home. That shows open weights are great." is a bit disingenuous if you're not also making it clear that you're not talking about Qwen 3.8 Max (or that you have a beast of a PC at home :D ).

woadwarrior01 5 hours ago||
I agree 100%. IMO, it all started with ollama misrepresenting the Deepseek R1 distills as Deepseek R1, all for hype and marketing. I've had so many ostensibly technical people telling me: "I tried DeepSeek R1 and it was terrible", and every time when I probe further they'd tried the tiny 1.5B Qwen2.5 distill model that was further brain damaged by ollama's naive RTN quantization[1]. DeepSeek themselves were very forthright about it by naming it DeepSeek-R1-Distill-Qwen-1.5B[2].

I suppose the road to technical hell is paved with marketers and grifters. :)

[1]: https://ollama.com/library/deepseek-r1:1.5b [2]: https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-...

Pavilion2095 5 hours ago|||
> coding tasks

Exactly. There are common coding tasks that these models can adequately do. They are absolute trash for anything that isn't coding. And even with coding, they are so, so far behind frontier models.

dgellow 5 hours ago||
No one is running that locally because of the AI bubble consuming all the hardware in the industry. That won’t be the case long term though
mcntsh 4 hours ago||
So the solution to convince the public that these companies aren’t “looking for new ways to screw them over” is to try to go into biomedical research.

I guess it’ll be great for Anthropic to have the cure for cancer, but what’s that gonna mean for people with cancer? Funny how he doesn’t talk about that part.

fabsalvadori 3 hours ago||
I see little conversation about AI powering robots' core and the fact that we are going to have like a billion robots by the end of this decade. Yes, sorry if trust is low.
dude250711 6 hours ago||
So, what does a for-profit CEO, beholden to profit-seeking investors, have to say?
BrenBarn 5 hours ago||
> I think that ordinary people don’t trust companies, governments, or the tech industry and always suspect that we are cooking up some new way to screw them over.

It is so tragicomic to see people whose lives are built around companies coming so close to realizing that everything they do is bad, and then at the last minute swerving aside to convince themselves that no, if they just do more of it and somehow do it "better", then it will all be okay. The reason people don't trust companies and believe they are cooking up some new way to screw them over is because that is what they are doing. If Dario or anyone else really wanted to dispel that perception there's an easy way: do a total 180 and start fighting against everything you've been pushing. But none of them will do that because they still fundamentally believe that what they are doing is good, and are unable to see that fundamentally it's bad.

3dsnano 2 hours ago||
i’m all for ambition and tenacity, but this is psychopath level communication… blissfully unaware, simply incapable of reading the room

if claude still can’t yet write and refactor coherent code (on its own, replacing software engs, just like he said last year), why would we expect it to… fucking cure cancer?

cindyllm 1 hour ago|
[dead]
andrewstuart 2 hours ago||
Is he worried?

Dario is always deeply worried.

I read somewhere he is the Woody Allen of AI.

doubtfuluser 3 hours ago|
> „… it will actually be possible to cure most human disease in ~5-10 years, as crazy as it may sound to ordinary people …“

Nothing beats the hubris of the Silicon Valley CEOs and self proclaimed “elite”. I’m missing people calling out “we are better than that”… NOT

More comments...