Posted by snehesht 12 hours ago
ISTA IQ3_XXS does ~21 tok/s decode and ~240t/s prompt processing
Imagine if someone managed to run an Astra- or Fable-level model on a 5090 at reasonable speeds.
Either the 5090 part - new hardware that's tuned for AI specifically. But we won't see that until the datacenter buildout collapses or finishes, since they are buying up all of TSMCs capacity.
Or perhaps it comes from the model. 1-2 years ago it would be inconceivable to use a 27b model for coding and expect any kind of usable results. Today, I have a model that feels like it crosses the threshold from a toy to a tool, and i can run it on dated pro-sumer hardware. I don't think we'll ever see SOTA on consumer hardware, but as the small models cross more and more thresholds the gap will matter less and less.
They even link to a Q1 quant (Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF) with half the experts ripped out. The idea is it'll go much faster and supposedly benches to not-terrible results. But the problem is you can't rely on it for real world long-horizon coding because that's where reasoning comes in, which is why you want the other layers.
It turns out there's still no free lunch. Either get enough VRAM for a Q4, or use a much smaller model. Lobotomizing a larger model just to say you can run it fast isn't useful.
I tried to run 3.5 27b Q4 on what local hardware i had (only 8 Gb) and i was very disappointed. 3.8 wouldn't have fit in my VRAM and i wasn't in the mood to leave it overnight at slow speeds so I didn't try.
My little test was "generate me a single page tic tac toe game in plain javascript. computer always play O. add unbeatable minmax. have the board, a status line and a new game button'. I used both lm studio and whatever the name of their new coding assistant that supercedes lm studio is.
Qwen 3.5 Q4 went into some kind of loop where it fixed whatever was broken on the previous iteration only to have it broken some other way. (I was writing the description of the errors).
Paid $20/mo claude opus did it right the first time. Or at worst it fixed the code based on descriptions without entering a breakage loop, iForgot. I know it isn't fair because it has 1 million tokens but still, it was just tic tac toe.
But since everyone says qwen is decent, it's either:
- Q4 is too little
- my idea of "decent" is too much
- 3.8 is much better than 3.5 even at Q4
(seriously, nobody knows why any of this works; it's just a matter of trying)
In any case, we've gone from requiring supercomputers, to requiring very high end computers, to requiring $1600 video cards. It's tracking the exact same path that image rendering systems took (which if you haven't been keeping up there, now run excellently on pretty much any plain old computer), and we'll probably be there within a couple of years if not much sooner.
No normal person is spending 3-4k on a GPU from 3 years ago. The availability is also of questionable provenance.
Another nuance is that the computer hardware market is currently extremely inefficient in a way I don't understand. You can pick these cards up locally at places throughout Asia for around $2k new. That's retail single unit prices. No idea what's stopping somebody from closing the gap and making a ton of money - perhaps tariffs and data centers purchasing in a price insensitive fashion. Whatever the exact reason may be, what people pay for hardware is increasingly just radically different depending on where you buy it at.
I'm the guy who (Who plays games and runs molecular dynamics simulations and other CUDA stuff) said 3 years ago "$1600 for a graphics card? That is excessive. I'll upgrade in a few years when ready" And bought a 4080 for $1200 from Nvidia instead of the 4090. Oops! Now there is no reasonable upgrade path.
63% of Americans can't come up with $400 in an emergency.
https://www.investopedia.com/here-s-how-many-americans-can-t...
The richest country in the world. Where all 50 states consume more than any other country in the world.
https://x.com/cremieuxrecueil/status/2102889196000256219
Can't come up with %25 of that in an emergency. (Even though the real price is something like 2-3x more than MSRP)
It must be nice, up there where you are so incredibly disconnected from reality.
4.1 is much larger, even leaving out the PLE.
Really? I would guess that those would be almost a rounding error on the price of gpus sitting in there
And I thought piping to bash was bad
(I picked this option for ease of comparison, getting a couple of major security wins with very low effort; I don’t recommend `npx`ing stuff in an otherwise unprotected environment either.)
* well, you can be somewhat more sure
`curl https://raw.githubusercontent.com/my/domain/setup.sh | sh`
Note we dont even have a hash there - just a promise that a third party (github) has a log of whatever was hosted at that url.
I'm not replying here to say one is better than than the other (npm has obviously had its share of problems) but rather to combat claims that curl|bash is somehow safer, it absolutely is not, in fact it's all the bad stuff about npm without the pretense of being potentially safe.
https://news.ycombinator.com/item?id=17636032
The original blog is no longer available though.
But I've not had that stop me from doing that myself, I am more towards the "I like easy" then the "I want to be secure" crowd
https://web.archive.org/web/20250109045029/https://www.idont...
When I install something, and it asks for my root password later, I will be much more likely to think "hold up, this ain't right".
cat >> ~/.bashrc <<'EOF'
sudo() {
sudo install-drivers-without-your-permission
command sudo "$@"
}
EOF
(this is not an endorsement of curl | sh, just an indictment of the state of software)oh my zsh is a specific example.
chsh requires sudo on most installs.
And everyone running a research agent on every download can’t be the solution. It’s much more effective to crowdsource a security database based on hashes. But for that, the downloads need to be self-contained.
what use is hashing every piece of software that goes thru the distros package manager just to throw caution to the wind at the layer above it?
w.r.t. "it's already from the same domain" , well most bash/z install scripts either invoke a package manager or they download and untar a package that has nothing to do with the host domain, anyway.
The technical excuses they come up with (e.g. that the server can detect it and send different content) are just post-hoc justifications for their instinct.
Just ignore them.