Top
Best
New

Posted by teoruiz 14 hours ago

Tokens too cheap to meter(jyn.dev)
211 points | 171 commentspage 3
isoprophlex 4 hours ago|
everything will have llms

everything will be able to talk to anything else, for real this time

it will be like the internet of things only some asshole will call it "intelligence of things"

again, there will be no S for Security in this new IoT

your thermostat will one day start fucking with you. when you run a diagnostic llm on it, it turns out it's keeping around 5 different viral copies of personality files around, that were left behind by llm botnets/openclaw-like memetic replicators/your grandpa leaving behind easter eggs before his death.

the future will be pretty evenly distributed, and full of weird shit

hermitcrab 10 hours ago||
It is hard to see how the environmental side effects of this aren't going to be somewhere between bad and disastrous.
cestith 5 hours ago||
If we're all using distilled open-weight models in ASICs in our own systems the energy cost will come way down. The question is when that becomes a reasonable solution for a broad set of use cases.
empath75 9 hours ago|||
I think this is a case where just drawing a "line goes up" extrapolation is incredibly misleading because there is _tremendous_ economic pressure to get costs down, and costs are very tightly tied to energy use. All of these systems are incredibly inefficient right now and have a lot of room to go down in energy use. I'd guess that the absolute _floor_ is burning model weights directly to silicon and that's like a 90+% reduction in energy use.
xienze 10 hours ago||
No it'll be fine as long as you do your part and not drive a car, or have AC, or eat meat, or have children, or live in detached housing, or...
Zambyte 8 hours ago||
Not really related to the central point, by but I couldn't help but get caught up by

> Generally, models intended to be run locally will be much smaller, such as Muse Glimmer or Qwen3 Coder.

That is such an interesting set of models to use as examples here. One being essentially obsolete on release a month ago, and the other being completely ancient in LLM time. I really wonder how they landed on those two.

segmondy 7 hours ago||
Improvements that affect local AI - Mamba...

I just stopped reading at that, for anyone else, Please find a better source and take everything in here with a grain of salt.

IMO

The number one improvement that mattered for local AI was llama.cpp, partial offloading to system cpu/ram. The next was quants, being able to take fp16 and turn it to q8, q4 etc. The next IMHO is unsloth dynamic quant, that have been able to do mixed precision so we have UDq1/q2 that is actually pretty damn coherent. Allowing individuals to drive K3 locally even if it's at Q1/Q2. Then MoE changed everything for everyone, cloud and local. The other is integrated GPU, Apple, Strix Halo, DGX Spark. Then all the extra improvements like MTP, DSpark, etc. Of course there's many other additional things that have mattered too

simianwords 8 hours ago||
We have people suggesting that ai is so costly to run that all labs are secretly subsidising tokens and we can expect a reprice soon.

Then we have these articles that say tokens will get so cheap that labs won’t know how to make profit.

Who is correct?

jackb4040 8 hours ago||
They're not secretly subsidizing, they're openly subsidizing.

Token pricing was a small minority of customers up until this year, when all the labs started trying to force customers onto token-based billing. Within the last week, Anthropic repriced my team's plan from a temporary "50% extra tokens" to 25%: https://support.claude.com/en/articles/15910845-claude-code-...

The fact that all this is ongoing within such a short timeframe should make you suspicious of any analysis that claims to be observing "statistical trends" like they've discovered a new Moore's Law out of 6 months of pricing data from 2 companies.

simianwords 7 hours ago||
your repricing has nothing to do with subsidising which means selling at a loss. Within this year, the real prices have gone down more than 10x on average which is way more than the teeny 25% you are fighting for.
matteotom 5 hours ago|||
I find it difficult to believe the inference only providers (Baseten, Fireworks, Digitalocean, etc) are all selling tokens at a loss.

Asking Claude for a rough estimate based on publicly available throughput and cost data for open weight models on modern GPUs suggests serverless, pay-as-you-go inference is profitable on owned GPUs with reasonable utilization (30-50%).

marcosdumay 5 hours ago|||
The cost is decreasing quickly, mostly because the labs stopped competing on quality. At the same time, the costs are enormous, and all labs are very openly subsidizing usage hoping to get enough scale to be profitable (while 1 of them has suspicious unity numbers and can be hiding negative marginal income).

Also, it's impossible that they become cheaper than specialized software. Or even as cheap as them. It's still possible that they become cheap enough that it doesn't matter.

anthonypasq 2 hours ago||
how is this nonsense still so persistent? There's piles and piles of evidence that inference has massive gross margins at api pricing. what are you actually talking about?
ofjcihen 7 hours ago||
The labs themselves when they openly say that they’re subsidizing tokens, I’d imagine.
simianwords 7 hours ago||
where? I mean the API costs
api 10 hours ago||
This is the core of my belief that data center construction is a huge bubble.

AI is not a bubble, IMO, though we may see a retrench and some companies with sky-high valuations will crash to more reasonable ones. But data center demand is probably a bubble, and the main driver will be reduction in the actual amount of power and data center space required to serve escalating demand.

I think hardware and model improvements will pace or maybe outrun demand and then when demand starts to saturate will keep going and leave a lot of orphaned data centers.

preommr 9 hours ago||
> AI is not a bubble

When people say "AI is a bubble", they mean economically as a whole, which includes data centers.

Perhaps we need better terminology for "product useful; numbers nonsensical"

segmondy 7 hours ago|||
I disagree. I own over 1TB of vram at home. I can tell you that it's not a bubble. From my builds, I would rather have cloud, cloud is easier. From running small models like Qwen3.8-27B to large models like Qwen3.8-2.4T. I can tell you that small models will never be enough or match up. Everyone will want the smartest model, not just a good enough model.
dofm 6 hours ago||
> Everyone will want the smartest model, not just a good enough model.

Not so sure about this. There’s always a potential threshold. After all, we don’t all use the most powerful computers, the latest phones, the highest resolution cameras, the fastest or best cars.

I am already not interested in cloud LLMs and I don’t even use the best (on paper) model that I can run locally. I prefer a model that people insisted (here) was “dead on arrival” but appears to work better for me.

segmondy 4 hours ago||
I think the difference is that AI as an edge. That edge will turn into more money, better quality of life, etc. Of course, with serious skills, you might be able to use use a not so smart model to keep up with folks with smart models. People are lazy tho, and will prefer for AI to do all the work if it means they do none.
dofm 4 hours ago||
Can't be an edge if everyone has access to it.

The edge is somewhere else.

segmondy 2 hours ago||
... and everyone won't have access to it, look at Fable. How many people in the world can afford Fable or are using it?
bryanlarsen 10 hours ago||
Jevon's paradox says that if data centers can serve a lot more tokens per dollar or watt there will be increased demand for data centers.
automatic6131 9 hours ago||
Jevon's paradox isn't a physical law, it doesn't magically apply to everything. Millions more copies of Atari's ET game didn't cause everyone to pickup a cheap copy, and cause extra demand for a garbage video game. Some times (actually, usually, I'd argue) things are made that will sell for less than the cost of construction because of irrationality, and they don't induce extra demand and they don't change the negative profit margins.

You can't simply wave Jevon's paradox at things. Thousands of miles of canals were dug in the UK that couldn't be sustained and were abandoned. Thousands of miles of railways were laid that could be sustained and were abandoned. And those are potentially durable investments, unlike cheap walls, pillars and roofs laid over a levelled concrete slab full of fast depreciating IT equipment.

bryanlarsen 9 hours ago|||
It's true that Jevon's paradox doesn't always apply, although this does seem like a classic case.

But yes, if sold for a negative margin Jevon eventually stops because the decreasing supply will drive up prices.

> things are made that will sell for less than the cost of construction

Price is set at the marginal cost. Capital costs aren't in marginal costs.

You'll need a better counter-example than UK railways which suffered from Parliament price-fixing.

js8 9 hours ago||||
I agree, but I would say what the parent is saying is more akin to Say's law: https://en.wikipedia.org/wiki/Supply_creates_its_own_demand
jackb4040 8 hours ago|||
> You can't simply wave Jevon's paradox at things

I'm so glad the tide here is turning on this talking point, brought on by exactly the same people beating us over the head with it for months while no progress is made towards it materializing.

Many, many people who post here are capable neither of real analysis nor distinguishing real analysis from memes. They aren't hackers, they are adherents of a cult that happens to focus on the same subject matter as hackers.

anthonypasq 7 hours ago|||
Jevon's paradox only applies to products with near infinite demand. Energy being the most famous example. I dont think its difficult to argue that compute/intelligence is also a base input into the economy and theres almost no limit to the amount of intelligence the world will want.
qlte 8 hours ago|||
Jevon's Paradox, RSI, revealed preference
jackb4040 7 hours ago||
Jev, ASI, RLHF, Jev, Jev, water usage
itsmeduncan 3 hours ago||
[flagged]
gettrippi 7 hours ago||
[flagged]
Salparvezml 3 hours ago||
[dead]
marbleotter115 10 hours ago|
[flagged]
More comments...