Posted by ThibWeb 6 hours ago
> That model’s usage was well within our budget ($68, about 4kWh of energy use / 365 grams of carbon emissions).
The energy cost is literally 1% of the total cost. For context, 4kWh of energy would drive you about 15 miles in an EV, about half of the average person's daily driving miles. It's boiling 10 gallons of water.
With the talk of AI data centers' impact on the world, you'd think this would be 10x to 100x the amount of energy in order to get the effects they're using here.
My takeaway: the AI data center buildout is an overbuild probably at least as large as the fiber buildout that left us with so much dark fiber. If not even bigger. The only thing that will save the economy is the inability of NVIDIA and chip fabs to produce enough chips to match the buildout planned.
Then why would any AI company exist when they could just use their money to buy tokens from another AI company and make money for zero effort? There would be no incentive to be a provider.
Not to mention inflation would grow to match or outpace the rate you could earn on these guaranteed AI gains.
I have the feeling China is somehow ahead when it comes to energy (and cost) efficiency for AI usage. After all, the two are in a direct competition, and this difference is significant. Or is the "hyper" scaling of energy hungry datacenters in US part of a bubble?
My entire point is that 99% of the dollar cost of running these models goes to things other than the GPU power. The capex cost to building cost to GPU cost to storage/networking/chasses/wiring plus the other operation costs dwarf the electricity. Even the other electricity costs, lets say double it for all the supporting compute, plus another 25% for a 1.25 PUE, and you're at 2.5% of all-in cost of running these models is from electricity.
The non-electricity costs are massive and the constraints on fabs, etc. will drive the amount of the AI build far more than energy availability.
a.) we’re supply constrained
b.) only 3% of households pay for AI
Inference amounts will continue to grow heavily.
There are noise issues in some places. But the panic is excessive. Like nuts.
Maybe environmental panics are to the left what moral panics about stuff like trans people are to the right.
> Nonetheless, it’s a good reminder to be careful with model selection and with agentic patterns. We could have achieved similar results for most likely 5x less cost with not that much more effort. Lessons learned! We need to budget for this, and be more careful. Could have seen it coming, but now we know.
I don't get it. Why was it wrong? Which one would have been better? What was the lesson and how could you have foreseen it?
I am a huge Anthropic fan, USA fan, but are we cooked with AI? Sorry SI, that's the important thing.
GLM 5.3 flash has been good as my default profile Hermes bot, after I readjusted its memory to point to a couple of key skills.
I got some great coding results with GLM 5.2, and 5.3 Flash is supposedly almost as good, so I will be trying it out soon for day to day tasks as the post advises.
It's an excellent workhorse. When I am running out of my GLM quota I switch GLM-5.3-flash to DS-4.1-flash.
He said his experiment was a failure because:
1. He accidentally spent 450M tokens vibe coding with the wrong model, instead of GLM 5.3 Flash.
2. When he used GLM 5.3 Flash, it was sometimes slow. So he switched to other models (Deepseek / Qwen) instead. His guess to why it was slow: GLM 5.3 Flash was so good that the providers were congested.
3. He still needed to use other models besides GLM 5.3 Flash, for R&D and benchmarking.
His takeaways from doing the experiment were:
1. Measure local usage more.
2. Experiment with agent orchestration, with bounded goals.
3. Don't count other models that are used for R&D.
4. Play with Jev.
5. Include experiments with flagship models to compare with cheap open models.
His conclusion about GLM 5.3 Flash: Probably viable for day to day work, but he'll have more thoughts next month.
The second reason appeared to be simply "because we chose not to". The post seems to be pretty much content-less in any practical sense. I clicked on it because I do quite like this models average performance and I was hoping to see some kind of review content.
Makes me appreciate my ChatGPT subscription. I’ve had multiple days between 1B-2B tokens (now less so, models have indeed become token efficient) and regularly in the > 100M range. Even then, $150 sounds excessive. I wonder if their cache is getting nuked for some reason, or maybe they decide to use Cerebras that doesn’t subsidize cached tokens.
It really makes you see how heavily subsidized the subscriptions are.
Edit: Fixed my math. Edit 2: I was looking at the wrong model on OR. Either way, the math is within the correct ballpark.
Edit: I found a linked article that mentions the inference provider who does the measurements.
I'm mainly using flash varients, at least as the default, bump.up to stronger model as needed (less often these days)
It's my favorite model family to interact with, it's prose is the best imo, it makes me laugh from time-to-time (like when it said it would "crib" some code from another project, lul)
I currently have qwen-flash working on an NES emulator harness so qwen-little can play my first RPG (ff1)
(tho I have used all the others I mentioned, happenstance I'm using qwen this iteration/task)
edit: I subscribe to z.ai, I don't host.
What kind of hardware and what particular quant?