Posted by jonotime 2 days ago
Everyone does optimization of model serving because it's good for every player in there.
(Also the water consumption thing is not a real issue.)
I have deployed multiple setups with 2/4/8 x H100/200 to do data entry with LLMs at big companies. Trillions of tokens already inferenced ok those. The starting price is about 100k.
For coding, I rather spend 10x more than have even 1 bug but I'm only spending 2x 3x more if you count subscription cost.
This is a killer use case for something like customer support though.
"I have not used anything else but DeepSeek is definitely better than anything else."
Ok? How are you judging that? Am I missing something?
I switched from DeepSeek 4.1 flash about 2 weeks ago for my Hermes sysadmin/coding agents and I am seeing better intelligence and lower overall spend.