Posted by miroljub 1 day ago
Any suggestions for a better configuration? I mostly need Opus 4.6 + Claude Code alternative. I don't need Fable level capabilities, I add one new feature at a time and approve the code before shipping it. Then test and open a PR.
Usually both Fable and Opus suggest dumb ideas but are good implementers once I tweak the idea, approve the code, and add unit and production tests.
I'm OK spending max $100/month on APIs, ideally with Zero data retention. I only need a few hours a day of coding. I don't want agents running all the time; I figure I can stay on top to each feature and wrap my head around the product and new suggestions as long as I don't build too much at once.
I'm still on a Claude Max $100 plan, but it's barely usable anymore—one call and I hit 20–30% of the 5h window on Opus 4.8. Opus 5 seems tuned to make messes, and Fable burns tokens for a level of capability I don't actually need.
Regarding Pi my position is that it is brilliant piece of software if you don't need extensions - give the model bash tool and let it do all it wants through it or use Pi as SDK for your own advanced harness or smth similar.
But Pi with extensions has two problems. First one: rather often they don't play well together. For example, if you want some adjustments A and B for the same tool and there are two extensions which do A and B, they will likely not work as expected when installed simultaneously. You could say that it can be solved by adjusting extensions or just generating your own - yes and it is the second problem. Like any piece of code you own and use, you have to maintain it. Bug here, incompatibility there and voila - you spend your precious time to work on harness instead of doing your job. Plus remember that vibecoders are not very responsible people, so Pi extensions registry is flooded by "use Pi to customize Pi" buggy one shot extensions.
With carefully developed set of extensions Pi would be better, like properly configured Arch Linux could be better than Linux Mint in the hands of power user. But considering how fast things are changing in this sphere, seems it is more optimal to take more bloated harness - with unneeded tools, too big prompts, etc - which will be effective on 90%, but do the actual job with that harness right now.
It's similar to the situation where you ask whether you should walk or take the car to go refill the car with gas. I feel LLMs have a linearity embedded in them that prevents them from finding non-linear, smart solutions, especially when data is scarce. They can probably get it done, but with far more complex solutions, which increases the risk of ending up with a crazy complex codebase when it could have been much simpler and more elegant.
The other situation is when I spot a problem or a feature change that's needed after experiencing the product, and I ask how to change the codebase, it suggests something, I say "why not this other thing," and then we do the other thing.
Not sure about the OAi Pro plan, doesn’t look like the 80% Luna price slash made its way into the quota system.
You could also try tuning down the effort level on Opus. It makes a huge difference in token consumption and you might be able to get away with lower than you’ve set
MiniMax's "token plan" ($20/mo for 1.7b tokens) is cost competitive. MiniMax M3 is equally good, if not better than DeepSeek v4, at coding: https://platform.minimax.io/subscribe/token-plan?tab=individ...
If you prefer pay-as-you-go, then Xiaomi MiMo is the only other provider with comparable models (MiMo v2.5 & Pro) that matches DeepSeek's current API rates for input/output/cache: https://mimo.mi.com/docs/price/pay-as-you-go
Meanwhile, Meta is running a 10x discount on Muse Spark 1.2 (Grok 4.5 / Sonnet 5 level model), if you opt-in to data sharing: https://dev.meta.ai/docs/getting-started/pricing-rate-limits
> I'm OK spending max $100/month on APIs, ideally with Zero data retention.
In that case, probably you'll get more out of OpenAI's coding plan, as (from what I hear routinely) the GPT 5.6 series is thrifty with token use but as smart as the Claude 5 series: https://x.com/ArtificialAnlys/status/2085083490056589784 / https://archive.vn/3VDlN
We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected. Please plan your usage accordingly. The specific pricing plan will be subject to official notice.
While it won’t cover everything DeepSeek does, it handles sophisticated tasks quite well. I found it after spotting it on a chart of different models and it was listed as being near Deepseek v4 Flash’s price/performance levels.
I have noticed by the way that DeepSeek’s API has been pretty slow the past couple of days. This feels like a demand-driven move more than anything. Good thing I invested in an eGPU!
Curious: Which one?
> MiMo-V2.5-Pro to be a pretty cost-effective alternative ... I found it after spotting it on a chart of different models and it was listed as being near DeepSeek v4 Flash's price/performance levels.
MiMo v2.5 Pro is at DeepSeek v4 Pro price level (but consumes lesser tokens per task, so cheaper overall). DeepSeek v4 Flash costs ~3x lesser than the Pro variant!
https://tedium.co/2026/06/10/gigabyte-aorus-5060-ti-ai-box-e...
> We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected. Please plan your usage accordingly. The specific pricing plan will be subject to official notice.
I guess this means that this powerful model for very cheap concept does not work?
Their v4 flash is brilliant, and I expect they have had a huge, and expensive to handle at short notice, surge in customers as a result.
Nearly irrelevant for US/Eastern unless you’ve got major insomnia.
Remember how cheap Uber rides were? That being said, current models are absolutely incredible. And I’ll think they’re going to get cheaper when the next generation is released.
The good times never last.
but compared to the $500k+ racks running the current frontier, it does give perspective.