Posted by jonotime 2 days ago
I'm well aware that there's nearly infinite opportunities to yak shave "perfect" OpenRouter setups and some people appear to enjoy bouncing from IDE to IDE as though change costs aren't a thing, but I discovered that I genuinely like Cursor and at least right now it's insanely subsidized by Auto clearly defaulting to whatever Grok's most powerful model is.
I dropped my $200/month subscription to $20/month and stick to Auto for all but really important Plan tasks, and I have basically zero chance of using up my monthly credits even using it 6-10 hours some days.
You make Cursor sound like one thousand times more important than it is. It's a product in deep water.
I actually do use an agent harness to organize files on my desktop. They make a great fuzzy file renamer. Point it at a directory of disorganized files with names all over the place, give the directory layout and file name pattern you want it to have and it makes it happen.
But ever since I've switched to one of the $100-tier subs, I can see why a lot of the people on it don't really discuss the open models often. I'd still use it especially when it comes to sensitive inputs, but for most work, what you get on OpenAI or Anthropic is really more than enough.
It really got even better when they also made their cheaper models up to par if not better than the open models.
I do think the crowd for open models are out there, especially when you see trillions of tokens running for them on OpenCode or OpenRouter leaderboards.
But I run it locally. When I tried it on open router when my gpus were busy I must have gotten routed to some crappy providers, because it was pretty bad.
For me Glm and DeepSeek are nowhere near this Qwen model. I tried various harnesses including omp which I heard supposedly "makes DeepSeek 20 points better". The difference was in the noise (1 point). I run a bunch of benchmarks Terminal World 40, terminal bench 2.1,SWE Pro, GSO. Before those 3 there was no open model that scored more than 1 point on my subset of GSO. Glm scored 5, DeepSeek 3, but Qwen did 17 and opus 19.
Qwen is a small model so it fails on factual recall. But if you give it most of the info it needs it us amazing.
The token-equivalent monthly spend is > $5K+. If Deepseek's token cost is 20x cheaper, that's $250/mo, and I'd be spending a lot more of my brainpower babysitting it and getting worse results.
For business/team accounts that pay per-token, maybe I can see the "freaking out" being warranted on the part of the fronter labs. But as long as they're willing to subsidize their end-user subscriptions, I'm not going to move off of them until the alternatives are truly at their level.
So the industry is responding, where it matters. Which is on heavy API usage, not coding subs.