Codex has become my goto tooling. I used to be a Claude Max subscriber, but I was becoming disappointed with the quality of the output from Opus 5. Fable chewed through my usage too quickly to be practical. Moving to a Pro account w/ Codex was a big improvement. Sol had great output, and the usage was more than sufficient for most of my needs. However astra does tend to chew up usage, so when i've done to much of that, and it's became an issue Grok Build has beocme my second go to account. The output especially after the cursor purhcase has become quite good, and the usage has always been very generous.
becquerel 5 hours ago||
Try using astra as an orchestrator for deepseek 4.1 flash, it seems to work out quite well.
sparkling 5 hours ago||
I am using exactly the same flow.
Astra for deep dive investigations, Sol 5.6 at mid-level for day to day tasks, Grok 4.6 via Cursor for routine and low complexity tasks.
6thbit 6 hours ago||
( why is the x-axis on the first chart in descending order ? )
alansaber 4 hours ago||
As anthropic/openai subscription allocations get squeezed you'll see more people using "second rate" closed models like grok. The token allowance with a Cursor subscription is crazy.
sidgtm 7 hours ago||
In my experience Grok especially inside Grok build is pretty solid choice, it’s a no nonsense model and stays on its course. Another surface where I truly enjoy the experience of using Grok model is Grok bot
aschobel 3 hours ago||
Yah, I am pleasantly surprised at Grok Bot. Hopefully this improves CUA which has been a touch lacking w/ Grok 4.6. Grok 4.6 works but is slow compared to stuff like Astra Light.
guywithahat 6 hours ago||
I've had really good experiences with Grok 4.6 and grok build. I've been playing around with tscircuit and it can write code with an understanding of spacial reasoning, while also importing cad components from different file formats into tsx, I've been having claude come in and try to error check it and so far claude hasn't found anything to improve in my three projects.
I'm excited for 4.7 although I share skepticism with other users whether 4.7 will be significantly better, since they didn't raise the price.
simonw 6 hours ago||
$2/million inout and $6/million output but I couldn't see any pricing information for cached input tokens?
sejje 6 hours ago||
cached input tokens are $0.50 per 1M (prompts under 200k tokens) and $1.00 per 1M (200k+)
simonw 5 hours ago||
Do other prices vary for >200,000 or just the cached tokens?
btian 6 hours ago||
$0.40
AM1010101 6 hours ago||
Did 4.6 not have an x-high reasoning level? Why are they comparing 4.7 x-high with 4.6 high?
ssutch3 6 hours ago||
It did not. xhigh is new to grok.
forgot-my-pw 6 hours ago|||
Not sure on the API side, in Cursor you can always use 4.6 at xhigh.
ssutch3 6 hours ago||
We've only used it through API - but you're right, now API supports xhigh for 4.5-4.7.
everfrustrated 5 hours ago||
I think 4.6 got an xhigh after launch. The benchmarks seem to all have been against 4.6 high.
oh_no 4 hours ago||
the AA numbers are generationally bad. double token use (the one thing Grok was good at was low reasoning usage!) to gain 5% in the benchmark score. with reportedly a larger model. maybe it shows gains IRL but wow, I've never seen a new generation model look so underwhelming compared to the last.
sourcecodeplz 4 hours ago||
looks like token efficient/verbosity took a big hit.
Output tokens from Intelligence Index:
- grok 4.6 (xhigh): 97M (for 44 score)
- grok 4.7 (xhigh): 240M (for 46 score)
oh_no 4 hours ago|
which is crazy because this was grok's competitive advantage, worse than OpenAI models but better than everything else, now it's less efficient than Opus or Fable 5.1