Top
Best
New

Posted by meetpateltech 9 hours ago

Grok 4.7(x.ai)
480 points | 395 commentspage 5
thih9 7 hours ago|
I refuse to use Grok. Mostly because of the usual reasons - somehow this high profile AI model seems more disgusting than others and it is in a way impressive.

But also Xai doesn’t seem to care about user experience and long term support.

eknkc 7 hours ago||
I am subscribed to ChatGPT, Claude, Kimi and GLM coding plans. 200$ one on GPT and the 20$ ish ones on all others. Recently added Grok and it has somehow bacome my second most used model.

For daily one off questions I prefer it because it is fast enough and I like the way it responds. I also use it for basic research like “find me a battery drill for this and that”.

Kimi and GLM feel extremely coding oriented. I use them for code reviews basically. I hate the way Anthropic models talk. GPT takes too much time and effort for that kind of stuff for some reason.

Grok happened to be a nice middle ground.

brandonagr2 7 hours ago|||
You should try it, it is less sycophantic than other models and is faster and better at most reasoning levels, don't confuse the twitter bots and services also named Grok with the frontier model itself
venzaspa 3 hours ago||
Perhaps he doesn't want to use it because it's owned by human being who many people view as vile.
swozey 7 hours ago||
I can't take anyone seriously who uses grok seriously. I like to look at the cybertruck owners forum every so often because it's just... hilarious. And the amount of superfluous grok use over there is just insane. Half the posts I click in there will have a bunch of people dumping entire grok takes "why do people hate cybertruck owners?" "Because they're jealous and poor," sort of stuff that they just LOVE to post.

As a technical point of reference to compare against other llm stuff, sure, I'll glance at a report or benchmark but I really couldn't care less about anything to do with the project and it could blow other options away and I wouldn't touch it.

ElectronCharge 6 hours ago||
Possibly interestingly, I can't take you seriously for having such a superficial approach.

You probably shouldn't cut off your nose to spite your face.

totallymike 56 minutes ago|||
Having ethics is not generally considered superficial. Musk is a deplorable person, and choosing not to use his CSAM generator seems like a pretty good idea.
mempko 5 hours ago|||
I don't know man, Musk doing Nazi salutes doesn't seem that superficial. He did help get Trump in power and also killed a lot of aid to children that need it.

What's superficial about refusing to use a product from someone like that? Or are you one of those 'technology isn't about politics' people? That's a superficial take if you ask me.

All technology is political, and understanding that is a deep, not superficial take. It requires systems thinking which unfortunately many people building technology seem to lack, despite software being a sophisticated complex system.

Saline9515 8 hours ago||
I tried in Omp (Oh-my-pi), and so far it's really problematic.

It will loop in thinking mode ("Let me implement those fixes: Fix 1, Fix 2, Fix 3 .... Fix 80, Fix 81"), ignore the AGENTS.md instructions, corrupt plan files, etc etc... I have 5.6 Sol as advisor/watchdog, and it blocks every turn, I never saw this. Quite a shame, 4.6 wasn't so bad.

xmorse 7 hours ago|
OMP is a joke. don't use that garbage
Saline9515 7 hours ago|||
Can you explain your opinion? I'm curious but such vague comments won't convince me.
samtheprogram 7 hours ago|||
Probably the same reason as oh-my-zsh, you don't need 90% of it. Further compounding the problem in an agent harness is that you are polluting the context window by throwing the kitchen sink at it.
marwatk 7 hours ago||
I've been experimenting with omp because:

- it allows different models within one session via roles (I only have API, so pay per token)

- it's much more likely (ime) to use the LSP over grep for determining how code fits together

But I agree a 20k+ starting context is way overkill.

I find it's very hard to get information on harnesses people are using. I have to stay model agnostic so I avoid claude, codex, cursor, etc. I've used and tried opencode, which worked well, but obviously lacks the above features.

Does anyone have a resource for following what people are actually being productive with? With so much vibe going on it's hard to separate the wheat from the chaff.

xmorse 6 hours ago|||
this summarizes the average OMP user and dev

https://x.com/greg_horvay/status/2100764473392820433?s=20

raincole 6 hours ago|||
What a crazy thread lol. I really can't tell who is serious and who isn't there.
Saline9515 4 hours ago|||
Again, this is vagueposting, what do you have against snapcompaction? It seems to be a nifty way to save token costs. At least it's an interesting innovation. https://x.com/_can1357/article/2064802476742574459?lang=en
polytely 7 hours ago||||
what do you use and why do you prefer it over omp
raincole 6 hours ago||
Just pi. `pi install` the packages you actually need or ask LLM to write a package for you. Keeping the harness minimal is the point of pi.
unrvl22 7 hours ago|||
you are a joke if you think omp is a joke.
kristofferR 8 hours ago||
What's with the deceptive graph on top? Not including Astra can't have been an oversight, did the model compare poorly to it?
Jcampuzano2 8 hours ago||
https://openai.com/index/our-decision-on-cursor-following-it...

This explains why. Mentioned in another comment, but cursorbench explicitly tests with Cursor as the harness, and OpenAI doesn't allow them to use Astra in Cursor.

user43928 8 hours ago|||
> with a proposed shutoff date of November 12, 2026

That said, I don't expect them to benchmark Astra in their Cursor harness given the situation.

Jcampuzano2 7 hours ago|||
Cursor never added Astra to its consumer subscription plans. And it's likely exactly because of this announcement. Why would they add support for a model they would have to remove shortly after?
oh_no 5 hours ago|||
shutoff for existing models, new models stopped as of that announcement, astra will never be on cursor.
kristofferR 8 hours ago|||
That's not accurate. OpenAI doesn't allow Grok to provide Astra to Cursor customers anymore, but it doesn't ban anyone from using Astra via alternative harnesses.

If Cursor wanted to include Astra in CursorBench nothing would stop them, they could easily have spent half an hour vibecoding in OpenAI API key support - if it hadn't been convenient to neglect to do that.

andsoitis 8 hours ago|||
Even if they could do that (workaround to include Astra in CursorBench), that has no practical consequences for Cursor users and that's what I as a Cursor user (what I use for dev, though I use ChatGPT for non-dev stuff) care about.
kristofferR 8 hours ago||
It would make the benchmark way better obviously, by showing how their new model compares to their competitors, the whole point of benchmarks and graphs.
Jcampuzano2 7 hours ago||
The point of Cursor Bench is to show how models perform in Cursor. If 99% of their users won't be able to access a model unless they go out of their way to include setup an API key for it (which would be insanely expensive with Astra), why would they include it in the benchmark?
mh- 4 hours ago|||
If the benchmark is "what's the best model to use in your Cursor subscription", why would they do that? OpenAI knew what the effects of their decision were. Hard for me to have sympathy for either party here, honestly, and I say that as someone who is a customer of both.
scottyah 8 hours ago|||
Deceptive? An extremely quick google search would answer your question. OpenAI pulled out of Cursor before they released Astra so it never got that benchmark.
kristofferR 8 hours ago||
Pulled out from letting them resell Astra access, that's not a limitation on running a benchmark.
scottyah 1 hour ago||
It is, that's how that benchmark is run. Another very quick google search.
Iolaum 8 hours ago||
I wonder if that means that SpaceX evals show that they consider astra better than fable or that they hate Sam&co so much they don't want to show their stuff.
Jcampuzano2 8 hours ago|||
https://openai.com/index/our-decision-on-cursor-following-it...

Its because of this. You can't use Astra in Cursor, and cursorbench uses cursor as the harness. They can't actually benchmark it using their harness hence why its not included.

babelfish 8 hours ago|||
They have Astra in other benchmarks lower on the page. They just don't want to show it winning
Jcampuzano2 8 hours ago||
The chart is cursorbench though and they asked about the "deceptive graph"
ryeguy 7 hours ago|||
They can benchmark it because you can use an openai api key with cursor. Astra is just not included in the cursor plan.
bluecalm 5 hours ago||||
Elon posted on X that Grok 4.7 is behind Claude and OpenAI for agentic coding:

https://x.com/elonmusk/status/2102082011233931762?s=20

so it's likely about usage in Cursor specifically.

DavCreator 4 hours ago||
https://xxcancel.com/elonmusk/status/2102082011233931762?s=2...
babelfish 8 hours ago|||
this is exactly it.
usumgallu 6 hours ago||
[dead]
Tsarp 8 hours ago||
Waiting on simonw "Generate an SVG of a pelican riding a bicycle " benchmark to judge this model
forgot-my-pw 7 hours ago||
It might be more capable, but AA indicates it's a lot less token efficient than Grok 4.6: https://artificialanalysis.ai/agents/coding-agents?agents=co...
rvz 8 hours ago||
[flagged]
jcims 8 hours ago|||
We're allowed to have our ceremonies.
kridsdale3 7 hours ago||
Thank you. If this whole thing isn't fun, it isn't worth doing.
user43928 8 hours ago|||
You don't think it's useful to learn whether a model's "intelligence" generalizes beyond the tasks and modalities it is usually optimized for?
TylerE 8 hours ago||
Absolutely not. Makes about as much sense as judging a car based on how good an airplane it makes.
lumirth 7 hours ago|||
Have you considered that the single most impressive breakthrough of LLMs as a technology is their ability to generalize beyond what they were explicitly trained on? Great analogy, pal, but LLMs aren't cars.
user43928 7 hours ago|||
I disagree. If GPT-7 can draw the Mona Lisa in MS Paint via computer use, this would be interesting.

That it isn't the most efficient way to achieve the same end result is irrelevant.

bluepeter 8 hours ago||
[dead]
nicolamanzini 4 hours ago||
[dead]
mempko 5 hours ago||
[flagged]
blactuary 4 hours ago||
And poisoning Memphis. So disappointing that no one has principles anymore
13415 4 hours ago|||
There is no need to use it anyway, it's always been uninteresting in terms of performance. However, for me the red flag was when Musk admitted he will personally interfere in its training and prompts to make it more of a propaganda tool. I'm interested in science and reality, not in the political delusions of elderly drug addicts.
Shiggy_ 5 hours ago|||
[flagged]
DaSHacka 5 hours ago||
[flagged]
felixgallo 7 hours ago|
[flagged]
knicholes 7 hours ago||
How do I obtain this morality build?
inferniac 7 hours ago||
[flagged]
oulipo 7 hours ago||
We do believe that Musk is fascist
Vaslo 6 hours ago||
No, we don't
More comments...