Posted by spIrr 9 hours ago
It's hard to understand what's going on with Grok. It's like it has capabilities in a theoretical sense but maybe the training is so focused on being in x.com/grok.com with the web search tool enabled for "is this true?11" type queries that with any API type usage with document workflow instructions, tool use, code gen etc it completely falls over
https://venturebeat.com/technology/cursors-composer-2-was-se...
Plus the model's capacity to take more context into account and actually integrate it to the output is simply limited by the number of activated parameters. If you give it a playbook, you are forcing to choose it between attending to the playbook and the task at hand.
If you want to force it to work step-by-step, you need to present the steps one-by-one. Ideally with rules for the current step at hand and maybe relevant input again, depending on overall task size.
Why did you think models love to re-read files before editing them? It increases recall quality and thus edit precision and thus benchmarks.
not sure I got it?
Separately, the frontier labs are kinda pushing us into that behaviour by releasing models with ever-larger context windows.
Glad to be able to put some numbers on it.
It does what its has been trained to do. So find out what its trained to do and just use it to do that. This is not general intelligence.