Incredible.
> On FrontierCode, which evaluates whether coding agents produce changes ready to merge into real codebases, GPT‑6 Sol improves substantially over GPT‑5.6 Sol, and is able to match Claude Fable 5.1 xhigh at much lower cost.
I continue to appreciate OpenAI's attempt at some honesty here, showing that they are capable enough and have skilled engineers to a point where they can recognize that slop is hated for good reason, and that there is a real issue. Compare this to anthropic, where e.g. in the Opus 5.5 announcement[1] one of the first points on the page is
> One tester completed a 680,000-line code migration in less than a day—work that would have taken an engineering team weeks. It’s good at finding and fixing inefficiencies in software: when we asked it to cut load times across every page of a web app, Opus 5.5 succeeded 39 of 40 times, while Opus 5 made smaller improvements that also altered the app’s behavior. A different tester had several Claude models build a game from a single prompt; Opus 5.5 scored higher than any other model on the strength of its graphics and polish.
This is the kind of shit that is the very reason why I stick to OpenAI and deepseek. OpenAI is simply more honest and reasonable about their models' capabilities, while delivering models that still have solid value.
Notice how the OpenAI announcement doesn't make use of anecdotes.
I guess the chinese competition spooked them.
Does anyone care about code quality anymore?
This will make me more valuable in the future when everyone has lost the ability to do anything on their own.
I respect that you want to learn how things are done, that is a great trait. But once you learn how its done, you should use the tools to free up cognitive load for more difficult tasks.
Have the frontier labs stopped trying to increase context window size?
OpenAI seems really competitive in most areas, and extremely competitive on cost, but still behind on coding.