Top
Best
New

Posted by plurby 1 day ago

GPT 5.6 Sol is the best "vision" model OpenAI ever released(blog.roboflow.com)
355 points | 164 commentspage 5
dangoodmanUT 16 hours ago|
I hate these "The best X thing Y has ever released".

Unlike when Apple says "it's the best iphone we've ever made", LLMs are more or less interchangeable. So "OpenAI's best model" means nothing if "Anthropic wipes the floor with them" or "[open weights model] is 10x cheaper for 1% less quality".

As a reader, it feels like these titles are click bait.

dzonga 14 hours ago||
the last image - it's barely visible to human eyes
sam0x17 16 hours ago||
> GPT 5.6 Sol is the best "vision" model OpenAI ever released

I mean I should hope so, as it is also the latest one

iamleppert 1 day ago||
Where are the Qwen benchmarks in this? I would be more interesting to see how Qwen performs.
SkalskiP 22 hours ago||
Hi! I’m the author of this blog. I regularly benchmark new VLM releases. You can check the results for Qwen3.8-Max and Qwen3.8-27B here: https://playground.roboflow.com/evals
ImageXav 23 hours ago||
Me too. This is an interesting comparison but in my experience Qwen and Gemini have typically been the top contenders for image related tasks. For that reason it would be great to have the comparison here, as I'm not surprised by Gemini's dominance over the other models.
RugnirViking 1 day ago||
It's really quite good! I was amazed recently by its utter inability to read some faded handwritten cyrillic on the back of a wood carving - 3 or 4 words only, reasonably clear letter forms I found recently, and then stepped back a bit and thought about how insane that was as a benchmark - I just expect it to work so reliably on other OCR and translation tasks that it was surprising to encounter such a failure
terhechte 22 hours ago||
Fuck ack. I'm working on a new benchmark that combines strong visual requirements with tool and coding requirements. I haven't even tested Sol yet, but between Sonnet, Terra & Luna I already see much better results from OpenAI's models. I'm not releasing anything yet as I still have issues in my harness that need to be fixed.
slybot 18 hours ago||
Am I the only one who cannot read the date on the blister pack even fully zoom in my phone?

If that is the full quality image given to the model, I think it's not surprising that the model confused with 03/2022.

fooker 23 hours ago||
I'm a little bit disappointed that vision seems to fall before language at scale.

It seems pretty counter intuitive that we can't do vision significantly better with specialized techniques.

TZubiri 22 hours ago|
Which is to say, still not ready for any production workloads yet. As in, it cannot reliably count the amount of objects in an image.

Still very impressive, but nowhere near the text chat revolution. OpenAI still trying to strike their second lightning

chistev 16 hours ago|
https://news.ycombinator.com/item?id=46444508
More comments...