Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • I agree. It did very well on an extremely challenging task.

    I asked it to recognize and draw the very faint reflection of what I was wearing, visible in only a tiny black part of a very brightly lit poster behind glass.

    In addition, the poster itself also happened to contain similar clothing.

    You can see the reference images and its output in my writeup here: https://medium.com/@rviragh/gpt-5-6-sol-very-good-image-reco...

    While a human can focus on the reflection easily, this is an enormous challenge for a vision model. It's very impressive.

  • Gemini 3 Flash should really be included in this comparison. Or at least 3.7. In most of my testing, 3.5 and 3.6 were both a downgrade in terms of vision capabilities, relative to 3, and at a much higher cost. 3.7 is slightly better than 3, finally.
  • In the third vision bench result, Sol is 100% correct but the expected has 1 error. Seems like an oversight.

    In the next bench, Sol looks like it’s correct again but the bboxes are rotated 90 degrees for some reason.

  • Ironically, the pill counting example selected to showcase "the best vision model" can be easily solved with OpenCV template matching, a technology created 25 years ago.
    by mv4
  • It is funny to me seeing Sol used for what a "traditional" AI model can do already (counting pills).

    We have vision models for our pharmacy and I could never imagine taking the latency hit to use a Sol in our robotics, it would be likely 25-50x slower.

  • Penny sample shown looks like failed EXIF orientation registered by the model/harness. The coins are correctly marked, it's rotated 90 degrees.
  • Anecdotal, opinion:

    Gpt is really good in vision stuff, or at least their MoE seems to be really cohesive. From my experience Claude models can be really good at language but the moment they need to look at a picture and decide why the design is not good what parts need improvement it degrades a lot. My easiest benchmark is giving them a screenshot of a feature in my app and tell it "identify non-normative UI blocks and improve readability and consistency". Sol does a great job at re-structuring the page into composable units that build upon each other and the general looks and feels of the app. Claude tends to over-focus one one part while completely forgetting about the rest or the cohesion as a whole.

    by weli
  • The summary "There are still clear limits. Gemini 3.5 Flash remains a better practical choice [than GPT 5.6 Sol] for high-volume detection and counting in our benchmark, especially at its price." seems rather understated !

    GPT 5.6 Sol was outperformed on all benchmarks by Gemini 3.5 Flash, apart from a single exception (OCR) where Fable was the winner.

    Gemini 3.5 Flash not only outperformed GPT 5.6 Sol, but did so at 1/3 of the cost.

Explore Birbla archives

GPT 5.6 Sol is the best "vision" model OpenAI ever released · Birbla