Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • I’m honestly surprised this is better benchmark wise than the text only model. I figured the addition of vision would take away from some of the text capabilities.
  • I'm not that impressed with this version. When using it, it often acts as if it has no visual capability and refuses to recognize images unless I remind it.

Explore Birbla archives