Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Random question, but has there been any improvement in OCR/document understanding in these newer models? Last time I checked (1mo ago) SOTA was still sadly Gemini, unless you wanted to pay $$$ for e.g. Sol
  • Not sure about OCR specifically, but the newer (past quarter) vision language models all have a lot of post training on detecting garbled text specifically. You can feed some of the old stable diffusion outputs into a modern model and they can figure out pretty reliably if the text gen is mangled. I think the feature is probably used as part of the RL for the image gen to correct for bad text rendering.
  • I truly wish these models were available when I was in University. As a visual learner it would have been much easier for me to understand certain topics with illustrative diagrams rather than reading a wall of text.
  • My wife has taken all her recipes and fed them through ChatGPT image gen to make zine pages and they’re really cool! She’s building a recipe book for the kids so they’ll know all the recipes from their childhood.
  • …but the illustrative diagrams are a simulacrum; if you ask Qwen, or any image-generator, for an “accurate” poster-design featuring a representation of a model of an atom and explaining its constituent parts I expect you’ll get an imitation-airbrush rendering of red, blue, and grey table-tennis balls orbiting in perfect circles; you might get an electron-shell diagram if you’re lucky. What you won’t get is anything remotely related to probability-clouds.

    Edit: just to test myself I asked Nano Banana 2 to generate “an undergraduate infographic poster about how atoms work” - and the result was something right out of a middle-school science textbook and very Bohr…

  • > As a visual learner…

    This may be a myth. Big paper in 2009, but nobody's proved it's real in the 15 years since.

    In a 2009 review paper entitled Learning Styles: Concepts and Evidence, researchers investigated the “meshing hypothesis,” which is the idea that students learn better when instruction is provided in a format that matches their learning style. Their conclusion is a hard pill to swallow. “The contrast between the enormous popularity of the learning-styles approach within education and the lack of credible evidence for its utility is, in our opinion, striking and disturbing,” the researchers wrote. “If classfication of students’ learning styles has practical utility, it remains to be demonstrated.”

    2009: https://www.psychologicalscience.org/news/releases/learning-...

    2022: https://fee.org/articles/learning-styles-don-t-actually-exis...

    APA goes so far as to say (2019) believing in learning styles is detrimental:

    https://www.apa.org/news/press/releases/2019/05/learning-sty...

  • The "piss filter" is still everywhere.
  • The examples posted on their launch blog page are quite impressive, especially for fine details, multi-panel/multi-page and text rendering.

    But: not open-source/open-weights, and no indication that weights/source will be released either.

  • I am surprised by the rather bad output. It doesn't achieve qwen image 1 quality in composition or anatomical correctness. Tested on chat.qwen.ai

    I am seeing third legs and glowing eyes. It's a Microsoft Lens level of quality and that one was pulled.

  • > Supports up to 4.5k token input, effortlessly generating complex layouts such as newspapers, storyboards, and exam papers.

    Impressive.

    Btw, what is currently the best model to run locally on a 16GB Vram? Is it Z-Image Turbo?

  • It depends on what you need, but Krea/Klein9b/Ideogram4/Z-Image are among the best right now for text2image and Qwen Edit and Klein are probably still the best at editing.
  • Krea-2-Turbo. I've even got it working locally on my M5 iPad Pro.
  • Not a single word about when/if they'll actually release the weights for this, or am I missing it somewhere?
  • Why release it just so Cursor/Azure/Amazon can profit off of it? Unless OpenAI actually opens it up fat chance.
  • After Z-Image and the original Qwen-Image - I think they've pivoted completely to closed source. From the images I've seen, Qwen-Image 3.0 is just a subpar equivalent to other proprietary models like gpt-image-2 and nb-pro.
  • Considering Qwen-Image-2.0 weights have not been released either, it unfortunately looks unlikely.
  • > to precisely describe the full 3×3 grid takes a full 3.7k tokens

    It's a shame they didn't share that prompt - it would make that demo more convincing.

  • Very cool but the Arabic text in the title image is obviously and hopelessly broken, which is oddly not the case when actually using the model. Could it be that the hero image was not generated by Qwen Image 3.0?
  • that's a very interesting point to test new models. I speak an Indian language called "Tamil" and have always tested new models with Tamil but also have tried little bit of Arabic (quranic verses) with previous GPT image 2 and Nanobanana pro and they have nailed it. Don't know if it was because of extensive training data.
  • You clearly haven't met a Chinese RedNote user.
  • AI-generated images are part of the web now, if you're doing ordinary web scraping, you can't avoid training on generated images.
  • yellow/red tint is an extremely common problem not matter the photograph source you train on

    source: work at a photograph start, even training on raw images things get tinted, it is an uphill battle

  • Maybe it was trained in a Mexican data center https://en.wikipedia.org/wiki/Mexican_filter
  • GPT Image 1 ended up with a yellow tint without training on another image-generation model's output. It's just that humans like pictures with a soft sunset glow, and this is a very easy global signal for a preference model to pick up on, and for a image-generation model to imitate. So optimizing for aesthetic appeal makes everything slightly tinted by default, unless you make sure to countersteer.
  • The meta keywords in the HTML is very interesting. 100+ references to NSFW topics such as hentai, nudes, etc.
  • Interesting too that they would include "ai friend" when China just added restrictions on AI "partners" such that many services stopped offering them rather than try to adhere to the restrictions.
  • Wow you really weren't kidding. Fully automated AI slop tentacles for everyone!

    https://pastes.io/uenL6X9K

    It also seems to have an obsession with this celebrity, based on how many times ctrl-f for "stefani" turns up a result.

    https://en.wikipedia.org/wiki/Gwen_Stefani

  • That is hilarious as porn is illegal in China. But I guess pron SEO is allright, if it is against the west.
  • I think the Qwen team is probably aware that the NSFW community is very quick to adopt any new image gen model (see: Civitai). So, it seems like a good SEO approach to try and surface their model in search results that said community is likely already checking.