Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- Looking at the frontend design examples; why do these models seem to love the "01 - UPPERCASE TEXT" motif. It's everywhere now (see https://try.cloudflare.com/, which has '01 · QUICK TUNNELS', but no "02" anywhere).by nemothekid
- I don't trust any of the benchmarks where Opus 5 surpasses Astra or Fable 5.1.
Maybe Terminal Bench 4.0 and ExploitGym are reasonable.
Terminal Bench 4.0
ExploitGymGPT 6 Astra 59.6 Claude Fable 5.1 55.1 Claude Opus 5 49.0 MiMo-V2.6-Pro 34.9 MiMo-V2.6-Flash 28.8 DeepSeek V4.1 Flash 26.8 MiMo-V2.5-Pro 1.5
DeepSWE v1.1GPT 6 Astra 42.4 Claude Fable 5.1 30.4 Claude Opus 5 22.1 MiMo-V2.6-Pro 17.8 MiMo-V2.6-Flash 6.0 MiMo-V2.5-Pro 0.1DeepSeek V4.1 Flash 74.2 Claude Opus 5 74.0 GPT 6 Astra 74.0 MiMo-V2.6-Pro 71.9 Claude Fable 5 70.0 MiMo-V2.6-Flash 67.9 MiMo-V2.5-Pro 19.0by user43928 - China will most probably win the AI race in the long run because of one major bottleneck the US has - energy. The electric energy and grid buildout in China has been massive since a long time and there is simply no way for the US to quickly catch up.
No matter how much cash you throw you can't just materialize a 100 nuclear reactors to power the data centers.
- Pelicans for Flash: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Pelicans for Pro: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
by simonw - Flash[1]: 309B total / 15B activated parameters
Pro [2]:, 1.02T total / 42B activated parameters
by stymaar - Anyone else more excited about Chinese models than American models these days? Big thing for me is affordability.by lwansbrough
- > This marks a key step in our exploration of the RSI path: scaling RL compute on verifiable, complex tasks, so the model can continuously expand its capability frontier through exploration and feedback.
Gotta love a capable open model. BUT, how can they just casually throw in that they're actively exploring RSI as if it's just another technique? Is this not alarming at all?
by Zaraif13 - I know we have strong views on what a truly open model is (open weights, open training data, open training code etc.) but I really like how transparent they’ve been about the training of this model.
The realtime dashboard they shared during training (https://mimo.xiaomi.com/rl/) was an incredible learning and teaching tool for me, and they’ve been unusually comprehensive in sharing details about their methodology (check out that tech report - it's got lots of clever behind the scene tricks like Google or Deepseek writeups) and benchmark scores (even the stuff they didn’t do well on).
If you’re releasing an open model going forward, please consider offering the community more of this transparency!
by rao-v