

Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- Checking openrouter (it's not available yet) and, uh, what's up with the spike in Qwen usage from early april here? https://openrouter.ai/qwen
Is this normal humans kicking the tires on a new model, or a few whales doing serious benchmarks?
by eleventen - personally seen a lot of people switch to Kimi and Qwen after Opus 4.7. Kimi 2.6 feels like Opus 4.6 which, to me, was a great model for 98% of coding tasks
- Qwen 3.6 Plus released and they offered it for freeby d2kx
- Any reports from people using their coding agent(s)?by bsenftner
- I'm using the pi-mono coding agent (open source, free) without any extensions and very simple prompts. The 3.6 27B model (BF16, 250k context) uses 67GB VRAM on an RTX PRO 9000.
It's very capable on almost any coding task I've thrown at it, and very good for easy-to-medium hard scripts, new code bases.
It struggles on some complex tasks in larger code bases, e.g. using to debug and fix bugs in llama.cpp it gets close to working code but often introduces errors. For such tasks its still very useful as a search/explore tool and drafting fixes.
by vibe42 - I'm running Qwen 3.6 27B Q5 K M GGUF on a Tesla P40 and koboldcpp using pi.dev as the harness, I gotta say I am impressed. Took some setup and configuring but I already have some code it has made commited and pushed. It can be slow on my hardware at >50k tokens, but the fact I bought this one P40 for like $150 back when the LLM trend started I can't complain. (I have a second one too but I couldn't physically fit the card in my server unfortunately.)
The setup I had to do was important and I had to compile koboldcpp with a few special params for my hardware, I mostly just had Claude figure it out. I don't remember everything I did now but it was very slow and would often stop mid task, it seems it was mostly a parsing issue. It made the model seem broken/dumb, but once I had all that settled I actually am able to use this how I use Claude Code. Disclaimer, I am pretty explicit with requirements, I imagine this fails more when you leave it to figure out things on its own but for my flow its pretty rad.
Currently setting it up as an automated agent now to pull Trello cards, create PRs for them, and move the card to be reviewed.
Command I am using to run: python koboldcpp.py \ --port 61514 --quiet --multiuser --gpulayers 999 --contextsize 262144 --quantkv 2 \ --usecublas normal --threads 4 --jinja --jinja_tools --jinja_kwargs '{"enable_thinking":true, "preserve_thinking":false}' \ --skiplauncher --model /data/models/Qwen3.6-27B-Q5_K_M.gguf --smartcache 5
by rayboy1995 - QWEN really hits the sweet spot it's cheap, fast, and actually good.by jdw64
- No opus 4.7 , gpt5.5 , Gemini flash 3.5 in benchmarksby maxdo
- TBF, Flash 3.5 was released 2 days agoby piyh
- Downloading this and cancelling Google Antigravity Pro at the same time:
I had a Google Pro account that I inherited from buying a Pixel 9 XL - it's free for a year after a flagship Pixel phone purchase. After a year they started charging for it, and i tolerated it, because Flash was usable in Antigravity for dumb auxiliary tasks that I did not want to waste GPT/Opus on. It had a separate generous quota from Gemini 3.1 Pro. Now with Flash 3.5 they combined the quotas with Pro, such that on a Google pro account you can work 4-5 hours per week in Flash. And by the way, 3.1 Pro is useless for programming, compared to Codex/Opus
by cft - same boat. Google Pro AI quota became barely useful for anything meaningful.
I think they envision Pro plan as "just a taste of AI, enough to lure folks into the Ultra plan" but that won't work for me when Codex is half the price and DeepSeek 4 Flash is 1/10 of their price per task.
So I'll downgrade just enough to keep my Google Drive space. And use DeepSeek 4 as workhorse plus Codex or Copilot for advanced stuff.
by bel8 - I'm using pi agent and love to try qwen models (hosted). What are the good options? The official provider doesn't include Alibaba. Is OpenRouter etc. fast enough?
(As a reference, DeepSeek v4 is severely throttled on these proxy services.)
by flakiness - i use opencode zen as a convenient pay-as-you-go way to try out all these new models. it doesn't have 3.7 yet, but at the rate they usually update it probably will tomorrow.
I couldn’t say how throttled it is, but it seems fine?
by notatoad - I use pi + openrouter (with qwen3.6-max-preview) a lot. I never hit any stability or performance problems yet.by atilimcetin
- Is this one of those ones where they'll drop the huggingface release a week later? Or do we know for sure that this is staying proprietary?by ndom91
- someone correct if i'm wrong, but I think the max models are usually non-openby Davidzheng
- Looking forward to more open weight releases from Qwen, especially 122B and 397B.by tarruda
- I am still waiting for qwem image-edit 2.0 open weightby guitcastro
- Ouch. I'm just getting into tinkering with these things - mine is running on a vanilla gaming desktop with a 12gb 3060 and 32gb of ram. Even going above Qwen 9B risks completely locking up the machine.by Pxtl
- I'm more excited for qwen3.7 9b and 72b, these are usually so good for their size
- Personally even more a lower quantized model like 9B.by ricardobayes
- Yeah that 60-150b~ range is such a sweet spot for current 'prosumer' hardware, I'd love to see something like a 120b-a14b or there about.by smcleod
- These are very good numbers. I still don’t get why they don’t compare against latest competitor versions in these posts, it’s not like we’re all not going to notice.by goyozi
- Marketing.by maelito
- this puzzles me too, I want to knowby hmokiguess
- honestly, initial version of Opus-4.6 was much better than whatever we are being served right now as 4.7. If it performs same level to that, i'm totally willing to switch.by beydogan
- I think its part of the expectation setting (with a side of we did our distillation/ eval harness on a specific model).
if they say it's 4.7 comparable, it anchors that into your head as the model to evaluate against.
by htrp - I think the argument is that trying to suggest that they’re close to N months from SOTA.
Realistically I assume they hope readers don’t notice the fine details.
The Qwen models are great for open weights but for every past release they haven’t performed as well as the benchmarks in my experience. They’re optimizing for benchmark numbers because they know it works.
by Aurornis - Nobody releases numbers that show them to be worse than competitors lol.
This even applies to OpenAI & Anthropic who don't even eval on the same datasets a lot of the time.
by Eridrus - I find it forgivable if it's within minor version bump. (NB that x.5 is now a defacto major-version bump for LLMs for whatever reason).
Even with LLMs, posts like this don't just fall out of a coconut tree. If you have a set of target benchmarks for your own model, then keeping "the set" of side-by-side comparable models is its own maintenance headache.
by NiloCK