

Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- clef from cloudflare runs on llama.cpp - being locked-in to strands cli would be a bummer and will slow down adoption.
Since it's a LoRa on Qwen, I assume this is runnable via llama.cpp. Pity that the PEFT/LoRa->GGUF translation is left to the user. Anyone got past:
$ uv run --with transformers==5.19.0 convert_lora_to_gguf.py ~/Downloads/lora --dry-run --verbose [...] File "/Users/user/repos/llama.cpp/conversion/base.py", line 630, in map_tensor_name raise ValueError(f"Can not map tensor {name!r}") ValueError: Can not map tensor 'layers.0.linear_attn.in_proj_a.weight'by hrpnk - This is a great model. I've been running it on device in chrome extension to filter things like email.
It is just the right mix of size, capability and speed to make it generally useful for adhoc bulk classification tasks.
For those wanting to run it in browser: https://huggingface.co/alxnahas/strands-decider-2B-webgpu
by miguelspizza - The ~2-week-old Intern-Decision family of models (0.8B, 2B and 4B) have the same Qwen3.5 base model family (albeit the instruction-tuned variants) and pointer head architecture.
- What I don’t get from this article is why strands-decider-2b v10 and strands-decider-2b v11 are shown as “other systems”, and with v11 notably outperforming v19 — why are these other systems, and why is v11 not the path taken?by gsnedders
- Fantastically well written. It's rare for me to be able to understand what the AI gurus are talking about, and this was written by humans for humans.
It can technically be used for a lot of use cases, I'd like people to chime in on ideas on this?
by keyle - Why is everyone calling binary choices `noul`? Does this have some meaning or is it just copying Jev’s API?by mattvr
- At this point I can't wait for a comedian to release a decision model backed by humans.
Meet Jerry- it's literally a guy named Jerry answering your questions.
by adenta - As a regular user of a bunch of specialized micromodels, I'll tell you this: you won't be happy with such a model (and its JEV counterparts) running permanently in the background on your PC's CPU. You need to offload their processing to the NPU. There are many pitfalls along the way, but the result is worth it.
NPU performance will be twice as high, while power consumption will be four times lower. No additional fan noise (if you know what I mean).
I'll wait another month until the first phase of the =battle royale= among models of this kind wraps up, put together a solution for the NPU/iGPU, and post it on HF.