Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- 100 tok/s on a 4090 could unlock entirely new use cases. What becomes possible when local inference is this fast?by c4pt0r
- Why isn't this type of expert caching in the native llama.cpp yet? Why do we need a separate codebase?by lxe
- Dwarfstar already supports this, curious how it compares, but I use the q4 quant daily and it works really well.
https://github.com/antirez/ds4/blob/main/docs/MODELS.md#qwen...
by kamranjon - LLM threads the world over are spammed with Strata links, it remains to be seen how much of the breathless hype remains standing once the honeymoon period is over. I've tried it but so far I have not seen anything that overly impressed me in terms of accuracy, though the speed is definitely there. I'm sure there are applications for LLMs where the quality of the answers is less important but I don't have any of those. YMMV.by jacquesm
- I've been working on support for this model in ds4 on the RTX 6000 pro - it's been really great for my use cases. The ds4 q4 quant performs a lot better than other similar sizes that I've seen.
Using the Q4 quant on an RTX 6000 Pro Workstation Edition at 450 watts:
Most important for me, I can run 4 concurrent streams at 400+ tok/s.Code: prefill 1,251 tok/s decode 255.26 tok/s Prose: prefill 1,251 tok/s decode 198.78 tok/sby AntiRush - I tried it and it worked surprisingly well. On my machine (Nvidia 4090, 128GB DDR5, Ryzen 7950x3d) I'm getting 124 tokens per sec, thought to share it here.by snehesht
- I've just tested Strata on a simple 50 image vision benchmark. The task is to output the exact coordinates of a requested object. The result via Strata had a median error distance of 154.8 pixels, avg of 168.8. Running the exact same GGUF and vision adapter weights on llama.cpp gives me a median error of 46.5, avg 81.4.
To put that into perspective, here are some more numbers from other models via llama.cpp:
Median/Average
Qwen 3.5 9B BF16: 46.5 / 193.3
Qwen 3.6 35B Q4 K XL: 38.4 / 76.4
Qwen 3.5 122B Q3 K M: 32.9 / 68.6
The difference in vision performance is as large as the jump from a 9B model to a 35B model. All tests were performed at temp=0.
I have done no further testing, as these results line up perfectly with my expectations.
by Jackson__ - I'm a little skeptical of going below 4-bit quants due to the potential for significant degradation in quality. I'm running 4-bit quants on an RTX Pro 6000 rented for approximately $1/hour and getting about 1.2 million tokens out and 40 million tokens in per hour with caching. The quality of 4-bit quant is good enough for difficult but well-scoped coding tasks. Here is the inference stack I am using: https://www.reddit.com/r/BlackwellPerformance/s/FrKwk3GoDKby a11r