Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- I've been using it for the past few days, and it runs really well!
I usually get 7 token/s using llama or lm studio, but this inference recipe runs at a smooth 80 tokens per second.
Genuinely very usable, and fully local!
by Pragmata - I'm not sure if I want to trust yet another loader with secret sauce, there are already a lot of those around.
What I'd really like is a simple utility that, given my system and a model, will tweak llamacpp to run decently (or tell me it can't be done).
by toyg