Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Is it non-autoregressive?
    by k__
  • Isnt this obvious ? I would have thought people would try such things before deciding they need something like Jev
  • how is Jev cheaper if I can run locally. 0.5% prefill, 0.1% decode, 99.4% cached, latency is <20ms
  • My question is why not use Jev instead? It's faster and cheaper.
  • You can turn any sufficiently smart LLM into yes/no decision model or equivalent. I already have an existing workflow with a two paragraph detailed prompt, that sends pages of stuff to an LLM and asks it to return only 7 JSON objects. Several of those objects are binary "yes or no" choices of like, whether the content contains certain things.

    You can even do it with small not particularly hard to host local LLMs like a variant of Qwen 3.6 35B A3B or 3.8 27B.

  • Qwen3-Next-80B-A3B can already run on a 16GB M1 MacBook at around 3–5 tok/s using aggressive memory management. Could a Jev-style controller push that to 100 tok/s on an M1?
  • I have run tests with qwen 3.8 and gemma 4 in a way similar to this post (based on an open source project that also does this with gemma4).

    Getting competitive accuracy with Jev is fairly easy, if by accuracy you mean that the highest weighted answer is the right one. GLM 5.3 is complete overkill, much smaller llms will do

    What Jev brings to the table, beyond speed, is that the reported probabilities match actual likelyhoods. If you present three options, with A and B equally likely and C impossible, jev will approximately answer with 0.5, 0.5, 0. Stock LLMs don't

  • Everyone is doing this to emulate Jev, but...

    I took a random book excerpt with 23,000 words (±30k input tokens) and used it as context. Jev still responds in 800ms, sometimes 500ms. That's in the neighbourhood of 20-50,000 tok/s prefill, which is obviously not possible with normal LLMs, not even Cerebras is this fast.

Explore Birbla archives

Turning GLM-5.3-Flash into a Jev-like decision model · Birbla