Open-sourced jev architecture last year with model,paper and dataset

Open-sourced jev architecture last year with model,paper and dataset

21 pointsby nandakishor_ml18 comments

Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Pity because without the frustration it would’ve been a good post. Catch is you’ve got intuition, but looking around instead of forward. Do it again, open-source it, either you’d quietly bring down few companies, or, when you’re close, you’d get offers from them. The who’s done what is a dying paradigm.
  • I've been saying it all year. Brand name is the only thing that truly matters.

    humans like simple, simple is shorter, shorter is easier to remember.

    if anyone can build anything, what gains value... IP/domains.

  • Just like the algorithm for short videos, even if you implement its prototype, it is hard to attract attention when you lack marketing and do not have an out-of-the-box product (you have the model weights, yet few people will deploy them). It will not become an instant hit like the Jev model; even others' secondary development may be more popular than your prototype.
  • Really interesting setup. Given the recent interest in non-autoregressive fast probability prediction, do you plan on extending your framework into a more horizontal, general-purpose open-source model?
    by 8MEM
  • I tested your model quickly and it seems that the model does not do what is claimed here (that it can do the same as jev). It relies on an LLM, while jev does not need an LLM. While it is interesting and novel work it is not a replacement for expensive LLM calls.

    For anyone interested in playing around with it, it requires python 3.11, else it will segfault.

    According to the code the conversation turn index is a signal that enters so it is judging on something that does not transfer to other tasks. In the sampling it moved the score more than the actual content.

  • The world's heavily about marketing, resources, connections, and signaling, unfortunately. You probably needed to market it in a bigger forum with shinier claims to attract attention (I don't think a paper on Arxiv is enough).
  • > It's incredibly frustrating that the thing that you made with months of hard work, sweat and sleepless night is architecturally similar with the vertical use case and don't get the support you deserve because frontier lab build something horizontal.

    Totally get the frustration.

    Don't get too bent up about it -- take it as the market validating your hypothesis in a way you didn't expect. You should feel proud! That skillset - finding niches that could be humongous given the right cultivation, luck, funding, and marketing - is amazing.

    I feel strongly that, if JEV and/or your model prove valuable as many of us are already thinking and hypothesizing, people will come knocking.

    While LLMs are a neat space, this is a desperately missing component, and as someone who has built conditional choice probability models for almost two decades, I'm wildly interested to dig in this weekend and review the usefulness, the applicability as a homo economicus level of automation in a noisy prompt space (e.g. ensuring transitivity and IIA), and looking at this as a strong evolution.

    I've added your model to my eval list!

  • From what I can tell these are classifiers trained for a single task. What excites people about Jev is that it can do zero-shot structured responses for arbitrary prompts. Now, this isn't new either; models like GLiNER have been around for a while. But Jev appears significantly more flexible and polished while still being cheap and fast.

Explore Birbla archives