Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • So, instead of resolving it privately like you were asked, you come here and complain? Was the bot posts on Reddit not enough for you? Pro tip: There is a right way and a wrong way to approach these things and you are MOST certainly approaching this the wrong way. And before you go accusing me of anything: I am NOT associated with that project, but I DID see what happened. You are acting like a damned child and should be ashamed of yourself.
  • How much does this speedup inference for the end user in terms of tk/s ?
    by S0y
  • As an update to this, Lemire absolutely crushed it! I benchmarked his change today. Here is my PR comment:

    https://github.com/jadidbourbaki/llama.cpp/pull/12#issuecomm...

  • Fun update to this: Daniel Lemire added another optimization to make this even faster. https://github.com/jadidbourbaki/llama.cpp/pull/12

    I’ll benchmark his change and add it to the article, crediting him for this improvement.

  • Btw, if anyone has experience with the open source community in general and llama.cpp in specific, I would greatly appreciate some advice. I’m facing a bit of an interpersonal issue that I really hope is resolved without any ill will. Here is the context:

    https://www.reddit.com/r/LocalLLaMA/comments/1wr5ylm/comment...

    Any advice for what I can do? Due to this, I cannot create a PR or issue in the llama.cpp repository. However, I am worried about bothering the maintainers on other channels in case it aggravates them further. Thank you for your help!

Explore Birbla archives

42x faster prompt lookup drafting in llama.cpp · Birbla