Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • This is so cool, and having a live demo is just chefs kiss
  • I didn't expect the 2,000 connection sweep to stay flat, since all of them are sharing one stream. What does per user latency look like at that end of the sweep?
  • Amazing project, I love it.

    How about using these Cactus models?

    Would it make sense for you to collab with those guys (1) for you to design a cheap but improved, commercialisable version of your $250 chip and (2) for them to tailor their runtime and quantizations to such FPGA hardware?

    https://github.com/cactus-compute/cactus

    Also have you thought about using a Alveo V80? Still not crazy expensive and could fit bigger models with same approach

    by gbxk
  • But the problem is not that your model is fast.

    Sure, you can go ASIC and go even faster, but the thing around GPU is that they scale well for both training and inference, and the technical floor is low.

    The level to get into FPGA design is insanely high, you've got to read timing diagrams, you need to know combinatorial and sequential logics and good sense of boolean algebra, you need to have an asynchronous signal based mindset which is vastly different from CPU/GPU, you need to know netlist and you need to endure the time it takes for the EDA to finish generating it. Yosys is still years behind Xilinx

    There is a reason GPUs are called accelerators; it sacrifices and does not try to really specialize on one particular thing, except high parallel dataflow and branch-free calculation. Otherwise we will all be using DSPs

  • Really cool project. I wonder what the future of LLM inference will look like. The Talaas demo is promising, but using an ASIC with weights in ROM means you can’t update the model (weights or architecture) without replacing the entire chip. SRAM isn’t dense enough to store model weights, but DRAM has bandwidth issues unless you use HBM which is expensive. Maybe novel memory technologies are the future (there are a number of emerging technologies in R&D), but they likely require breakthroughs to become commercially viable. Systolic arrays could work, one could imagine architectures where routing (architecture) is fixed but weights are programmable, or architectures where the weights and routing are programmable but the compute units are fixed function (coarse grained architecturally reprogrammable), or maybe the weights are in ROM but can be hot swapped easily with into fixed compute elements. Definitely an interesting and emerging field.
  • Ignore the naysayers!

    Any article, even the really good ones on HN, while they get positive comments, for whatever reason, always get a lot of negative ones, too...

    That is, the negative comments are absolutely unavoidable, even for people accomplishing great things!

    I personally think that what you've done is brilliant, absolutely brilliant!

    I can't wait to see more in this space...

    Brilliant, absolutely brilliant!

  • Hi everyone, I’m a friend of Mike’s; he’s having issues replying to the post at the moment, but hopes to post a thorough reply to the comments as soon as possible
  • Crazy how little traction this kind of project gets on here. Posts about squeezing a 1+T model to seconds/token and you have this wave of optimism like "It's the effort that counts! We'll get there!". Sure, this particular project isn't really scalable in the same sense (PL fabric/use what ya got/cost/power) but IMO it's conceptually a brilliant thing to showcase comparatively. I have a strange feeling a decent chunk of people dismissing this project are the same who spent small fortunes on hobby llm inference setups/investments and see this as a useless exercise. Meanwhile dozens of $$$M startups in the CIM/analog compute/etc have been R&D'ing for years now that will make this same outcome a reality before we know it (crazy inference speeds on usable models within local reach). Anyways kudos to OP and really enjoyed the documentation and findings of this!

Explore Birbla archives