Running PrismML's Bonsai inside DRAM by breaking DDR4 timing rules

Running PrismML's Bonsai inside DRAM by breaking DDR4 timing rules

2 pointsby pcdeni0 comments

Comments

Hacker News

Interesting project, but the slop readme made me quit reading halfway

by deivid

You are saying this is 13 times faster? More proof please. How do we set it up? I really want this to be a real thing.

by ilaksh

Very interesting. Do you have any descriptions or writeups written by a human?

by Retr0id

Very interesting. I don't see it mentioned so I'll ask:

Would being able to alter voltage levels on the fly (eg, cells x y and z get 1.25 while abc get 1.2) expand the ability here?

by butvacuum

Jebus, that is some sloppy prose. Can people not even be bothered to write the summary themselves, anymore? AI doesn't want anything, so they can never have a point of view, so their prose rambles incoherently across all the various prompts they've seen in a project. This project sounds like the ramblings of a crazy person. Even though the fact that DRAM can do any computation is interesting, nobody should have to read this mess.

"why are we still transferring data to the compute? Why not execute AI inference natively within the memory?"

You already answered that question: 47.5 seconds per token from a tiny 2B 1-bit model model.

by SwellJoe

Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Interesting project, but the slop readme made me quit reading halfway
  • You are saying this is 13 times faster? More proof please. How do we set it up? I really want this to be a real thing.
  • Very interesting. Do you have any descriptions or writeups written by a human?
  • Very interesting. I don't see it mentioned so I'll ask:

    Would being able to alter voltage levels on the fly (eg, cells x y and z get 1.25 while abc get 1.2) expand the ability here?

  • Jebus, that is some sloppy prose. Can people not even be bothered to write the summary themselves, anymore? AI doesn't want anything, so they can never have a point of view, so their prose rambles incoherently across all the various prompts they've seen in a project. This project sounds like the ramblings of a crazy person. Even though the fact that DRAM can do any computation is interesting, nobody should have to read this mess.

    "why are we still transferring data to the compute? Why not execute AI inference natively within the memory?"

    You already answered that question: 47.5 seconds per token from a tiny 2B 1-bit model model.