Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • That's next level masochism.
  • Should LLMs be designed to be modular, so that instead of needing access to the whole model, for a given prompt, only a small subset of the model would be used? If knolwedge was sufficiently modularized, most of it could be ignored.

    Maybe a hyopthetical model of 5T of indexable weights could be used with only 50 GB of GPU ram, efficiently, because it stays resident in the GPU.

  • There's a tendency to cite only decode speeds, but in practice, an LLM generates far fewer tokens than it has to read (unless you ask general knowledge questions). So the effective performance is much slower than the decode rate suggests, because a 512-token prompt already takes 6 minutes to load
  • I missed the explanation for how the SSDs are connected.

    Maybe a dumb question.

  • The SSDs are necessary because Apple's architecture doesn't allow RAM upgrades. Reminds me of someone who said "640KB ought to be enough for anybody".
  • You currently can't run a 2.8T locally; there's just no way. So, it's a good start.
  • A medium prompt in only 11 days.
  • Reminds me of Deep Thought from Hitchhiker's Guide to the Galaxy

Explore Birbla archives