Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Need one for vision models too tbh. Token/s doesn't really map easily
  • Could you add cost estimates too? I usually care about the tradeoff between VRAM/throughput and roughly what the same workload would cost on different providers.
  • How about the other way around? I'll tell you my hardware and you tell me the possible stats? I've been having some problems with bigger models, using MoE to make it work for longer/tedious work like nightly sweeps that don't really on speed.

Explore Birbla archives