Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • FYI, the title makes no sense in isolation.
  • some real numbers on my m4 max -

    image gen (zimage-nano), 1024x1024: 58s

    image -> textured mesh (trellis.2): 2m 49s

    sfx generate (5s clip): 3.6s

    music generate (8s, ace-step): 15s

    speech synth: 13s | transcribed back: 2.2s

    video gen w/ audio(ltx unified-av) 4s,768x512: 2m 48s

    text chat (laguna xs2.1): 102 tok/s - https://mlx.fast leaderboard

  • I think it’s generally a good idea to use a standalone, local solution. I also like the approach of combining multiple models with different focuses under a single interface. On the other hand, I usually use very specialized models (for programming, for example) and rarely end up generating music or a video on the side. So I can’t think of a reason to use the system right now.

    What target audience do you have in mind?

Explore Birbla archives