Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • By next month the competition for TTS will be even more!

    Voice models are not winner take all market unlike LLM APIs

    Coming here as Developer Relations at AssemblyAI

  • > and Qwen3-ASR

    Is the ASR inference engine open source as well?

  • All TTS generations are too fast. It's almost I'm listening to a podcast on 1.25-1.5x speed.
  • This is awesome. Thanks for pushing the audio pareto frontier forward.

    Probably far fetched for now, but I think the next big evolution is building the pareto/much cheaper alternative to GPT-Live-1.

    The STT/TTS market is quite saturated, while today, there's almost no cheap/open source alternative to GPT-Live-1.

  • For some reason it switched voices half way through a 33 second clip.

    For OP the clip name is nari-nina-01a0a12f-980a-765e-8029-fa56bd23210d.wav

  • https://apimade.com/audio-compare.html

    Added it to my blind TTS model comparison leaderboard. So far Darwin TTS is the open model leading the pack, ElevenLabs is at the lead.

  • This is really cool work! I'm curious like what do you see as the biggest lever for speeding up TTS models or from a technical perspective that this was a promising direction in the first place to push on. If I were to guess, some distillation but I'm certain there are probably TTS model aware architectural changes that just make inference wayyyy faster?
  • Question: How do you plan to differentiate, because there are so many TTS and its constantly changing every month who would become better

Explore Birbla archives