Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • This is cool
  • Can you change the voice? Also, is it possible to vary the emotion in one audio?
  • This is cool! What model are you using?
  • Thanks! Airy runs on a proprietary TTS model that we built in-house, rather than a third-party model.
  • Great, Working fast and cool. You can also try to adding some more voices from different parts of the world.
  • You know, I never did find out why people dub anime with such unnatural voices.
  • The voice actors usually try to match the timing of the mouth flaps in the animation. It's hard to do that while also conveying the meaning of what is said correctly (enough) -- and practically impossible to do that while sounding natural in English too.
    by Figs
  • Why did you pick such creepy voices? They are very childlike and weird.
  • They just sound like anime voices
  • They're definitely anime-style voices, though most of the female ones are very annoying. I couldn't stand to watch an anime with them.
  • As with seemingly all AI these days - it seems to fail with prosody and simply speaks the very next word with zero regard to the nuance or cadence that author intended, or an understanding of any the words being spoken.

    When AI achieves the ability to deliver some of Shakespeare’s greatest soliloquies or monologues, then I’ll pay attention.

  • Prosody is hard. No AI voice I've heard really nails voice generation with fully smooth and human-like cadence and quality, but they're gradually getting better. Airy voices sound a bit tinny and childish, but even so, they're better than many I've heard, including for the elusive "humanness" quality.
  • All but one of the voices sounds tinny. The timing and cadence was good though.

    Is this intended for anime autodubbing? There's only one voice (Rowan) that is remotely traditional broadcaster-style.