Launch HN: Speko (YC S26) – OpenRouter for Voice AI

Launch HN: Speko (YC S26) – OpenRouter for Voice AI

13 pointsby abdik2 comments

Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • I made a completely free 100% on device translation app https://apps.apple.com/us/app/arda-translate/id6778970560 and hard to image a world where TTS and STT will not be done locally in the future
  • This looks really interesting.

    I feel like there is a lot of room to build great voice-based agents that don't exist right now.

    I have found that ChatGPT voice mode is unusable (e.g. hallucinates me saying things); Claude voice mode is usable, but very buggy around tool calling, and it often mishears things. And it only supports Opus, not Fable (though it looks like you don't support either of those). But I use it anyway.

    Question, do any of your TTS options support increasing the speaking speed?

  • Answering your question from the small end: I picked my STT by testing self-correction handling. My tool cleans up spoken drafts, and the failure that mattered wasn't word accuracy, it was "meet Tuesday, no wait, Wednesday": a raw transcript of that is worse than useless, and models differ a lot in how gracefully downstream cleanup can recover. Your spontaneous-speech testing sounds close to this already. Do the boards score corrections and disfluencies specifically, or do they fold into overall accuracy?
  • Does this include a turn taking API? It'd be great to have one API that could do "Conversation in a box". One of the biggest annoyances is daisy chaining many models together for turn taking, dumb models for immediate responses, with smarter models returning and taking over after.
  • The benchmarks page seems interesting and something I can use to help make an informed decision. Can you talk about how you're measuring some of these? I imagine it needs to involve some human input.

    https://benchmarks.speko.ai/turntaking

    by webo
  • Routing makes even more sense for voice than text, but the constraint space is trickier: latency budget (round-trip vs streaming), WER on accented speech, and TTS naturalness all trade off non-linearly. In our voice pipeline, DeepSeek-V4-Pro plus a small dedicated STT beat an end-to-end frontier voice model on cost-per-minute by ~10x while staying inside a 300ms added-latency budget - purely because we could mix and match components. The hard part is honest benchmarking: WER numbers are only comparable within the same eval set, so a "find the optimal combo" service lives or dies by its methodology. Do you expose per-component benchmarks (STT WER, TTS MOS) separately, or only the combined scores?
  • The piece I'd want to see in the constraint solver is effective cost rather than list price. On the LLM leg alone, current 5.6 pricing splits at 272K input tokens with the long band at exactly 2x, and cached input runs 90% below standard input, so two stacks with identical list prices can differ several-fold depending on how much of the prompt is a stable prefix.

    Does the optimizer model cache-hit rate and context distribution, or does it score on list rates?

  • Two things bit me when I was choosing a voice stack, and I don't see either as an axis in your benchmarks.

    First: whether a model is "suitable for realtime" turned out to be a property of the transport, not of the model. I had written one TTS off as too slow based on the vendor's own guidance, then measured it again over a streaming path and got about 0.9s where the non-streaming call had taken 5s. Same model, same provider. If a benchmark is run over one transport, a model can look disqualified when it's actually fine for the way you'd ship it.

    Second: non-English breakage doesn't show up in aggregate quality numbers. The speed-optimised tiers - the flash/turbo class - were fine in English and fell apart in Japanese. Not "slightly worse": confidently wrong words that changed what the sentence meant. What made that expensive is that the vendor's own docs said their turbo tier was equivalent to their flash tier. That's true in English and wasn't true in the language I was shipping in, so the documentation actively pointed me the wrong way.

    Both of those mean the thing I'd actually pay a routing layer for is per-language and per-transport measurements, rather than one quality/latency/cost point per model. "Best TTS under 300ms" has a different answer in English than in Japanese, and I couldn't get that out of any vendor's published numbers.

    So: are your benchmarks per-language, and do you measure over the transport people actually ship on? If yes, that's a bigger deal than the routing itself for anyone shipping outside English.

Explore Birbla archives