Discussion summary

The discussion covers GPT‑Live's real-time translation capabilities, concerns about energy use and ethical issues, and the potential for local voice model deployment.

What the discussion says

  • Some users praise the translation quality, while others criticize it as still imperfect.
  • Concerns about energy demands and ethical implications are raised.
  • Interest in running voice models locally is expressed.
“With this, human translators have been totally and absolutely a solved problem.”
— rvz
“watched the live translation video very impressive.”
— zuzululu

Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Gemini live has been able to do this for over a year now. I can just activate it on my phone and it really works surprisingly well, especially the interruption. I've tested it with my 95 year old Dutch grandmother and it switched seamlessly between English and Dutch with her and handled her poor hearing very well, including her asking for repetition.

    I'm a little surprised by how much OAI is playing catch up here.

  • Much better than it was before but it’s still significantly weaker than a direct chat.

    For example I asked

    “Why should LLM attention use dot product instead of cosine similarity, being that we often hear vector magnitude does not encode most of the useful information needed”?

    The voice response was directionally right but lacked detail and was a little hand wavy.

    The answer to the same question in a text chat was much higher quality.

    The voice response replied “let me think about that…” so it appears to be invoking 5.5 as advertised, but it’s definitely weaker.

    I had reasoning set the same for both.

  • (Atty from OpenAI here)

    GPT-Live-1 is the first version of a new generation of models, and we believe the full-duplex architecture + delegation enables entirely new ways of human-AI interaction.

    Would love to hear your feedback!

  • Once this gets video capabilities and is ported to glasses, it'll be a major revolution for blind people (and I say this as a blind person).

    People have tried "smart <thing> that helps blind people navigate" since the 80s, many, many, many times, and all such projects failed. The cycle of "wow, blind people could benefit from a navigation aid, why don't I make one, if there's none around, I must surely have been the first bright university student to think of this idea" is pretty well known in the community, and I'm personally quite tired of it. Nevertheless, I think this may be the one.

    Circa 2020, I have said that people who are getting a guide dog now are probably getting their last one. I think we aren't far off from that prediction coming true.

  • Like a lot of AI things, this seems both cool and kind of creeps me out. I've never used voice interfaces in the past (siri, the google one, whatever is on my tv) so I'm probably not the target market, but this does seem like an improvement.

    The part that creeps me out is, we're living in an era where we're more disconnected from each other than ever before. Do we really need to be replacing conversations?! The demonstration video of old ladies sort of hints at something for me, which I think we already have a societal problem with the way we treat the elderly (and a massive elderly-loneliness issue) and there's kind of a sadness of imagining people becoming really close with this machine that doesn't really think. Definite ick factor.

  • What I’m missing from this announcement is the capability to use connectors and tools. I don’t really get it - NONE of the frontier assistants can use tools / connectors while in voice mode - Claude, ChatGPT, Gemini, Grok. It seems so obvious: I want to be able to research stuff, pull up documents, jot down notes and do productive work while I’m talking to it, and not end voice mode whenever I need to connect to an app or service.

    It’s weird. The old Claude voice mode WAS able to use tools but when they revamped it, it lost that capability and is now pinned to Haiku :(

    So, yay for finally a voice mode that’s powered by a frontier model and hopefully as good as Grok voice, but sad to still not see tool use while in voice mode.

    (I haven’t tried it yet, only read the announcement)

  • This is the opposite direction AI should be going. Human relationships are the most valuable thing we have, and so, naturally, technology seeks to intermediate and now replace them.

    I'm not Catholic, but this podcast presents a very interesting argument against talking to AI as if they were human: https://newpolity.com/podcasts-hub/debate-chatbots

  • I had preview access to this one for a few weeks. It's very good. I had one conversation that lasted a full hour while I was walking the dog, got some good brainstorming done against one of my projects.

    The best feature is that it can delegate questions out to GPT-5.5 in the background, so you're no longer restricted to a voice model that's several years behind the frontier.

    I did report a fun bug with it though: it was interrupting me and laughing at my (not really intended as) jokes while I was still talking! They seem to have clamped that behavior down thankfully, it felt a bit rude and condescending.

Explore Birbla archives