Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Even if really great, it still can’t use tools. The only consumer Google thing with access to MCP tools is Gemini Spark and that has lots of other problems. I wish they would combine their efforts on a great consumer product but it’s Google we’re talking about…

    I still dream of the day that we get full tool parity in voice and text mode so your voice assistant can do everything you connect for you. Grok and Claude are btw almost there, only a very few minor built-in tools don’t exist in voice mode, but I already use both to connect to heaps of things! It’s so valuable to verbally discuss something with an agent, have the agent pull in context from GitHub, Notion, email, and then create artifacts somewhere

  • Our company's Google Workspace Business only offers 3.6 flash & thinking in the Gemini App. Has anyone else seen 3.7 or 3.8 roll out?
  • I have access to both, benchmarks are actually better on 3.7 for my task, but happy improvement over the others.
  • I'm still only seeing 3.6 Flash / 3.6 Thinking in my Google Workspace for Education account, and 3.5 Flash-Lite / 3.6 Thinking in my "Plus" plan Gmail account.
  • 3.6 Flash and 3.1 Pro are included in the basic Workspace subscription. The Workspace admin has to upgrade your seat for the access to newer models ($17/mo now, $24/mo starting Jan 2027).
  • Google takes a long time to roll out models in general. Both my personal and work accounts still have only 3.6 Flash, even though 3.8 was released 2 weeks ago and 3.7 more than a month ago.
  • I'm wondering if Google intends to drop the next major version of Gemini Pro as a total bombshell drop to make Anthropic and OpenAI panic. They seem to be taking their sweet time on frontier model updates.
  • They promised 3.5 Pro at their next or i/o event, but the rumor is that is never going to be released because it would have been embarrassing. They just started letting their engineers use Claude, so it sounds like things may not be going so well with Gemini
  • Google need to allow saving history and exclude it as training data. I will not use it seriously until this is resolved.
  • Agreed. They really are the greediest when it comes to data for training (unsurprising for google I guess)
  • Gemini's Live Mode is already much better than GPT Voice in my personal experience, even though it was much dumber. It really does feel like talking to a real person. ChatGPT keeps humming to whatever I say and has some weird voices.

    Excited to try this out! Shame on Google for not releasing Gemini 3.8 for Google AI Plus users yet, though.

  • OpenAI just released the new full duplex mode to the API as gpt-live-1 or something like that. Very realistic.
  • Not a great impression to have your demo video demonstrate how one of your 'most advanced' AI models loses to the most common check-mate pattern in all of chess.
  • Ah, feel the AGI man!
  • agree that it’s a weird choice, but more because i don’t need my chat model to play chess at all when chess engines exist.
  • I am an adult, and i lose to elemental school kids in chess.

    I am not so advanced enough as a human being.

  • Seems more than good enough for a live model though! I can imagine this demo being extended to be a lot nicer to play with. You can just feed the model engine analysis and it can make as high of quality moves as needed. No longer any correlation between the model's understanding of the position and the moves that would be made but I think that's still a really nice improvement when thinking about this as adding live voice interaction to existing chess vs computer functionality rather than adding chess to possible interactions with the latest live voice model.
  • Playing a legal game of chess without using a guided decoding technique is a massive achievement. Ref https://aclanthology.org/2025.mathnlp-main.11/.
  • For an LLM, just being able to play an entire game of chess without illegal moves and without inventing pieces that aren't on the board is an achievement. Even more so for a live model. Then again, who knows how much the harness helped here
  • The voices sound really pleasant and realistic. Sadly it doesn't support SIP. I found forwarding streams with websocket for phone calls tend to introduce some unwanted latency that's quite noticeable in a conversation. I've found a lot of success with GPT-live-1 so far, but the generated voices are lacking something I can't put my finger on.
  • What's your use-case for automating phone calls?
  • Gemini is underrated in that it produces the only prose that is somewhat bearable to read.
  • I find its style the most sycophantic and annoying personally.
  • Chatgpt doesn’t seem so bad lately. At least as a Claude refugee.
  • I can’t agree more! It feels really smooth to drive, kind of buttery compared to other frontier experiences, at least in antigravity 2.0 or whatever. I’ve been doing some web app coding with it and I’m happy with the results.
  • Try Gemini live in a multi lingual environment. It can pick out speakers and live translate to you. Truly underrated for its capabilities.
  • I work for Google and we have the choice between Gemini models and Opus. Opus is slightly better than Gemini flash but I find the style unbearable.

    It reminds me of a pedantic grad student.

  • I was surprised when (finally) trying out Claude how much I preferred Gemini's way of communicating. I wont argue Claude is better at coding, but for knowledge work, I had to dig through Claude output to find what I actually wanted. At times, it even felt borderline incomprehensible.
  • For heavyweight work I have been using Astra, but for rabbit holes and brain storming Gemini is far more enjoyable to interact with.

    I'm worried in their push to catch up on the SOTA front, it's going to lose that natural sounding touch it currently has.

  • It's also the only model that generates accurate translation and localization. No other frontier model comes close. Although Gemini's coding capabilities are subpar, its natural language processing is top-tier.