Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • So Seedance is good primarily because of TikTok and this because of YouTube. I wonder what portion of all recorded video is privately held in hard drives at people’s homes or Apple photos. Of course there is data labeling and cleaning but is the next evolution just a question of access? Same goes for LLMs. Would people be willing to sell their data? Kind of a messed up way to make yourself obsolete. Or there is a limit to scaling?
  • Draft videos more efficiently in 360p

    While it sounds great you're quickly disappointed after you run the same prompt at standard resolution only to get a different result because it's non deterministic.

  • Can you hold the seed constant?
  • It certainly makes for easy demos, but I always struggle with the practical application. As in, what work or enjoyment does someone actually get from this? Ads and media pre production seem plausible, but it fails the 'how can this enrich life' in a way most other AI tools don't. Maybe for them that's not a consideration, if their only interest is the other meaning of enrich that might flow from ads and numbing rivers of slop.

    Why do we look at art, watch videos/movies? Is that replicable as a function of text, other existing media, and 3-30 cents of compute per second? I'm pretty functionalist about these things, and at some point it probably won't be possible to tell the difference. But until then, at which point we might just say 'death of the author', it seems like a category error.

    I do work with artists that use video and image generation models to create stuff, but from what I can tell they're interested in faster iteration and controlling a lot of intermediate steps (their graphs can get pretty labyrinthine).

  • Google's main source of revenue is advertising not enriching people's lives. It's not a charity.
  • Yes, most people using AI for creative projects spend a lot of time and attention mastering their tools, and figure out how to adapt them into their creative processes. AI can dramatically lower the cost of indie productions, while also a allowing a broader range of stories to be told. Even the most successful film makers need to bow and scrape to get their projects funded, democratizing visual media can be a good thing, even if you, personally, are no more likely to do this than you are to pick up Photoshop or record a podcast.

    I enjoy making short films with AI. When my latest short screens at a festival in Ocotber, alongside traditional and AI films, hopefully the audience will like it too.

    The quick "one shot" video generation might be slop to you, or I. But if someone wants to send it as birthday greeting to their aunt, and they both enjoy it, what business is it of ours?

  • > As in, what work or enjoyment does someone actually get from this? Why do we look at art, watch videos/movies?

    I like to generate songs from obscure poems

  • > As in, what work or enjoyment does someone actually get from this?

    I have been pondering this and have come out with a single statement. It allows you to create the missing piece in the art you want to create. Assets for games, music for the lyrics you wrote, or just a whole song to justify a crazy dance you want to do. Whenever AI art is discussed, people tend to be so purist about art. One person cannot do everything in a project they take on to express themselves. Usually people come back with then get someone else to do it for you. I think that is an economic argument as people are trying to protect artist pay. I get that. But is it fair to just not let something they want to create because they don't posses every talent needed to do this?

  • Kids love image and video generation! The former is cheap enough to do just because it is fun.

    I have a young boy, and whenever he builds an impressive "scene" from LEGO (like a diorama or whatever), I take a couple of reference pictures with my phone and make it into a "real" movie scene, cartoon, or whatever. He loves it, and this motivates him to build more and bigger things out of LEGO.

    If he builds something really special, I might actually fork over the $5 to use Omni to turn his LEGO creation into a 10-second video instead of a still image. It'll blow his mind!

    PS: There also are cheap and even free phone apps that make stop-motion animation trivial. We've already made a couple of videos of his toys moving around that way.

  • Before AI, cool and interesting shots carried the promise that it was reality - even if it was perhaps exaggerated.

    The amazing videos and photos carried the promise that I could experience that for real. They were aspirational.

    Today I suspect every cool shot is made of pixels arranged on a 2D screen by an algorithm. It doesn't do it for me.

    That makes me sad...

  • >"Draft videos more efficiently in 360p

    Generate lightweight previews in 360p resolution up to 60% faster

    and at a third of the cost compared to Omni 1.1’s standard 720p resolution. This is helpful for rapid prototyping, storyboard iteration, and quick rendering in developer platforms."

    This is a great idea, to have a low-resolution mode for additional speed to create previews, do test runs, create rapid prototypes, etc.

    My curiousity is, what's the absolute useable minimum that this could be?

    That is, would/could 240p resolution work? If so, what about 144p? How about lower? Then, could those images be upscaled quickly (and is the result still usable?) with a faster image upscaling-only neural network?

    The reason why knowing such lower numbers / lower bounds -- is because they could be important for additional cost/time savings and/or running derived LLM's on local resource-constrained hardware...

    Anyway, great post, great idea, and we welcome Gemnini Omni 1.1 Flash to the ever-expanding list of LLM/AI's!

  • Surprised they don't have something in between a video generation and image generation model (literally like storyboarding) which still took into account physics and world knowledge. You could probably (relatively cheaply) generate decently high resolution storyboards and then have something to directly feed a video model at the end.
  • I let myself get mildly excited with the last Omni release, but it turns out it (and this one) can't do the one practical thing I want - Sync generated video to provided pre-existing audio.

    Meanwhile, I'm happily using Minimax H3 locally on my 12Gb 4070RTX to finally finish the lip syncing to recorded dialog on my abandoned 20 year old Flash animation hobby projects.

  • Pretty sure I heard one of the PMs in a podcast a few weeks ago say they are intentionally not building support for it out of concerns of enabling deep-fakes.

    I agree though. My issue is the cost for using AI video models is way too high for anyone not building anything serious with them, at the same time they are too restricted for actually using professionally. Prompting them with text to get something generated is cute, but then you just end up creating slop that everyone hates, ultimately devaluing the power of these things.

  • Yeah, that'd be a great feature.
  • > Minimax H3 locally on my 12Gb 4070RTX

    Minimax H3 is about 240Gb alone, how do you do? How much quantised is it, and how good are the results?

  • Anyone else becoming numb to these updates?

    I feel like I should be excited about being able to generate almost perfect videos but, I just don't care anymore.

  • Google does anything except launch a new version of Gemini Pro.
  • New version of Gemini Pro has now become like GTA-6
  • Why do they need to? For search, instant models are more important and fit the use case better.

    Pro models are mainly for coding agent work; it doesn't necessarily make them any money.

  • I have a pro subscription, I think they have just given up. Likely because when they test their new models against the other frontier models they are so bad, they just pull it back. This leads them to try and innovate in other areas where there is currently less competition so they can compete. Not a bad play.
  • Just because Anthropic and OpenAI really want there to be an arms race justifying the outsized investment, doesn't mean the optimal play is to build larger, more expensive, models.

    The capital infusion the frontier labs have received has gotten to a size where many believe it may not be possible to recoup this investment without some very unrealistic things happening.

    I think it's reasonable to not completely drain one's cash reserves trying to stay ahead in a race where participants may very clearly be about to run straight off of a cliff.

  • > What stands out about Gemini Omni Flash is its accuracy: the details hold up under scrutiny.

    Quote under the video of a short Argentinian footballer wearing no 10 with "RESSC" on his back. Can't make it up.

  • Interesting that OpenAI abandoned Sora entirely but Google are continuing to invest heavily in their own video generation.

    Maybe because they see video generation as key to developing "world models"?

  • Also Google already has a huge built in training corpus with YouTube, gphotos, and geospatial data
  • They're main edge has been multimodal. I think they're still the best overall on multimodal? If I were them I would try to be the best at at least something.

    I never personally got the least bit excited about Sora or nanobanana or whatever video/audio generation thing. But I guess I'm just not their customer. I do love the read-side of it though.

  • This is paid API access only. They are here to make money not to position themselves for an IPO. Not a value judgement only an observation.
  • 10% of their revenue comes from YouTube, so they need to make sure that they own any technology that might disrupt that platform.
  • YouTube. And video ads.

    Previously when making video ads you'd need to actually create the video. Actors, cameramen, editors - you name it. Now a new video ads is just a prompt away, directly inside the ad-spend web UI too no doubt.

    People say Google have lost and that they're having their lunch eaten by anthropic, but I am not so sure...

  • Sora was a social network type thing. Google sells their models on a PAYG basis - and makes money off them. Nano Banana alone has changed advertising 2D mockups and Photoshop like tasks forever. Notice how GPT image 2 is now available also on a PAYG basis.

    I work in advertising and some days I spent a lot of money using these models. The amount and rapidity of prototyping using them has changed everything about advertising pre production.