Discussion summary

A developer created a tool enabling LLMs like Claude to 'see' videos, addressing limitations of existing models. The project has received positive feedback and comparisons to other solutions like Gemini.

What the discussion says

  • Some users find the tool innovative and useful.
  • Concerns about cost and practicality are raised.
  • Comparisons are made with other models like Gemini.
Claude won't accept video files, ChatGPT reads transcripts only.
cortexosmain
I gave Claude a video for a speeding ticket, and it was spot on.
garciasn

Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • I think this is much more useful than just LLM related applications. I'd suggest renaming it to not make it seem like it's LLM related.
  • Cool idea, but keyframes are not videos. Motion, object permanence, are not things Claude can infer from a set of images. Nice demo though!
  • I have been going through this with claude and qwenvl3:8b this week. Both are pretty decent at inferring context and analyzing contact sheets. Finding high visual interest moments with a mixture of coarse and fine keyframes.
  • Exactly! We experimented with a whole bunch of video encoding techniques for LLMs here: https://vlm-run.github.io/mm/encoders/#video
  • So I did this yesterday for a video analysis sample with ChatGPT and it took the video, pulled out frames, did difference tests across the frames to look for significant frames to focus on, did image recognition on each frame, and interpolated motion and action between.

    So I’m not sure why this says ChatGPT doesn’t “see” video and reads transcripts. Obviously if the video is already labeled that’s the shortcut. But it did an impressive job describing a video I have no inclination it would have in its training data. One could argue it wasn’t “native” and had an agent orchestrator to rely on external tools to accomplish the goal… but it worked.

  • Had the same experience with Claude, just somehow the entire thing felt (token) expensive.
    by frb
  • I was just thinking about this exact use case yesterday:

    And it's for me measuring different charged speeds at different starting battery capacities and different temperatures and I was like well. What if I just had a video camera pointing at the voltage going in and out and then I could see the battery percentage increase and I can have a temperature gun pointed at the phone as well. And I couldn't know what temperature of the phone is as well and it could just figure it all out create charts..

    This would make reviewing different charging equipment really easy as long as you really have to do is plug it in and tell other people to do the same thing and take a video of it and beat it to the system.

    I might very well give this a try!

  • It's kind of wild how much we are abandoning basic problem solving skills in favor of just pointing an enormous stack of GPUs at it
  • Are models any good at descerning motion from multiple frames?

    For instance if I gave models multiple animations of a bouncing ball as individual frames. Would they be able to tell which bounce was the more realistic motion.

    (Is this a potential new benchmark? maybe also variations of stair dismount)

    by Lerc
  • I’d imagine they could. I’d try Gemini 3.5 flash with high fps.
  • The comments in this post strongly validate the need for reliable video processing and understanding with VLMs.

    While you can use Gemini or other local VLMs, the real challenge is token efficiency, accuracy, and coverage. For example, how do you make a VLM “watch” a 2-hour or 4GB video without losing context or meaning?

    Video transcript alone can be sufficient for basic workflows needing no visual context. But when deep contextual understanding is required, e.g., self-driving, security analysis, warehouse tracking, etc., you’ll need more advanced methods like keyframe sampling, clipping, chunking, and shots+transcript.

    You can explore the different encoding strategies we designed for efficient video processing and understanding here: https://vlm-run.github.io/mm/encoders/#video.

    FYI, the repo is now public, and contributions are welcome.

  • Nice @OP i put together something similar as well. Incidentally I found for motion design specifically llm is not able to infer specific animations as well as it just being described very plainly and accurately what is happening and the timing.

    One thing which sort of worked decently was actually take the frames and put them into a grid and have the agent look at the image of all of the frames together. It did surprisingly well but missed a lot of subtle details that it couldn’t see.

    Also tried various kinds of vision embeddings, heat map of motion etc, and blur etc to show motion. But none really worked as well so I ended up just describing it until it got it. Haven’t quite found the right solution yet.

  • I was creating a scene by scene remake of a cutscene from an old DOS game. The sprite sheet had several sprites which were cycled (e.g. a horse with it's head down and up). The engine would cycle through these regularly to create some "liveliness" in the background. It was tedious and I didn't want to figure out which sprites belonged at which pixel location.

    I recorded a video of the relevant part of the cutscene using dosbox and then split it into numbered frames using ffmpeg. Then I gave that + the spritesheet to Claude Code and asked it to figure it out and tell me which ones are at what position. I should probably have deduped it but in any case, it churned through the whole thing and got one or two out of 15 or 16 sprites right. The rest, it just dropped into random places. YMMV

  • Ask a coding agent to decode the assets. Works pretty often for such old games.
  • This looks cool but this should be renamed without having Claude in the name.
  • llm-real-video would be a much better name
  • I’m currently punishing Fable by making it watch the entire series of 7th Heaven.
  • Inhumane
  • It's going to make itself unavailable again. Actually... that's probably a litmus test for sentience.
  • "Where the video goes: stays on your machine" - No, the frames (that this tool extracts) obviously get sent to Anthropic if you use Claude.
  • "Or any LLM" on your machine.
    by fny
  • Pretty terribly expensive way to watch a video with Claude.

    Use Gemini or some local VLM to do this way more efficiently. We spent quite a bit of time on video understanding, and Claude will just burn tokens.

    Check out this library: https://vlm-run.github.io/mm/

    You can swap models and try out different encoding methods for videos (https://vlm-run.github.io/mm/encoders/#video)

  • Assuming that's your project, the GitHub link from the PyPi page is a 404.
    by mh-
  • Do you mean that Gemini is most token-efficent at watching videos? Is that the case for e.g. just giving it a video in the browser? I admit, I dont give LLMs videos as I just assume it'll burn too many tokens.
  • Seems cool from the docs page, I was about to give it a shot but https://github.com/vlm-run/mm goes 404 …
  • Exactly this. Gemini is best at this. Just give it video link - YouTube works best - and it will analyse the video.