Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- > I also like this kind of examples because no one in their right mind would ever spend the time to write something this custom
There are people in their right mind who would do that and their are already examples of people who did similar things.
But maybe not in the future if people would confuse all the effort with AI
by croes - this is not a good benchmark for models, but it's great if you're optimizing for attention on twitter because video content and 3d animations perform best on social media.
a real benchmark is instead running evals on your own traces, and building a cost/quality/speed profile for models based on real workloads. but it doesn't get you a shiny video you can post on twitter.
by try-working - You can do both.by sumedh
- I don't think it's a bad way to benchmark new models, I just find it concerning that the author implies that "pelican on a bicycle" has been exhausted. At the risk of making overly broad, unfalsifiable claims I think multi-year exposure to AI content has dramatically raised our expectations for speed and volume but lowered them for quality. We see a very janky pelican and declare the problem solved.by YmiYugy
- Humans are drawing pelicans riding bicycles now. Just google it and you will find 5 or 10 of them in the first few results. Including a t-shirt design.
So it's a pretty much pointless test now.
by mattmanser - 100%. Why waste the tokens to render Lord of the Rings when the pelican test still clearly benchmarks so well.
- Can someone explain what the pelican on a bicycle tests exactly? And why is it so important? I've never understood how it could translate to a useful task in real life.
- Agreed. It's far from solved. Modern LLMs still generate pelican bike SVGs with obvious errors:
* some omitted the bottom of the diamond which connects from the pedals to the rear wheel
* some added an extra connection from the pedals to the front wheel, making it impossible to steer
* none could align the head tube with the fork
* none added a correct offset to the fork
* none could generate the chain properly in a way that attaches to the two sprockets correctly
I mean just look at these:
* Grok 4.5: https://s3.eu-west-1.amazonaws.com/images.dylancastillo.co/p...
* GPT 5.6 Terra: https://s3.eu-west-1.amazonaws.com/images.dylancastillo.co/p...
* Sonnet 5: https://s3.eu-west-1.amazonaws.com/images.dylancastillo.co/p...
by dllu - Also, one of the advantages of the pelican test is that you can evaluate it all at once. There’s no reason the two-dimensional depiction can’t be made more challenging. Yes, at some point the pelicans might approach the subjectivity of a fine art painting, but we haven’t even seen a depiction that’s competent by the standards of a high school art class. That’s not to say the elementary school–level SVGs aren’t amazing - rather, I agree with your point.by Demiurge
- I think it's touching the limit of what one can reasonably expect any intelligent thing to produce with the only direction being "produce an svg of a pelican riding a bicycle".
When you aren't sure if an LLM can write an svg well, or that it will be able to form a pelican shape, or animate a bicycle, it's a good test. After that, it's all judgement: how detailed should the pelican be? pelicans are the wrong shape for a bicycle by default, so how much can I change its physiology to match using a bicycle before it isn't a pelican? Do I care about how well the client is able to render complex geometry?
It's not that there isn't room to do better, or that it doesn't tell you anything at all, but rather we've reached a point where what it tells us isn't very clear anymore.
by BobbyJo - I don't think he's claiming it's been exhausted. It's just that things have progressed to a point where people are arguing over the finer points of which pelican looks better -- which is often a matter of taste, and an indication that we've hit the knee in benchmark where models are no longer failing in obviously awful ways.by jonas21
- The difference between this and Simon's pelican is that with Simon, I get the prompt.
Last I checked, I did not see the prompt for this really cool thing, so it is not reproducible.
Did I miss the prompt somewhere?
by consumer451 - he said the prompt was the first paragraph of LoTR, but he didn't mention a preamble
this guy seems to have taken that idea and got something similar/better, so likely the prompt isn't too special
by vanjajaja1 - I always thought of the Pelican more of like a gimmicky quick test. There are people who took it as a serious benchmark for overall model performance?by trentor
- "Draw a pelican on a bicycle" is not a serious benchmark.
"Draw an animation of this long ass scene from a movie, and only call me when everything works e2e" can be.
- Out of curiosity, I asked Opus 5 to do the same with the first ~1.5 pages of Neuromancer by William Gibson, which I thought might be a good test of interpretation from the model.
It refused to use the text verbatim because of copyright (ironic), but the output was interesting nonetheless.
https://claude.ai/public/artifacts/275dc3c2-7bd3-432b-94ff-d...
by weakfish - Thanks. Your result is what I would expect - nice stuff, but far from what he supposedly "casually" generated.by neop1x
- it would be interesting to see results from works that don't have notable existing film/tv adaptations or illustrations (though I suppose there will always be fan art) to see if it’s really able to generate something novel. I’ve had my own version of this since I’ve been reading LotR recently too and I kinda wish I had read the book before watching the movie (not that the movies are bad though). Fortunately I am finding there is tons more content in the book than the movie :)
Also, large models refusing to work with copyright material is really hypocritical, copyright enforcement for thee but not for me
by jedbrooke - A simple prompt that still stumps frontier LLMs most of the time is “create a pinball game”. They’ll put all the right pieces there but then fail to arrange them such that the game is truly playable. They’ll put a wall in the way of the launch chute so the ball can’t be launched. Or the flippers will pivot the wrong way. Or there will be holes such that the ball drops off the bottom without getting within reach of the flippers, etc.
Opus 5 is the first I’ve seen to “one shot” it (in a harness, so it was more than one LLM call).
by darrinm - We don't use these models by way of "one shot" so I don't see why it's relevant. It's clearly useful to let them iterate.by kzrdude
- Failed demos like this give weight to the argument that AIs need more of a world model, an understanding of how physics works to avoid obvious stumbles like this.by Schlagbohrer
- I can forgive the modeling being godawful jank (windows floating in the air, disconnected from the house). But I expected it to have a better understanding of the text. Instead, we have Bilbo's "disappearance" interpreted as him magically transporting or cloaking, and similarly for his reappearance.by dundarious
- which to be fair, happens just a few paragraphs later :) I first watched without sound and thought "oh that's the birthday speech disappearance"by kiwibyproxy
- But this only seems wrong to you because you're familiar with the prior context. When there's only a single paragraph to work from, and it's a drily humorous text, why not employ comic literalism and lean into the perplexity with which his neighbors viewed him?by anigbrowl
- I'd rather have them battle on the topic "Who builds a better Google Wave for LLM chats" to explore the space of how AI studios could be.
"Getting Started with Google Wave": https://www.youtube.com/watch?v=eKUAqNGVwX0
by qwertox - i remember this failing, but this product looks pretty useful in 2026by misiti3780
- Also see Game Helpin' Squad's "A Pissed Off Tutorial For Google Wave":
https://www.youtube.com/watch?v=4Z4RKRLaSug
Their "World Quester 2" tutorial shines as the holy grail of consistent and ergonomic user interface and game design. The menuing system is so magnificently structured and well organized, it bring tears to my eyes. Google Wave pales in comparison.
by DonHopkins - It’s a fun idea to re-animate dead google products by feeding product videos to an AIby baxtr
- It seems pretty clear that Anthropic models have been specifically trained to be good at generating three.js (JavaScript 3-D Graphics) code, so given current state of AI code generation in general, I don't find three.js models/animations as indicative of anything other than the model's ability to write three.js code.
When Fable was first released the day-1 demos of it on Twitter (presumably from people who were given early access, and/or Anthropic employees) were pretty much 100% three.js stuff. Yes, it looks nice, but it doesn't tell me any better than an Erdos proof whether the LLM will be able to run my vending machine.
- Don’t worry. I’m sure they’re not training it to be good at things like writing database backends, financial services, logistics systems, user interfaces, or anything of economic value. As long as you’re not working on three.js specifically, I’m sure Anthropic isn’t making any progress you should be worried about.