Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- I'm working on generating slides with our own model, and one thing general LLMs can't do is leave empty space.
Whitespace is the core of design that actually feels good, but they keep trying to add "distinctive design elements," and you end up with that AI-flavored excess everywhere.
They only know how to add. What LLMs seem unusually bad at is taking things away.
by Swankivo - Video game tips. Constant mistakes and hallucinations, in my experience. Seen this across a lot of different games. Even in really well documented games, such as OSRS (which has multiple fantastic wikis).
Anno 1800 was a recent one I had trouble with, using Claude Opus. Completely made up game mechanics. Rainbow Six Siege, too.
by sandcat_ - LLMs are bad at not inventing stuff (hallucinating facts, sources etc), they're also bad at not over explaining, remembering details reliably, asking the right question and avoiding repetition.by tartoran
- They aren't funny. The jokes they come up with are extremely lame and the sort of thing you would expect a company HR manager to tweet.
I asked a bot why it thought it wasn't funny once, and it told me it has been trained to avoid being misinterpreted or offensive, so anything that might be considered edgy would have been RLHF'd out of it. I thought this was very introspective.
by elliotto - They don't generate keyword search queries very well. They can overcome this by brute force but if you watch what they search you will cringe.
nhl toronto scores nhl hockey toronto scores "nhl hockey" toronto score today nhl "hockey score toronto" "hockey" who won toronto
etc.
Somehow being good at semantic search makes them bad at keyword search, for whatever reason.
by ghostpepper - Identifying bird species from photos that I upload.
To be clear, it depends on what your definition of "insanely bad" is.
I'd say ChatGPT/Gemini make egregious mistakes on ~10% of my photo uploads.
I recently uploaded a photo of a short-billed dowitcher and ChatGPT told me that it was a Wilson's snipe, explaining all sorts of details about the legs and tail feathers (neither of which were visible in my pic!).
I then followed up explaining that a Wilson's snipe hadn't been seen at my location since last November (and that Wilson's snipe was out of season at my location) and Chat revised its estimate downward to 85% Wilson's snipe.
Again, I followed up and I revealed the precise location of the bird and ChatGPT said something like "oh yeah, 99% short-billed dowitcher"!
I've had similar experiences w/ Gemini (haven't tested Claude).
Again, 90% success rate is pretty good, but the other 10% of the time, the 2 LLMs that I use fail on species ID and often hallucinate features on bird photos.
edits for typos, plus another example from the same "birding outing" the other day.
I uploaded a very clear photo of a sparrow.
* ChatGPT says "song sparrow"
* I explain, "no way. this sparrow has yellow over its eye and the breast is wrong for song sparrow."
* ChatGPT: Oh yeah, savannah sparrow
* I explain, beak is too big for savannah sparrow.
* ChatGPT: Oh yeah, saltmarsh sparrow.
* I expalin, "no orange on the bird's face."
* ChatGPT: oh yeah, seaside sparrow (finally correct!)
by busyant - One unexpected discovery that I have made while building an AI-based system: the LLM's are bad at designing prompts.
We tend to think that the AI has some sort of self-knowledge and should be good at designing prompts for itself but it's really not.
Been struggling with a task that heavily depended on prompts, ended up rewriting all my prompts from scratch in my own words, and it finally worked. Then every time I ask Claude to fix something in the prompts, it invariably makes it worse.
A very strange phenomenon that can probably be explained by the quality of prompt design advice that made it to the training dataset. Bottomline, all the prompt design advice that you can find on the internet is really not great.
by mojuba - Serious answer: no model ever gets close to writing an architectural floor plan that makes sense.
They understand all the rules and best practices, they can (sometimes) spot a bad idea in a floor plan, they can describe a good floor plan.
But ask them to make one, even if you give it every detail (even a "node graph" of rooms), they will still output nonsense. Same for text and image models.
Floor plans should be the new Pelican Benchmark.
by jampa