Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- Suggesting business names for businesses, I mean they are great, but they already exist, multiple times even.by TZubiri
- True, tried it so many times and every time I come up with something myself though sometimes inspired by the AI's ideas.
Verifying trademarks and domain name availability is usually an additional step you need to ask it to perform. Trademark DB searches by the way are intentionally made difficult to scrape so most of the time it's a manual process anyway.
However, once you give it all the information (TM search results, domain name availability) it can help you with the judgement of how safe the name is from the legal perspective. With the obvious caveats, but still a good starting point if you are serious about the name.
by mojuba - Having a spatial understanding from an ASCII map, while doing long term planning. Just try making an AI play nethack or similarby TiccyRobby
- Convert it to an image on the fly to feed it into a vision language model and I expect it would work just fine.by lrvick
- Here is a system I developed for my own projects:
1) hand-write a simple 2d TUI-based rogue-like in Rust using pretty much just the std; 2) grab opus 5.0 (it used to be opus 4.6, 4.7) and give it some vague "requests", and ask it to make this game "production-ready" and "blockbuster", but keep the 2d and TUI aspects so I can actually run it. 3) now the fun part, take a test subject, say GLM 5.3, and ask it to find code smell, architecture issues, duplication and all sort, and *simplify the code*
compare the result to my original version.
It's not a simple thing, but the concept is simple: can an LLM remove all the mud?
The winners so far are (ranked by the quality of the final result, not by token cost)
GPT 5.6 sol (extra high thinking); GLM 5.3; Grok 4.6; Qwan 3.8;
(fable could not make it to the list because it simply cannot follow the instructions)
by drpython - These models seem to be bad at writing prose or text. Many of the sentence structures seem to be unvaried.by kanzure
- If I am relying on the model to do the writing without any context or learning on how I want it to write then yes. However if I build skills that have learnt how to write in the way I want them to then I find they write very well, or at the least how I want them to as opposed to how they do natively.by NoPicklez
- I'm working on generating slides with our own model, and one thing general LLMs can't do is leave empty space.
Whitespace is the core of design that actually feels good, but they keep trying to add "distinctive design elements," and you end up with that AI-flavored excess everywhere.
They only know how to add. What LLMs seem unusually bad at is taking things away.
by Swankivo - LOL: from the RFCs: “In protocol design, perfection has been reached not when there is nothing left to add, but when there is nothing left to take away.” (from https://www.rfc-editor.org/info/rfc1925/ )by lanstin
- Video game tips. Constant mistakes and hallucinations, in my experience. Seen this across a lot of different games. Even in really well documented games, such as OSRS (which has multiple fantastic wikis).
Anno 1800 was a recent one I had trouble with, using Claude Opus. Completely made up game mechanics. Rainbow Six Siege, too.
by sandcat_ - I used ChatGPT on nfs heat and was fineby skeptic_ai
- As a noob, how does the end user improve this? What's the best way to make the knowledge from the specialised wiki available to the LLM?by salamandars
- Probably not benchmarked on games. Doing so might sacrifice quality on other more important benchmarks, and when it's used for legal and medical purposes, it's the right choice.by TZubiri
- I've experienced this also, sometimes I ask it about WoW stuff, e.g tips for arena or which enchant to get and it makes a lot of mistakes in regards to which spells or enchants are available in which phase or expansion. I guess the source material is quite bad.by wiper88
- LLMs are bad at not inventing stuff (hallucinating facts, sources etc), they're also bad at not over explaining, remembering details reliably, asking the right question and avoiding repetition.by tartoran
- They aren't funny. The jokes they come up with are extremely lame and the sort of thing you would expect a company HR manager to tweet.
I asked a bot why it thought it wasn't funny once, and it told me it has been trained to avoid being misinterpreted or offensive, so anything that might be considered edgy would have been RLHF'd out of it. I thought this was very introspective.
by elliotto - You can get better results from less aligned models like Kimi K3. Still not actually funny, but at least it’s able to produce some unhinged stuff and I guess shock and twists are kinda related to humor?
Still missing the human connection of cause, so im not sure if this is a technical / skill issue in the first place.
by nilsherzig - Oh, that we have already figured out. It's because...
https://www.scribd.com/doc/290970915/The-Jokester-by-Isaac-A...
by Yizahi - It makes sense. It goes father than that. LLMs generate probability-based tokens. What is the most likely next word?
Humor goes against what we expect. A punchline works because you don’t see it coming. It’s not funny if you’ve heard that one before.
LLMs are, by design, going to be shitty comedians. They don’t have unique perspectives and their own voice
by mingus88 - They don't generate keyword search queries very well. They can overcome this by brute force but if you watch what they search you will cringe.
nhl toronto scores nhl hockey toronto scores "nhl hockey" toronto score today nhl "hockey score toronto" "hockey" who won toronto
etc.
Somehow being good at semantic search makes them bad at keyword search, for whatever reason.
by ghostpepper - that and always putting the “current year” at the end of the search term (so the results are more recent, I guess?), except that “current year” consistently ends up being 2-3 years ago since I guess that’s what’s in the training data (even on a harness that injects the current date)by jedbrooke
- Can confirm; Claude is quite bad at this by default. Need a special skillby nunez
- Reminds me of using AltaVista search back in the day. Yes, it was that bad.by mthoms
- I’ve noticed this too but it hasn’t been obvious to me that this style of search is not a learned behavior. Tool calling is very much part of the post training phase, I would expect that these style searches just naturally emerge during training. This is just my prior though.by astro1234
- I suspect that this behavior is a learned adaptation. And that it's most likely a feature not a bug.
Based on personal usage, I think it reflects functional degradation of search engines. I've found LLM keyword combinations are more likely to find the results I want with most search engines than mine. Including the big one.
The big one had solved this issue a long time ago by generating those associated keywords based on your input keywords, but somehow, something, somewhere has degraded that system to the point of inanity. And so here we are.
by areoform - Identifying bird species from photos that I upload.
To be clear, it depends on what your definition of "insanely bad" is.
I'd say ChatGPT/Gemini make egregious mistakes on ~10% of my photo uploads.
I recently uploaded a photo of a short-billed dowitcher and ChatGPT told me that it was a Wilson's snipe, explaining all sorts of details about the legs and tail feathers (neither of which were visible in my pic!).
I then followed up explaining that a Wilson's snipe hadn't been seen at my location since last November (and that Wilson's snipe was out of season at my location) and Chat revised its estimate downward to 85% Wilson's snipe.
Again, I followed up and I revealed the precise location of the bird and ChatGPT said something like "oh yeah, 99% short-billed dowitcher"!
I've had similar experiences w/ Gemini (haven't tested Claude).
Again, 90% success rate is pretty good, but the other 10% of the time, the 2 LLMs that I use fail on species ID and often hallucinate features on bird photos.
edits for typos, plus another example from the same "birding outing" the other day.
I uploaded a very clear photo of a sparrow.
* ChatGPT says "song sparrow"
* I explain, "no way. this sparrow has yellow over its eye and the breast is wrong for song sparrow."
* ChatGPT: Oh yeah, savannah sparrow
* I explain, beak is too big for savannah sparrow.
* ChatGPT: Oh yeah, saltmarsh sparrow.
* I expalin, "no orange on the bird's face."
* ChatGPT: oh yeah, seaside sparrow (finally correct!)
by busyant - One unexpected discovery that I have made while building an AI-based system: the LLM's are bad at designing prompts.
We tend to think that the AI has some sort of self-knowledge and should be good at designing prompts for itself but it's really not.
Been struggling with a task that heavily depended on prompts, ended up rewriting all my prompts from scratch in my own words, and it finally worked. Then every time I ask Claude to fix something in the prompts, it invariably makes it worse.
A very strange phenomenon that can probably be explained by the quality of prompt design advice that made it to the training dataset. Bottomline, all the prompt design advice that you can find on the internet is really not great.
by mojuba - Can you share some tips what worked for you?by kubelsmieci
- > LLM's are bad at designing prompts.
You are going down a maddening rabbit hole. Prompts are the things humans write, you are building gas town but unironically
by TZubiri - Serious answer: no model ever gets close to writing an architectural floor plan that makes sense.
They understand all the rules and best practices, they can (sometimes) spot a bad idea in a floor plan, they can describe a good floor plan.
But ask them to make one, even if you give it every detail (even a "node graph" of rooms), they will still output nonsense. Same for text and image models.
Floor plans should be the new Pelican Benchmark.
by jampa