Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Voting on this is ridiculous. Obviously we should have AI decide which AI challenges have been met.
  • Based on the votes, I can only assume people are still deluding themselves on LLMs capabilities. Is it doing amazing stuff? Yes. But it seems like people still think coding is the ultimate and hardest possible job and so if it can do that it must surely be able to do everything else. My personal experience has show that it still regularly makes up garbage and throws in nonsense sources that do not back up its claims.

    Yeah maybe if your topic has 2 decades worth of text material to absorb it will get it mostly right like with coding, but anything that is less common? Complete crap shoot.

    Just today I wanted to know if platinum cure silicone will be inhibited by plaster. The first 20 results are all AI spam with 30 pages of fluff and thus unreliable at best, so I asked AI directly. At first it says sulfur and calcium will inhibit the reaction, which is bad because plaster contains those elements. Then it says it will be fine according to X sources. Check the sources, none of them have anything at all to do with curing silicone on plaster, the articles are about using silicone molds to cast plaster. Failure.

    Eventually I just had to search youtube videos until I found someone doing it in real life.

    I see the same bad, and sometimes catastrophic, takes on things I have a lot of experience in, like agriculture, construction, and mechanics. It is completely worthless for anything mechanical unless you are trying to start something extremely simple from the 40s or earlier, and even then it will still tell you stuff like "clean the carburetor" on an old hot bulb diesel.

  • Very pleased one of my predictions was totally wrong: https://news.ycombinator.com/item?id=23252711

    Sure, sure, what LLMs make still isn't "efficient bug-free code": my prediction is falsified because while LLMs can write and train new models with machine learning, ML is fundamentally not advanced enough to throw arbitraty new tasks at like this.

  • Too many questions. I bailed after about 10, with no idea how many more there were.
  • Heh, there's one of mine: https://stoppels.ch/goalposts/?c=39727943

    "GPT-4 looks at original ASCII art of a foot, not copied from the web, and says it is a foot."

    The vote is currently 64% yes, 18% no.

    Just now I asked Opus 5.5 to generate an ASCII art foot, and it did a passable job. It's not great, but it's a foot. Then I pasted it into ChatGPT (whatever they're serving to the free tier by default, which seems to be 5.6 Luna), and it said it was a "train/locomotive": https://chatgpt.com/share/6abeaa39-cc80-83ed-851f-29370db089...

    Maybe it's Opus's fault for drawing a bad foot but I think it's fair to say LLMs are still pretty bad at ASCII art (without additional tool calling etc).

  • At least 1/3rd of these predictions aren't clear enough to determine exactly what is being claimed/predicted. Even after reading the full comment multiple times, on a lot of them I couldn't tell where the author had set the goalposts well enough to say whether we've crossed it or not.
  • There's one of mine in there where I predicted in 2023 that it would be 20 years until AI would be reliably able to entirely build and deploy arbitrary applications from a prompt. I was off by about 18 years on that one!

Explore Birbla archives