Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Haha this is spot on how I've been feeling lately. I find it unbearable to work with this model for this reason... any trick out there you can do to steer it not to overcomplicate things? I guess Codex here I come
  • > Worth naming: the Add to cart button is still black.

    Got an audible guffaw out of me. This really is what the experience is like sometimes if you're just giving it a result without being specific in implementation, and it comes out of nowhere, some days much worse than others.

    I've become patient with it, but whatever this style of output is called or doing - it is both condescending and entirely unhelpful, and it seems designed to frustrate.

  • That's so weird... This doesn't at all match my experience with Claude. I've never seen it behave this way.
  • > Why is half the site blue now? I asked you to change one button.

    > Half the site is blue. I asked for ONE button.

    Those are my only options when the site is clearly not blue, two buttons are.

    There is a reason for why I am much more specific than this.

  • At least with Codex, this has not been my experience at all. It still screws up sure, but in every case I can ask "why did you do this" and it can trace back what made it take that particular decision. Typically it's always that I either didn't specify the problem correctly or made a really dumb mistake (executing the task on the wrong project....did this one yesterday) or it's something within a skill file that instructs it (at which point I fixup the instructions).

    Once in a blue moon it's actually the model making a material error in it's thinking and I have to go back and redo it.

  • This is actually what keeps people using AI: variable reward schedule. It's basically gambling.
  • I got way too annoyed at this before realising it was an optional game and I could just close the tab
  • Great site, triggered memories! haha.

    To try to add something to this discussion -- I think that while I've seen these sort of loops less --- what I have seen is "overly helpful".

    Models nowadays want to double-triple-quadruple check things. I'm being silly but it verges on "I have a working solution but let me write a variation in Rust to ensure a convergent solution and prove this works".

    I've had to stop models nowadays mostly because they're being agonizingly pedantic in their validation. Opus is actually one of the most pedantic and "off track" here. But again, not in a bad way. I'm usually like "stop testing latency between 50 runs of this app... this is version one.. we're going to make a million more changes.. you're not buying us anything".

Explore Birbla archives