Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • When this article talks about ONA it means https://ona.com/ - a cloud agent service that was acquired by OpenAI a couple of months ago.

    (I wouldn't suggest basing my evaluation of the entire field of coding agents around that particular product.)

  • > I was optimistic, but once I had it wired up to one of my projects, instead of making magical hands-off progress on my todo list it spent nearly my entire $20 worth of “ona compute units”, whatever those are, thrashing and trying to get a hold of the todos from linear just so it could pick one to start.

    This is one of the reasons that I've simply not bothered with a lot of these types of AI products. It feels like gambling. Maybe I'll spend $20 on tokens and end up with something awesome. Or maybe I'll spend $20 on tokens and end up with nothing useful and then I'll be glad it was only $20 I lost.

  • I was fairly skeptical of agentic coding before I used it for a real product. Although I still have to be heavily involved in planning the code that LLMs write for me, they can write code much faster than I can, and they know more about edge cases than I do, so they can handle edge cases/subtle bugs that I would have missed. I have been paid to write code at every level of the stack from assembly to frontend javascript, but I'm not equally good at all those areas. In some areas I can still outperform LLMs, but for areas I'm weak they do a much better job than I would have.

    I still think of what I'm doing as software engineering, and I'm glad that I had many years of professional and hobby development before using agents since I think that's given me the ability to make good architectural decisions (and helps me resteer the LLMs when they want to do something suboptimal), but my involvement in actually writing code is quickly going to zero. That said, they aren't perfect and they still introduce bugs, but I believe the quality of my current product is higher than what I would have created pre-agentic coding.

    Things I've found helpful in keeping quality high:

    - Visual regression tests (detect UI bugs before you commit them)

    - Fuzz testing of interfaces and app behavior

    - Automatically add regression tests for any bug that I/the LLM fixes

    - Logging/alerting that tracks an errors/invariant violations triggered in the app

    - Performance metrics that are surfaced in a dashboard.

    All of these are very easy to add since the LLM can create this infrastructure for you. The fuzz testing in particular is something very few products I've previously worked on have since most people don't know how to implement it. I ran the fuzzers for a few minutes and they quickly caught multiple subtle bugs that I was not aware of.

    This is a real product that helps a real, non-VC funded service business, and although I could have made something similar myself it would have taken me a lot longer, be harder to use, and probably be less reliable.

    Edit: while it's true that you can quickly blow through the $20/month plan, the $200/month plan allows you to get a lot done and is basically sufficient for my needs. It's also very cheap when you consider what it would cost to pay someone to do similar work.

  • Hey Kira!

    I'm Matt, I run Product & Engineering at Ona.

    First, thank you for trying the product and for taking the time to write this up. I shared the post with our product engineering team. There are several things in your experience that simply aren't good enough, and we're working on them. I've also credited $200 to your account in case you do want to explore further.

    To be concrete:

    > My very first encounter with the app was that I was unable to login on their desktop version at all. Auth is hard, so I can empathize and forgive this.

    Whilst I appreciate the forgiveness, we hold ourselves to a higher bar than this.

    We've done extensive testing of the desktop authentication flow and haven't yet been able to reproduce the failure you experienced. If you're willing, could you email me at matt at ona dot com? I'd like to grab some logs and work out what went wrong.

    > it spent nearly my entire $20 worth of “ona compute units”, whatever those are, thrashing and trying to get a hold of the todos from linear just so it could pick one to start.

    Looking through what happened, there were a few different things going on here.

    - Roughly a quarter of the OCUs were spent by our devcontainer setup agent. That agent created a PR which standardises the development environment and adds install, build and CI tasks so that future agents can operate in a reproducible environment. There is real value in doing that setup once, but we did a poor job of making it obvious that it was happening, why it was happening, and what you were paying for. We'll fix this. We've also switched that setup agent from Sol to Luna as of today after tuning it against our evals. This should make that setup substantially cheaper going forward. - The Linear flow also involved far too much friction. You went through multiple authentication and setup steps before the agent could actually get to the work you wanted it to do. That's not the experience we want. We have shipped a fix that lowers this friction already, and have a couple more planned (that will take a little longer). - OCUs themselves are our attempt to combine model and compute consumption into a single unit. If someone has paid us money and still can't tell what they're spending it on, that's a problem. We're actively revisiting how we explain and expose this.

    > This only adds more friction between me and my projects, which is literally the opposite of what I want when I pay for developer tooling.

    I agree completely. Today, Ona asks a lot of a new user before it has earned the right to ask for that investment: connect this integration, authenticate that service, let us configure the environment, understand what an OCU is.

    The thing we need to get better at is making the initial experience more boring: sign in, point Ona at something useful, and see it accomplish something valuable before you have to think about any of the machinery underneath.

    > The thing is I think ONA is a good idea. I am evidently willing to pay for this kind of tooling. But I want a version that works without lighting twenty dollar bills on fire.

    Based on your experience, I can see why you reached this conclusion.

    If you're open to it, I'd love to spend an hour with you to help you get Ona working on something useful and show we can achieve what we outline in our marketing. No expectation that it changes your opinion but I'd like the opportunity to learn from what went wrong and show you what the product should have felt like the first time. Email is as above.

    I look forward to hopefully connecting soon.

    - Matt

  • The churn in this space puts javascript to shame. As an example, its only been a few months and AFAICT no one is even talking about openclaw anymore.
  • It's August 9, 2026 and if you're a software engineer who hasn't had multiple "holy shit, I can't believe it just did that" moments, it's time to consider a new trade.
  • I don't know how you can claim it's all vapour ware.

    Two years ago, I couldn't just roughly describe my backlog and then have the code fixed. I had to type it out myself, run it, look at logs, fix toolchain issues, and so on. It was tedious. Or I could get a junior to do it.

    Now can get these things done quite fast, without concentrating nearly as hard.

    Clearly, it isn't vapour.

    It delivers something. That something we have yet to figure out the best way to use, but there's definitely something there that works.

    I get the feeling a lot of people are frustrated because the little gains are lost in organisational chaos, rather than the tools not working.

  • >If agentic development actually worked the way any of them say it does

    I think its fascinating just how much of a gap there is between what's being claimed, and the verifiable observable data of the open source world. Major open source projects are by and large starting to ban LLMs now, because the contributions made by LLM users have been universally terrible and unhelpful. There doesn't appear to be a single major project that's found generating code to lead to major productivity speedups, and the consensus appears to be that its just lead to a lot of crappy contributions that are harder to spot immediately as being obvious crap

    I regularly see people claim that they are now 10x more productive with LLM code generation, and I just wonder where all the code is. Is it somehow true that these gains are only being realised in proprietary projects, and not a single one of them has put even a small fraction of their new found engineering powers into eg Godot? Why do only the poor quality LLM code generation users make PRs to open source projects, and never the engineers that know how to really use it correctly?

    If you look in the open source major project space, you can find almost no evidence that AI code generation exists at all. Go browse your favourite critical tool and look for AI generated PRs that have landed in the codebase, its probably a tiny handful of them in comparison to the human written PRs prior to an LLM ban. It turns out that once you have a verifiable, open quality review bar, for some reason almost no LLM commits really meet the level of quality necessary

    I strongly suspect that what we're seeing is that much of the tech code-writing economy had already become completely performative prior to AI turning up. It no longer matters in the current age if your code is good, or works, because your job is to give the illusion of product development while the stock market price gets pumped, until you all cash out your share value, get bought, or hop jobs in 2 years. For many companies it literally does not matter if you produce anything that generates value (or works), because the illusion of progress is all that matters. AI is absolutely incredible at creating the illusion of progress, because it looks a whole lot like real code, it just appears to have failed the bar of making actual projects that work. If that was never the goal in the first place, it probably really is a 10x productivity boost

    by 20k

Explore Birbla archives