Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Post made by account 2 days ago.
  • Is there an implication here I'm missing?
    by wy35
  • I think I'm probably in the minority here, but for my line of work, the software engineering and coding is only a small part of the work. I write simulation software, so a deep understanding of physics, math, and how they can be applied to the software is absolutely crucial. I'm assuming this model is tuned to be more focused on SWE topics, and the very reason we seek "multidisciplinary" hires is the also why I actually need a jack-of-all-trades model to back my coding agents.
  • I'm also in simulation software! Wondering which models you are finding helpful, the models I'm using for general SWE skills are horrible at our simulations and even basic physics/engineering calculation and intuition
  • I am skeptical. Lived experience is what matters and I don’t have anyone in my life (Devin shop) saying good things about SWE other than it’s free. Hope I’m wrong and it’s not so bad this time
  • SWE 1.6 was great for small tasks. Very fast and good enough. 1.7 was unusable for me. Took more time thinking than GLM 5.2 and seemed to be generally running in circles. I tried it but abandoned it.

    Looking forward to 2 -- maybe it'll be usable

  • Seems like benchmaxing? For example for Terminal-Bench 4 it doesn't have great results. And why not show other benchmarks?
  • Probably, FrontierCode is made by Cognition itself. The model also seems worse in every way than DeepSeek v4.1 Flash, launched today.

    Also the submitter's account is very new which makes me suspicious of self-promotion.

  • At work I setup a cloud worker, where i can spin up as many concurrent agents I want, with unlimited fable 5.1 (thanks employer!!).

    I now just work from my phone, and speak into the agents as they run. I dont write code and I dont write documents. I work on very complicated distributed systems. I dont open my laptop most days. Its a legacy brick I carry around.

    Some of my coworkers are still doing things by hand, and are working long hours to produce 25% of the output (when considering hours worked). I stay quiet with my setup. We are in the end times for this job for the people that can see clearly how to automate their own job

  • that's why after 1 year of product development of these AI 20x maxxed speed, we reached AGI 'wizards', there's really no difference in output, outstanding bugs no longer get solved and sites still suck, even doing things that were just regular development 20 years ago. Are you sure they aren't only producing 2.5% of your output that you manage just by farting into your phone? Are you sure it's 25% really? Seems way to high, days when I have diarrhoea my AI agents move even faster
  • > SWE-2 is post-trained from Kimi K3

    On the one hand I would have expected a completely new model, on the other hand it's an RL-ed K3 go Fable 5 capabilities, which demonstrate that this is probably possible, which is nice.

  • I thought it would be GLM based, as Devin has had free GLM-5.2 for a while now
  • why are all American AI models basically Kimi in a trench coat
    by htrp
  • Not sure it matters when devin is the most consistently shit product I've used. And yes, I tried again, they wasted the money on the billboards.
  • Well, where VC money is involved, the aim of these ads placed in SF is not exactly to convince any potential users.
  • I have not met a single person/company that uses Devin… does anyone here actually use it?
  • > And yes, I tried again

    This made me laugh a bit. I was forced to do an evaluation of their shit product twice due to being backed by the same PE firm; "take a look at it again, it's much better now". It sucked the second time also...

  • Where are the model stats? Is this open-weights? If not, why would I use this over DeepSeek Flash 4.1?

    I think these competing labs need to realize that no one wants another closed-weight model provider... We aren't even happy with the two we have right now, and their days are entirely numbered. If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the closed-weight models in favor of something they have more control over (or is cheaper).

    The big labs love to release their new model and quantize after the first week. You don't have that problem using dirt cheap API rates on OpenRouter. DS 4.1 flash is also faster than fast mode Astra. OAI's subscription rates are good value, but now these new open-weight models are nearly as cheap on API usage rates. I honestly can't wait for the day we're not beholden to the two big labs anymore. No wonder there's so much fear pumping happening at the moment from Anthropic and their funded NGOs.

  • > we're probably a month away from 1/3rd of users moving off the closed-weight models in favor of something they have more control over (or is cheaper).

    DS 4 Flash requires large amounts of memory to run at reasonable quants (I think a system with 160 GB or so). DS 4.1 Flash is even larger, I think around 250 GB.

    Any DS version is dumb when compared (in realworld tasks) to Astra/Opus 5, which means, one would spend thousands of dollars, and still need to rely on cloud services to do jobs that are non trivial.

  • > If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the closed-weight models in favor of something they have more control over (or is cheaper).

    Only in the world where the incumbents don't react. Eg if they saw lots of users moving away, they'd drop prices or do something else.

    by eru
  • This really just exists so cognition can stop spending API tokens with Anthropic or OpenAI.

    Basically any successful AI based service will do this because at scale the frontier models are expensive and you’ll have enough data to fine tune your own.

    Same reason Harvey is doing models now and basically every other provider

  • Cognition, the same company that a few years ago demoed a coding bot purporting to be able to autonomously complete upwork tasks, but upon closer inspection was going off the rails and not even completing what was asked?

    https://www.youtube.com/watch?v=tNmgmwEtoWE

    As others have mentioned this is post trained from Kimi k3, which is already quite capable, so it can't be that bad, but any claimed improvements in performance should be taken with a grain of salt.

  • Previous versions were based on Kimi too. I'd consider if it I could access the model outside Devin. No lock in for me, thank you very much.
  • Well the models did get better but yeah their early product was godawful
  • I'm surprised they can post-train a closed model off of K3. Is that the norm for open-weight licenses?
  • As others have mentioned, it's really matured a lot and at this point is one of the best cloud-hosted, team-managed coding agents, when factoring overall UX, testing and QA lifecycle via its sandboxes, and its ability to be controlled with an API. We use it quite heavily.

    It's coming from a different starting place than Claude Code or Codex are as individually controlled single-developer tools. Devin has been more persistent in pursuing the direction of something that operates more autonomously at the team level, as a peer. And while it might be slightly behind in raw harness ability (maybe?) it's probably ahead on the team-focus.

    by deet
  • A friend recently pushed me to try out their coding platform, Devin, after I decided to move away from Cursor. I had the same reaction: "What, the con artists from like 2024?" But after some cajoling, I gave it a shot and was pleasantly surprised. I guess they learned their lessons, grew up, and are doing good work now, maybe?
  • If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%).

    Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"

  • > benchmaxxed

    An aside: When did talking like incels became cool?

  • When a benchmark becomes a target, it's no longer a good benchmark...