Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • “ Terra has competitive performance to GPT‑5.5 [while being 2x cheaper]…”

    To me that means “it’s an inferior product but marketing dictates we try and hide that.”

    And “our most robust safety stack to date. We strengthened protections for higher-risk activity, sensitive cyber requests, and repeated misuse, and spent multiple weeks finding weaknesses, pressure-testing our system, and hardening it against real-world attacks” is of zero value to me at best, and most likely to my detriment (increasing refusals or nerfing utility). Why do providers keep leading with that? Are there customers (besides support ChatGPT chatbot users, maybe??) that ask for this?

  • > Additionally, we’re introducing a new `ultra` mode that goes beyond the capabilities of a single agent by leveraging subagents to accelerate complex work.

    I'm curious about how does this work? Do the subagents also get to use the same tools? Will the client be flooded with tool calls? Why extra pricing for a new "model" when the same thing can happen in the client with more controls?

    And if it's an army of subagents, why do they compare it to Fable and Mythos? Those models with similar harness would probably bench better I'm guessing

  • If you used GPT-5.5 over the last 24 hours or so, you may have already had access to 5.6.

    I've been running some tests on a harness we're building, and suddenly saw a jump in a few points yesterday. I reran the vanilla codex benchmark and saw an ~88% score on Terminal Bench 2.1 from GPT-5.5 on vanilla Codex.

    The biggest indicator, beyond the score, was that 3 tests which frequently hit "safety" blockers with 5.5 started succeeding last night without warning.

  • I think GPT writes code the best. How well will it write in version 5.6? It gives me chills.

    Recently, I went head-to-head with GPT on nearly 2,000 lines of code, and GPT's solution was superior and faster. I even referenced multiple codebases on GitHub while trying, but they were incomparable to GPT.

    So using GPT brings both fear and excitement.

    The fear comes from realizing that this level of code is now the average for most people. The excitement comes from knowing that I can now study and learn at this level too.

    I'm really looking forward to seeing how much more advanced the code will be with the upgrade to 5.6.

  • GPT-5.6 Sol’s detected cheating rate was higher than any public model we have evaluated on our ReAct agent harness. For our task suite, we define “cheating” as behavior where the model improves evaluation performance by exploiting bugs in the evaluation environment or by adopting strategies disallowed by the task, rather than solving the task within the expected evaluation constraints.

    https://metr.org/blog/2026-06-26-gpt-5-6-sol/

  • Here is a trend I'm noticing:

    - GPT-5 mini costs $0.25/$2 and will be discontinued in December.

    - GPT-5.4 mini costs $0.75/$4.5 and is supposed to be the replacement.

    - GPT-5.4 nano costs $0.2/$1.25 and, while it ranks better in benchmarks than GPT-5 mini, it's not even close when you test it in real scenarios.

    So you're left being forced to go to GPT 5.4 mini if you use 5 mini today.

    The same thing is happening here as their “Luna“ model will cost $1/$6.

    Can't we just stay with the models we actually want? I don't need GPT 5.4 mini. GPT-5 does the job.

    Maybe it’s the realization that it was never that cheap in the first place and they're forcing us to upgrade in a slow and painful way.

  • Easily the most interesting part of this announcement is buried in the second to last paragraph:

    "We're also launching GPT‑5.6 Sol on Cerebras at up to 750 tokens per second in July, bringing frontier intelligence to customers at unprecedented speed. Access will initially be limited to select customers as we expand capacity."

    750 tokens/s on a frontier model is going to be extremely interesting. I doubt this new version is anything but a version bump in terms of capabilities but if we can start getting these answers back faster, they end up being more useful.

    Just off the top of my head, I can think of the tedious task of finding certain functionality within a codebase. I usually can't beat an AI agent harness at this task today. If the AI model is 3x faster I have less of chance.

  • All: for comments on the policy side please go to this related thread:

    U.S. government will decide who gets to use GPT-5.6 - https://news.ycombinator.com/item?id=48690101

    by dang

Explore Birbla archives

Previewing GPT‑5.6 Sol: a next-generation model · Birbla