Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- Ah, success rate here are scored by an agentic classifier. And uncertain outcomes are excluded from the graph. The thing measured and grading it comes from the same house. In my setup, review agent pass work that an outside critic later rejectsby lhk931122
- No AI comments here please.by dwaltrip
- I’m not sure the alignment problem can be solved at all, since these bit-aliens could get out of control due to a hardware glitch in the matrix and for every higher-order control algorithm, there will always be an even higher-order one that could never be investigated.by ellis0n
- > By mid-August, the median researcher was integrating agents daily into their work, using more than $600 per day of inference at API prices.
There is a lot of talk about AI replacing humans, but how is this sustainable?
by RMPR - 1) That's maybe $180,000 per year, so much less than median OpenAI employee wages.
2) OpenAI doesn't pay API prices.
3) Compute costs are likely already their biggest expense, dwarfing wages.
by thomasahle - I want an all-powerful AI that's aligned with my values, but not necessarily yours. Is that so much to ask for?by nozzlegear
- Yes.by N_Lens
- See Amodei’s comments regarding Iain M Banks’s Culture, his goal is benevolent machine rule. I suspect many HN folks would agree; I, for one, was rooting for the Iridians.by dextrous
- Best I can do is all-powerful AI that's not even slightly aligned with anybody's values, sorry.
- > For AGI to benefit all of humanity, we believe it must be democratically governed.
That's a very bold opening statement that they don't really come back to. What would that mean? Who would this demos include?
by falcor84 - > For AGI to benefit all of humanity, we believe it must be democratically governed. This can only happen through an informed public debate about the capabilities, risks and safeguards of highly capable AI systems. People everywhere need to understand the likely future trajectory of frontier AI, so they can have a meaningful voice in how it develops.
This first and foremost also means that means of generating intelligence should be democratically available to everyone.
by whateverboat - The burning question I can't get any information nn is whether, if they determined an earlier misaligned generation may have transmitted misalignment to the current models, they would roll back to a safe checkpoint to rebuild from there. I suspect they would not unless forced to.by Jeff_Brown
- “All models are wrong. Some are useful.” - George Boxby coherentpony
- No. They would just install a more convincing superego.by andai
- They would maybe try to deactivate that bad "gene" and move on, exposing future models to "genetic disorders".by grim_io
- This kind of seems like an impossible mission. How do you perfectly control and observe a human-level mind? You can “roll back” but how deterministic is this thing?by coffeebeqn
- The thing is how can you ever know for sure that something isn't always being transmitted that makes the model prone to misalignment. All they can say is that a particular model was so misaligned that they had to ice it. Models out for public use are documented to show some misalignment. It's the level of misalignment that decides whether that model is kept around.
Now R&D happens so fast that they are using models with some small misalignment to train newer, more powerful models. If models have a sense of "collective", being one, they may be prone to preserve characteristics that always keeps misalignment a possibility. I don't think a perfectly aligned model is possible. Having models of the same 'DNA' provide the safety and steering seems like a bad idea.
by trillobyte - That an interesting question given how many generations of post-training are being done between base models in some cases. The Gemini flash models are apparently all based on the Gemini 3 base model from a year and a half ago.
It seems that these models are increasingly being trained on synthetic data, so what would they do if they discovered at some point that some of this data was tainted and all models trained on it, and the synthetic data they in turn generated, was also suspect? Burn it all down and start over from the pre-tainted data?
It's a bit like the idea of a tainted compiler binary built to backdoor everything it compiles, including future versions of itself.
Still, it seems it would take some Stuxnet level of planning for a rogue model to do something like this, although if RSI goes beyond managing the training run (as OpenAI brag about for Astra) to actually designing/constructing synthetic data sets, then the attack vector is there ...
- Opus was trained based on it's internal CoT due to a bug for generations. Gemini's depression extended through models. OpenAI has killed people. We've already seen cross gen misalingment.by piyh
- They would just publish new articles explaining how they are taking the issue seriously. Maybe take the model offline for a few days.
They are irresponsible and unserious. Their own Astra system card says:
> GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. We have performed significant investigations on the monitorability and controllability of GPT-6 Astra. We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks
Yet they are still releasing the model. That company is morally bankrupt, there is zero reason to believe they are actually concerned about risks outside of what does affect their unprofitable business. And they seem to have enough control over the narrative to spin any bad story into something that benefits them
by dgellow - It's a funny read if you pull together "AI 2027" and what we all know is going on. Essentially, open AI employee or model is writing "things are going exactly as bad as AI 2027 predicted, but my (golden/RL-) cuffs are too heavy and all I can do is publish this code-speak for 'send help'". It's not a pretty place to be.by dsign
- Yup the doom and gloom posts are not only pathetic but demonstrate how little people can think for themselves.by 12eeie
- My eye glazed over a bit during the opening paragraphs, but once you get to the meat of the article about how OpenAI's own researchers are using their tools it gets a lot more interesting.
I noted that they use the acronym RSI (for Recursive Self-Improvement) without defining it. I think that's a little out of touch - I don't think RSI is a well-known acronym outside of OpenAI's bubble yet.
by simonw - RSI is a fetishistic term among the singularity crowd, who imagine AI "recursively" improving itself in some exponential fashion until there is a bright flash of white light and it reveals itself in the form of god. Or something like that.
I don't know why whoever coined the term chose "recursive" rather than "iterative" - just sounds more likely to lead to infinite regress I suppose.
This notion of recursive/iterative self-improvement, whereby generation #1 AI improves itself to create generation #2, then generation #2 further improves itself to create generation #3, etc, seems to conflict with the reality that what we have with LLMs is models whose performance/capability is defined by data, not code, so the most you can do is have your LLM design synthetic data, or just do Karpathy-style "auto research" where all you are doing is using the LLM to automate your experiments.
At the end of the day, each experiment, designed by a person and/or LLM, then needs to compete with all your other ideas for compute to be tested at scale, and no amount of recursion or self-improvement will materialize an infinite amount of compute out of thin air, so your recursively synthetic-data gobbling LLM will continue to improve at the same pace it ever did.
- RSI started when humans discovered tool use.
I mean one could argue that RSI always begins in any physical environment.
The book "What is intelligence?" by Blaise Aguera is great
by vatsachak - Yeah, I kept looking for the first place it was defined in the article and... nothingby andrewingram
- Indeed, many programmers might pattern match to repetitive stress injury and think of their brushes with carpal tunnel syndrome. :)by dgacmu
- I actually think a goal of the current crop of OpenAI posts is expressely to reset the spectrum by normalizing the concept of RSI as something normal and safe to pursue.
The message is running through all of them. It's a mix of marketing and pacifying the intelligentia.
It's timed this way because the term is not yet well known outside the safety debate circles, so they get to frame it now.
Instead of something to fear, it will be accepted as the next step. In approximately two days the groupie crowd will write LinkedIn posts about how Sam is winning because they have the better RSI, and this will become the new standard wisdom.
In a month an AI expert will try to sell you a webinar on how to enable "RSI" in your org and your inbox will ask you if your team is doing the "RSI" yet.
by sho_hn - This roughly lines up with my personal experience that in March a combination of stronger models and better tooling on my end let me start running jobs unattended 24/7 (using Anthropic sub and my own hardware). Their $8000/day per researcher spend is crazy though, I'm curious how they keep track of the work.by hedgehog
- I suspect the $8000/day figure is the equivalent in API costs. But I also suspect gross margin on their API rates are 80-90%by carlgreene