Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- > We also tested Astra on SRE-Bench [15], a benchmark that measures whether models can reverse engineer software binaries to understand its core logic without access to raw source code. Astra solved 88.0% of tasks in a single attempt and 99.2% within four attempts, compared with 55.9% and 68.7% for GPT‑5.6 Sol, respectively.
So the closed source application should open its source in near future?
by quyleanh - I was listening to the Lex Fridman / DHH podcast last night [0], and DHH was saying that this is a new era for open source software. I'd agree, and also extend it open hardware.
Recently I've seen quite a few posts from people using AI to reverse engineer the Bluetooth protocol or such on devices that need a proprietary app. The same thing for firmware is surely coming, which is great as you can get lots of fun hardware from China, but it often has shitty firmware. Once that becomes the norm there's no reason not to make it open in the first place.
[0] - https://open.spotify.com/episode/45lhw2Adbrsw0xSCOgIeg3?si=r...
by fy20 - Not if OpenAI considers reverse engineering an offensive cybersecurity skill.by tintor
- That hero video is interesting.
A projector and speech.
Maybe I'm in the minority here, but I find speech to text / text to speech (but not live audio mode) is quite comfortable and effective for coding now.
The speech to text part can be frustrating if your local tts model does not have word match context for coding. Codex desktop does this remotely well but is slow. I've been experimenting with local software for myself to do this between different llms.
The wall projector is a cool idea because I think it frees the user from staring at a lonely little rectangle while sitting in their fixed office chair.
If done right, this could bring us closer to the dream of more natural, social computing.
Bret Victor's (failed?) project Dynamicland involving a projector on a desk had this goal. I hear he's not much a fan of LLMs. On the one hand, I can see why. But I think, used correctly, it might be the sort of thing that unlocks his dream and, really, my dream, too.
A here's a presentation of Bret's talk on it: https://www.youtube.com/watch?v=7wa3nm0qcfM
Slight tangent: using speech to text to ramble about your rough design for like 20 minutes to an llm produces surprisingly good results over short prompts even when you contradict yourself. They're so good at picking up on what you're orbiting.
- I ran into the same problem as you, so I ended up by coding a local app that is very similar to Wispr Flow, but uses the small english Whisper model on my low-end Windows laptop.
It is still a quite fast. In fact, I just typed this in using this app.
by starik36 - given the fact that we've moved in my office from 3-people offices to open-plan office to flex desk now I'm not exactly sure I would want my coworkers to speak all day to their computers and gesturing / walking in front of a projector (provided there will still be coworkers with IA)by makapuf
- which raises the question, is the model in the demo actually gpt-6? or it is gpt realtime 2.1? It's unclear how gpt-6 can interact at the realtime level and if so, how can developer get access to it?
- >this could bring us closer to the dream of more natural, social computing
What I saw was multiple people living alone in a small box in a warehouse (probably filled with other boxes) with all of their natural, social interactions directed at a wall. I wonder if this is foreshadowing for the future of work, at least it is what work will look like as envisioned by OpenAI.
by bloggie - It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?by XCSme
- The point is to inject something into the process that these AIs can't do for you.
People SHOULD feel like making a useless Mario Kart clone isn't worth the effort anymore. They should, instead, be trying to figure out how to actually use these models to make something that doesn't feel like a useless Mario Kart clone.
by nater5000 - You're not being ambitious enough! Spend your tokens now building the primitives and foundations of much larger, complex systems. No matter how much faster and more efficient models get, eliminating the gruntwork will always pay dividends.
- Yeah. I’m almost glad I didn’t invest any time in any of my 100s ideas for a startup. Most of them would be destroyed by AI by now.
But, you can create cool stuff just for yourself. That’s the upside. It’s just hard to make a living on cool stuff for yourself.
by bgarbiak - Maybe instead of creating cool stuff try to go and solve real problems? It seems to me that we are lacking in that department since all that LLM fuss has started 3 or so years ago.by maxnevermind
Live a life doing whatever makes you happy.> Like, what's the point, if the next AI can do it in 5 seconds?Post-work society is an inevitability if we don't destroy our planet.
by gavinray- It's less lack of interest in creating that bothers me, it's my lack of interest in learning – it would surely be crazy for a SWE to care about how some new framework works anymore? Even if someone could reasonably argue that it might be slightly useful today there's almost zero chance it will be useful in 6-12 months times.
But it's not just tech – my lack of interest in learning and creating is starting to generalise with the models. Music, writing, coding, maths, etc...
I need to get used to switching my head off and asking the AIs to think for me whenever I need to engage my brain. It still feels very unnatural.
by kypro - > Like, what's the point, if the next AI can do it in 5 seconds?
I built a phone app recently, not released to the public, just an idea I had for ages but could never spend the time actually building. Its 100% vibe coded, and took me a few weekends to build... I'm talking a few hours in total.
The point I'm making is that you now have the power to create stuff you would never have had the time to build. You can think big, wild stuff. Experimentation. Throw-away code.
What a time to be alive!
by Flere-Imsaho - „The depressing thing about tennis is that no matter how good I get, I'll never be as good as a wall.“ -Mitch Hedbergby billypilgrim
- - OpenAI claims Astra beats all benchmarks (compared to Fable and Opus, except "Humanity's Last Exam (w/ tools)"): https://openai.com/index/gpt-6-astra/
- Artificial Analysis scores Astra (max effort) as 61 points on intelligence, behind Opus 5. https://artificialanalysis.ai/models/gpt-6-astra
Who is wrong here?
Some benchmark results in Astra page for Fable and Opus are blank (-).
What is Artificial Analysis intelligence index measuring that Astra scores poorly on?
Can someone from OpenAI / Artificial Analysis comment / clarify?
Even OpenAI Astra page mentions the low scope from Artificial Analysis for Astra.
by tintor - Used curiously fewer tokens, however.by causal
- If you scroll down in the Artificial Analysis page you linked, you'll see all the individual benchmarks.by AnodicElegy
- > Humanity's Last Exam (w/ tools)
This is one of the only benchmarks that actually matters for testing the frontier however. Other benchmarks can be gamed by simply being more persistent, but HLE is a diverse set of open-ended research-level questions. It tests domain knowledge and problem solving skills. Burning more reasoning tokens may help somewhat but not as much as e.g. coding benchmarks.
by zarzavat - Many people claim that the Artificial Analysis Index is highly contaminated - I have not personally looked into it.
Though, unlike the creators of benchmarks like Terminal Bench or ARC AGI, the Artificial Analysis Index team does not seem to have deep technical or ML backgrounds. They are ex-strategy consultants, McKinsey, et. al.
by kubrickslair - I really, really don't find the Artificial Analysis Intelligence Index credible anymore. It's some weighted score of benchmarks, and benchmarks increasingly don't reflect how good a model is.
That should be obvious if you compare Gemini 3.8 Flash (which is an _excellent_ model especially for its price and TPS!! but 10min of prompting in any harness) will tell you it's nowhere near close to Sol/Astra.
But AA scores Gemini 3.8 Flash at 59, and Astra at 61.
by dannyw - Related: OpenAI begins rolling out GPT-6 Astra - https://news.ycombinator.com/item?id=49554273
How about we stick to that one for talking about the rollout, and this one for talking about the model?
by dang - Finally, OpenAI has a Fable/Mythos class model. 5.6 Sol felt like 5.5 on steroids, probably just a different checkpoint with a lot more RL post training.
I wouldn't be surprised if there are some conceptual similarities to the kind of latent reasoning Anthropic sees in claude's J-space, although those aren't the same thing.
Recurrent/looped transformers themselves aren't a new concept, but it's interesting to finally see this approach show up in a frontier production model.
Canceling my Anthropic Max sub when this ships.
by HAL3000 - Sol easily outperforms Fable on every task I've tried it on.by CSMastermind
- > Canceling my Anthropic Max sub when this ships.
At this point, it reads like people are cancelling old ones and getting new subscriptions every two to three days, whenever a new ,model drops, and quite possibly by the end of the week they are back to the old provider while still having active subscriptions with at least two to three others. Interesting times.
- yeah i'm wondering the same way... especially in light of the 20x debacle (where we found that 20x of Max vs 5x only applies to the 5hr limit, not the weekly limit, whereas OpenAI's 20x actually is 20x overall).
Also Opus 5 has been really tough to work with. I can't understand half of what it says, it's just so damn obscure.
by atonse - I feel that for some time now, the biggest constraint when working with models is not their intelligence, but their speed. It does not matter how smart the model is, it will make mistakes, because the instructions are ambiguous and new facts are found during implementation. The biggest problem I've had working with software developers has always been the lag between seeing the results and steering towards the right direction, not the skills of the developer (with many exceptions of course).
Regardless, working on the wrong things is time wasted. And again, I'm procrastinating here while waiting for Fable to run a benchmark on a few solutions to a problem I have. We can guess what would work, but we only know after the benchmark. A faster model, with fewer capabilities, would've been a much better choice this time... well, "git gud" they said... and live and learn! Faster model = less time for procrastination.
PS. AI models don't live and learn; the discussion about AGI is pretty pointless imo. It's a tool. Does it matter if it is AGI or not if it does what you want it to do? Does the IQ of your colleague matter if he's good at what he's supposed to do? Or bad? Well... I guess it does matter, as many people are up in arms about whether Astro is AGI or not. Personally, I think we're past the point for that debate. These are amazing tools.
by tappio - I'm sorry but no - output quality matters a lot more for me.
Just yesterday I tried to use Google antigravity to do a side project I've had on the back burner for 10 years now. Gemini flash is insanely fast - at first I was amazed at how quickly I was getting responses, and it seemed to hold it's own in technical discussion, although sycophancy is next level. But then when I actually let it do the coding part it was just drivel. I wouldn't even bother improving that code - like cleaning up after a lazy unskilled coworker - throw everything away and start over because the foundation is just leading in bad direction.
I spun up Astra on the same problem and although it was sluggish in comparison, and much more pedantic about irrelevant details - the feedback/pushback was actually meaningful. The implementation PoC also took tweaking but we got on the same page really fast.
Gemini Flash 3.8 was just producing garbage ultra fast, Astra could actually be steered into a direction I want and it provides valuable/insightful feedback.
I don't have infinite reading capacity/mental stamina - I would rather the model take it's time and let me see something high quality rather than get bombarded with garbage. If it can be faster that's great - but I'll always default to smarter model. The only exception is stupid trivial tasks like log analysis and similar.
by rafaelmn