Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- LLMs are an excellent tool for discovering which problems mathematicians thought were interesting, actually aren't. If the LLM's haphazard jaunt (somewhat exhaustively) through the problem space can 'solve' it, the solution was never going to be particularly useful or illuminating.by inboulder
- My take home from this entire drama is that one should not use LLM services for confidential or proprietary information as they all seem to be run by assholes. And you’re sending them everything you are doing. Would you send your lab notebook to an asshole? Hell no.
I say that as a mathematician (on paper) who perhaps surprisingly doesn’t give a crap about the problem itself.
by sdcfgy - > as they all seem to be run by assholes.
Wait until you learn how many other SaaS and web 2.0 and cloud based things are also run by assholes.
You know this metaphor? https://en.wikipedia.org/wiki/Turtles_all_the_way_down
But instead of turtles, it's assholes.
But more seriously, no, none of what I wrote above is an attempt to excuse or play down the specific role of assholes in large AI companies.
by walrus01 - Or if you're going to trust one of them, maybe it shouldn't be OpenAIby hn993302
- Or use Lumo from Proton. Are there any other privacy first companies offering LLMs?by johanvts
- Their privacy policy for normie subscribers says in plain English they use your Personal Data for research. I think it’s pretty unreasonable to use the service and expect otherwise.by junofan
- That they do it is just concerning to me in that it says that home-ran models just aren't good enough. Surely researchers like this have the processing power to run them at home, they just don't have the processing power to train models of comparable level.
This is something I feared would happen and where open source would be left behind. Maybe they can do something with crowd-sourcing computational power from volunteirs. They were after all able to get Leela Chess Zero to be comparable to AlphaZero by training from volunteer processing power but it seems to me we live in a world now where the best models keep their stuff closed.
In imagine generation too. I'm not sure how well Stable Diffusion can compete in following instructions with all those advanced models that are kept secret.
- > one should not use LLM services for confidential or proprietary information
That’s obvious, isn’t it? Just like you wouldn’t upload your confidential documents to an online spellchecker, or your proprietary code to an online compiler?
by pansa2 - I keep looking for a technical article to appear on HN discussing literally anything about the mathematical result--- Not fluff, not marketing, actual content.
Instead, all I read on HN about N.-S. is human soap opera, told from every possible angle.
In 100 years we won't care about the soap opera. The N.-S. result itself will still matter.
Someone, anyone, please, submit articles on the result itself.
by RhysU - If you want a level headed technical discussion instead of pitchforks and drama you have come to the wrong forum.by paxys
- If you'd like to explore what the Navier–Stokes blow-up construction looks like visually, I vibe-coded an interactive 3D visualization based on the published result for fun :) Demo: https://minfx.ai/navier-stokes/ Source code: https://github.com/minfx-ai/navier-stokes-blowupby michalsustr
- The main thing to take away from this whole story is that frontier labs have hired teams of mathematicians with the sole purpose of solving open problems.
If this doesn't show you that the fantastical claims of their LLM's ability to solve problems on their own are bullshit, I don't know what will.
It's pretty clear all of this is a marketing effort only, and they pass human results as LLM findings.
We already know that the supposedly industry-changing Mythos and Fable results were actually complete BS and they're just your run of the mill model. There's nothing at all to suggest this is any different, and once we get this "unreleased model" (aka bob from the math department), we'll see it was all lies again.
by iLoveOncall - What's missing from the story I think is the part about "...then had a breakthrough on August 15th. The mathematical rumour mill kicked into gear...". If only the two of them were working on the problem in secret, how did their breakthrough become a rumor?
- They weren't working in secret?by dboreham
- People talk. If you have a breakthrough solving one of the most famous problems outstanding, you're going to tell people.
You'll say, "Don't tell anyone", which they will ignore because they get a rush and perceived status by sharing it. So then they tell someone, along with "Don't tell anyone", etc.
It's a small enough world (both in academic math, one at Anthropic) that you get to OAI in very few hops.
by tyre - Having in mind the allegations by Apple against OpenAI, I don't find it unthinkable that there could've been some form of misconduct happening there.by geraneum
- > My two favourite hypothetical questions regarding this used to be:
> If I'm running Codex and one of my API keys accidentally gets consumed in the context, what are the chances that someone else might ask for an API key in the future and get mine back? (I asked someone at OpenAI once and they called this the "regurgitation" problem and assured me that they take great pains to prevent that... but wouldn't describe how.)
> If I brainstorm with ChatGPT about potential new directions for my company, what's the chance that information might be exposed to a competitor in six months' time who asks "what might company X plan to do next"?
> My new preferred hypothetical for this is:
> If I use ChatGPT to help me partially solve a Millennium Prize problem, what are the chances that my work will influence training such that a later model helps someone else solve it first?
by tosh - Maybe just a rumor of a high value target having their API keys accidentally consumed in the context...by weinzierl
- LLM can't be trained that easily. More like actual human are checking your logs and stealing valuable things from you.by feverzsj
- LLM’s seem very good at solving mathematical problems of which there is an enormous amount of exisiting work/attempts in their training data. This is an amazing capability, but does not convince me that these models are «thinking» or «reasoning» in the way a human does. A human mathematician could in theory categorize/discover an entirely new field of mathematics tomorrow, based purely on their «human intelligence», I wonder if we will see similar examples by LLM’s soon. It seems to me currently impossible that LLM’s can replace human mathematicians, because of their (assumption) likely dependence on human input in the sense of enormous amounts of pre-existing attempts/data.
If an entirely new problem, within a new field of mathematics were to appear tomorrow, I highly doubt an LLM would be useful at all on their own. Is this the «ultimate ASI test»?
by civvv - A lot of what humans do is combining old ideas.
And an LLM could in theory also stumble upon entirely new ideas: there's randomness in how they generate their reasoning and answers after all.
I suspect that we are seeing a lot of advances coming from the combination of existing but somewhat obscure knowledge coming from LLMs at the moment, because LLMs are really good at this. At least compared to humans.
Even before our AI friends became good, they were already known for having read approximately every paper and every textbook published in any language. You only need to increase intelligence a fairly small amount from there to get to something like the 'convex hull' of human knowledge.
Compare https://slatestarcodex.com/2016/11/17/the-alzheimer-photo/
The gist is that basically whenever anyone comes up with a new method you get a big burst of activity of picking up all the now lower hanging fruit, that was previously out of reach.
by eru - The kind of humans who invent entire fields of science on their own come by a few times a generation. It’s fine to say AI isn’t anywhere as close to them in intelligence, but instead is comparable to the “average” mathematician who is building on the work done by others and taking it a bit further.by paxys
- Extreme temperature levels (>2.0) can push a GPT of its manifold, essentially producing predictions barely distinguishable from random noise (it flattens the probability distribution of the next token). In theory this could predict anything including the next field of mathematics (infinite monkey theorem) but realistically that would never happen.
However, how to we know the next field of mathematics isn't a novel combinations of several other sub-fields? That level of mathematics would be indistinguishable from magic to most people and so in their eyes the GPT did something truly inventive.
by krona - If an entirely new problem, within a new field of mathematics were to appear tomorrow, I highly doubt a human mathematician would be useful at all on their own.by bonplan23
- This is also what I've been thinking. The result itself is amazing but it's not like this was completely unexpected. There has been a huge amount of progress on the problem in the last 10 years without which it seems unlikely today's full resolution would have been possible. It is not clear what strategy was taken but it sounds like it borrowed heavily from the two spanish mathematicians. Experts will scrutinize the proof and it will be interesting to see if anything truly original or unexpected was done, outside of known techniques, a move 37.by lhd1
- I find the claims from OpenAI somehow more relatable and reasonable.
- They threw compute on a problem another team/company was rumored to have solved to see what their secret model could do.
- The texts I read do make it seem like OpenAI wanted to talk and share credit generously.
- Imagine working on a frontier math problem with someone at Anthropic and not only do you use Codex but also through a non-business account that allows training on your data.
- Timeline-wise, if they mainly used GPT 5.6 it's unlikely any meaningful data made it into an model that's being internally validated right now.
by pietz - It’s fishy though that they heard one of seven problems was about to be solved and threw perhaps 15 million bucks at the right one.
[Edit: they said "two of": "On Tuesday, September 1, we heard rumors that two Millennium Prize problems had been resolved. .. we launched an effort ... on all open Millennium Prize problems".]
by piker - >While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.
What do you mean, as OpenAI employee, you cannot tell that his work has entered the training data ?
But also correct me if I'm wrong, if the two mathematician were really close to finish this problem, and their conversation were used by OpenAI, shouldn't the Agent have succeeded way faster/efficiently instead of using "4.9 million messages and used about 300 billion output tokens."
by burrish - It's almost as if it was actually found by manually written brute-force algorithm running on OpenAI's massive computer cluster.by feverzsj
- "shouldn't the Agent have succeeded way faster/efficiently instead of using "4.9 million messages and used about 300 billion output tokens."" No. Have you seen LLM output? The fact that it managed to converge at all in the problem space is amazing, an LLM is a 'book', the context window is the page number, the page has a token, that's it, that this huge book of weights already contains a method to search a problem space at all and actually go down a 'good' path eventually is remarkable, but it's going to be incredibly inefficient at doing so, a book doesn't have any kind of real-time model building like a brain.by inboulder