Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • I don’t feel the existential dread of mathematicians is correct. It seems to me in fact these results are bringing math mainstream. I now personally look forward to the interpretations and discussions of the significance of such results by human mathematicians.

    Now I understand that it’s mostly the super stars benefitting from the increased attention. Folks who are less established don’t share in that glory. But on the other hand it seems like an exciting time to go even deeper for in various specialties of math by deciding where to focus these powerful tools. For every conjecture defeated some seven or eight new ideas open up. Our path through that combination will be set by creative and curious human mathematicians.

    [edit: deleted a distracting comparison to Chess]

  • As in chess and go and also coding for the past ~year there are two groups of people: the disappointed and the enthusiastic. The disappointed are sad that they lost their advantage and that the craft they honed for years or decades has rapidly lost its value; the enthusiastic are excited about the future and what computers can bring to their domain and how it will evolve. I’m a bit of both if it comes to programming, more enthusiastic than disappointed, but also more than a bit terrified about the pace of it all. I imagine that’s how Kasparov felt back then, that’s how Lee Sedol felt and now that’s how Terry Tao feels.

    The most disappointed folks will simply drop out, but the enthusiastic ones will keep going and with luck make up for the ones who decided to quit. Chess and go certainly went this way.

    by baq
  • i think i agree, we are going to have /more/ math and now need /more/ mathematicians (we are seeing https://vibemathed.com/)

    these LLMs are great are generating arguments but they don't ask questions, we will need mathematicians to shepherd them into more discoveries

    i really want to see open weight models crack some breakthroughs

  • Do mathematicians have the right to say "no AI PRs please, the volume is too much" just like how some open source maintainers do it? I guess they feel a loss of control, there is no way to turn the hose off.

    Thinking of this a little bit with the perspective of every new proof as a burden, dumped for review by actual mathematicians.

  • Every time someone makes a comparison to chess I die inside. Chess is a spectator sport primarily funded by a few eccentric billionaires. Players artificially constrain themselves in timed environments knowing that they will never be able to produce better moves than a smartphone because a select few people find it interesting. Only ~30 top professionals actually make enough money to have a full career playing chess, maybe a few hundred more can sustain a meager lifestyle with coaching gigs. I shudder to imagine what will happen to the tens of thousands of non-Fields medalist caliber mathematicians if math goes the way of chess. Perhaps Terence Tao and a few other famous mathematicians will be funded by Peter Thiel to report on how well humanity can keep up with the machines? How do you expect any mathematician to be optimistic about this comparison.
  • The chess analogy is awful. If you simply want to know the answer to a chess problem, give it to the engine. Chess only lives on because it's a competition between humans to test their skill (just like bicycles, cars, trains didn't eliminate foot races) ... the computer is largely factored out, but not entirely -- people train with the computer, use it to check whether they played correctly, ... and they cheat. A lot. Thus there are more and more sophisticated mechanisms to detect and prevent cheating.

    If you translate that to math, then all you get is math competitions, not math as a career. Of course the translation isn't nearly exact ... there's a lot more room for professional mathematicians because the math space is far more vast than the chess space and can't generally be cranked out mechanically (we have proof).

    P.S. The response is nonsense ... I explained exactly why it's awful (others have too) and the response doesn't in any way refute the explanation ... rather it offers up a ridiculous strawman.

  • Given that we were nowhere near this state even two years ago, I think it’s a question of velocity more so than just distance.
  • I find this whole way of looking at things weird. Did maths exist just to entertain and employ mathematicians? Surely maths is, like, useful? Not immediately, not predictably, but in the long run? In which case, whether mathematicians feel bad about it is mostly irrelevant - it's like complaining about the railway because it may put coaching inns out of business.
  • The old way of establishing career credibility is being destroyed, for better or worse. Accomplishments that used to be career-defining are hard to distinguish from AI, and correlate more with access to compute. Think about Bill Gates's math paper he wrote in college. That kind of thing is gone now as a path to credibility. There's still competitions and grades, but the diversity of paths is going away. Maybe new ones will open up. This is a competitive advantage for old people who have credible pre-2025 accomplishments they can point to.
  • problem number 1 and 9 are surprisingly very intuitive

    check here : 1. high dimensional sphere packing https://muchmirul.github.io/conjectures/sphere-packing/

    2. multicolor ramsey number https://muchmirul.github.io/conjectures/multicolor-ramsey

  • The first link is very sloppy and doesn't actually explain why the "certificate" proves anything about the sphere packing. Or if it did, I couldn't understand it.
  • My main gripe here is the lack of transparency around the total experiment and construction. I doubt that they simply pointed their model at these ten specific problems alone and gave the model one shot; therefore the $2000 number could be completely misleading, similar to P-value hacking by not disclosing the total experimental setup.

    I want to know:

    1. How many total problems were given to the model, and what percent were left unsolved at what cost before giving up? 2. How many attempts did you give the model at solving these problems? 3. How expensive was the harness, e.g. did the model have access to a job cluster?

  • > therefore the $2000 number could be completely misleading, similar to P-value hacking by not disclosing the total experimental setup.

    I don't think that comparison to p-hacking is fair. I mean not reporting price of all run is nothing like committing scientific fraud and fake results.

  • Yeah I remember reading about something along the lines of Mathematics is now about the scaffolding around you find the problems/solutions not just the problems and solutions. For teaching purposes. This was before this ai craze
  • I don't think you want to bring cost into this argument.

    Even if the cost was $1 mil for these 10 problems, that's maybe 10-20 math researchers for a year.

    Do you really think that if you paid that to humans, they will deliver the same results?

  • I guess people will always find something to gripe about.
  • I believe we're seeing a new kind of mathematics that will require completely new formats for publication, a bit similar to those used in experimental sciences. AI-powered mathematics should be fully reproducible, so it's the authors' responsibility to disclose the exact model type, inference settings/seeds and the full prompt history leading to the result. Of course that would ideally require open weights models.

    It's not just about requiring to disclose AI use. AI-powered mathematics is a completely valid discipline that doesn't need to be shy, but it should develop its own publication culture.

    by c7b
  • It seems like they threw it a decently large battery of open math problems and probably limited it to something like $200-500 per problem:

    https://x.com/polynoamial/status/2083478171975082334

    As a complete guess, it seems like they tested hundreds to thousands of problems with a relatively low per-problem budget

    --

    The linked tweet from Noam Brown at OpenAI reads:

    > And yes we did try other major problems without success. Sadly no Millennium Prize problems (yet).

    > But also, we didn’t spend a lot on each problem. It’s possible to push test-time compute much further.

  • Can’t wait for this stuff to have quality of life increases for the average person. So far all I see is that AI has made owning a computer more expensive, made some jobs redundant, increased spam and distrust with questionable authenticity of content and of course made some Americans very rich.
  • Downvoters mind to explain?
  • If it doesn't help average people, why do millions of them pay for it?
  • It has had significant quality of life increases for me. I use LLMs for everything from:

    - travel and restaurant recommendations. my last few outings have been entirely LLM-advised and they turned out excellent. LLMs seem to have ingested every single Google review, photo, and menu of every business on Earth and can answer very nuanced questions like "is the garlic chicken at <restaurant, city> garnished with coriander?"

    - fitness, nutrition, accounting, therapy, medical, legal, immigration advice (sure it's not a real professional but you know what, it's pretty fucking close, and any capability gap is made up by having perfect two-way communication which you don't get when talking with a human)

    - coding (work, side projects, personal tools, documentation & pricing questions, "review this code", etc).

    - I start reading most articles with the prompt "Summarize this article: <url>". I just started a non-fiction book by pasting into Claude: "There are 12 chapters in the book <book-name>. Can you give me a 2 sentence synopsis of each chapter?". It reduces the "activation energy" hump and screens if it's worth reading at all.

    - I use the LLM in my Tesla for on-the-fly advice for parking and other things. You can simply ask "what's the best Boba place around here?" and it will give you a decent recommendation. You can also follow up with "does this place have ample parking?".

    - I use the LLM in YouTube to summarize videos and ask specific questions and/or get timestamps to the parts I care about.

    If your critical thinking skills are strong then LLM is a literal superpower.

  • Not everyone works for Evil Corp. I work in the public sector and my work supports public health and safety initiatives. AI has allowed my team to get much more done than we would have otherwise which improves the quality of life of the people in my community.

    So I would like to counter your cynicism with a “YMMV” depending on who you work for.

  • Henry Yuen's (whose work problem 6 builds on) comments on this are worth reading IMO: https://bsky.app/profile/henryyuen.bsky.social/post/3ms2jpch...
    by gpm
  • It sounds like he hasn't verified the results of a problem that he has personally worked on, so how many of these problems have actually been verified?
  • This starts to feel like chess engines. It’s obvious their play is superior but it’s impossible for humans to understand the moves.
  • Replace philosophers for mathematicians and Douglas Adams was spot on again.

    Whilst current models can't 'intuit' and come up with conjectures, they can certainly disprove some of them very quickly through the kind of grind that humans can't do. I suppose there really are some mathematicians out there today, whose last few years of study, have just been up-ended by this.

    --

    "Yes we are," insisted Majikthise. "We are quite definitely here as representatives of the Amalgamated Union of Philosophers, Sages, Luminaries and Other Thinking Persons, and we want this machine off, and we want it off now!"

    "What's the problem?" said Lunkwill.

    "I'll tell you what the problem is mate," said Majikthise, "demarcation, that's the problem!"

    "We demand," yelled Vroomfondel, "that demarcation may or may not be the problem!"

    "You just let the machines get on with the adding up," warned Majikthise, "and we'll take care of the eternal verities thank you very much. You want to check your legal position you do mate. Under law the Quest for Ultimate Truth is quite clearly the inalienable prerogative of your working thinkers. Any bloody machine goes and actually finds it and we're straight out of a job aren't we? I mean what's the use of our sitting up half the night arguing that there may or may not be a God if this machine only goes and gives us his bleeding phone number the next morning?"

  • > Whilst current models can't 'intuit' and come up with conjectures

    People keep saying this. Why?

    Surely the AI can complete the prompt “Generate new research questions based on these observations”?

    When I read the reasoning traces of coding models they are constantly asking themselves questions and attempting to answer them.

  • Ahh, but you missed the continuation, where they get to the heart of the matter: money.

    "Excuse me, We demand rigidly defined areas of doubt and uncertainty!"

    DT: Might I make an observation at this point?

    MT: You keep out of this metal nose.

    VF: We demand that that machine not be allowed to think about this problem!

    DT: If I might make an observation…

    MT: We’ll go on strike!

    VF: That’s right. You’ll have a national philosopher’s strike on your hands.

    DT: Who will that inconvenience?

    MT: Never you mind who it’ll inconvenience you box of black legging binary bits! It’ll hurt, buster! It’ll hurt!

    DT: [Booming] If I might make an observation …

    “All I wanted to say,” bellowed the computer, “is that my circuits are now irrevocably committed to calculating the answer to the Ultimate Question of Life, the Universe, and Everything.” He paused and satisfied himself that he now had everyone’s attention, before continuing more quietly. “But the program will take me a little while to run.”

    Fook glanced impatiently at his watch.

    “How long?” he said.

    “Seven and a half million years,” said Deep Thought.

    Lunkwill and Fook blinked at each other.

    “Seven and a half million years!” they cried in chorus.

    “Yes,” declaimed Deep Thought, “I said I’d have to think about it, didn’t I? And it occurs to me that running a program like this is bound to create an enormous amount of popular publicity for the whole are of philosophy in general. Everyone’s going to have their own theories about what answer I’m eventually going to come up with, and who better, to capitalize on that media market than you yourselves? So long as you can keep disagreeing with each other violently enough and maligning each other in the popular press, and so long as you have clever agents, you can keep yourselves on the gravy train for life. How does that sound?”

    The two philosophers gaped at him.

    “Bloody hell,” said Majikthise, “now that is what I call thinking. Here, Vroomfondel, why do we never think of things like that?”

    “Dunno,” said Vroomfondel in an awed whisper; “think our brains must be too highly trained, Majikthise.”

    So saying, they turned on their heels and walked out of the door and into a life-style beyond their wildest dreams.”

  • > they can certainly disprove some of them very quickly through the kind of grind that humans can't do

    Of course computers can grind in a way that humans can't. But now we have systems that convert the human-comprehensible ideas into a computer's plan of attack, in a way that greatly expands the frontier of ideas thus treatable.