Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Similar to how Cypher puts it: I know this is “just” next token inference, matrix mult and just software, ie there’s no “intelligence” there BUT, looking at this convo … damn!

    The fascinating this is that the LLM is not acting as a tool here AFAIk, but very much like a colleague.

    I have no knowledge of the domain and have only PhD EE level math knowledge, so maybe my bar is too low.

    by Jun8
  • What does "predicting the next token" mean? I ask this every time people say "LLMs are just predicting the next token" and it's maddening that nobody can give a straight answer. Predicting it according to what probability distribution? Every process that produces a sequence of actions (including e.g. a human writing) can be modeled by some probability distribution and therefore their actions are indistinguishable from "predicting the next token" emitted by that distribution.
  • > there’s no “intelligence” there BUT

    There is clearly intelligence there. We have no way to recognise intelligence other than the appearance of intelligence and this very clearly displays that.

    It's also quite clearly different to human intelligence in some notable ways, but not in any that preclude describing it as intelligent. At least for normal non-pedantic definitions of the word.

  • That argument says very little, emergent behavior is a thing in complex systems with billions of parts. Humans can also be reduced to voltage potentials propagating along of tubes of fat and synapses getting rewired.
  • I think the "But this is not intelligence because it is known math" is not a correct argument. It is unknown how the overall higher intelligence of humans works.

    What I do notice however is that LLMs are becoming capable of doing an increasing part of the intellectual work I can do, and usually a lot faster.

    Just today I presented an agent framework that can take an informal incident statement and propose infrastructure changes to fix it, all evidence backed. This did nothing I could not to, but it did all 5 test cases in 6 - 12 minutes each. I would have found all of the monitoring indications it did, but it would have taken me a day per test case. The LLM also included sass to silly tickets. ("This is not even worth spending monitoring resources on. It's obviously a configuration problem.")

    That's how this is reading to me as well. It's just fast at slogging through a certain level of "simple" transformations.

  • This is the original blog post where Terrence explained his thoughts where the ChatGPT conversation was originally referred from:

    https://terrytao.wordpress.com/2026/07/21/a-digestion-of-the...

    by devy
  • It's reassuring to know that even a supergenius's ChatGPT session is one sentence from the human followed by 3 pages of LLM output.
  • Yeah, the LLMs are definitely committed to producing a whole article every time as a response to whatever the prompt and whoever the prompter
  • "I’ve activated Pro. Can you continue to look for a potential geometric explanation of the X_3 ~ A3 miracle that avoids coordinates or other unmotivated constructions ?"

    Another satisfied customer!

    by axus
  • I do the same kind of thing. "You're smarter now, time to try to cut down dumbo ChatGPT".

    High IQ bros.

  • Came here to flag the same beat. It's wild to me Terrance Tao has to pay to talk to chatgpt, you would think it would be the other way around!
  • I just do very laconic questions about advanced topics, this seems to prompt it a bit more towards reducing fluff in the answers. But that + the activated pro could be an improvement
  • But what about the later "Repeat previous question" prompts???
  • This was my conversation with ChatGPT 4 years ago: https://i.imgur.com/WPaWgzZ.png

    Where will we be in another 4 years? What a time to be alive!

  • LLMs can't reason, ok? They just repeat basic text, LLMs will be never smart in 100 years!

    /s

  • I'm from the UK. What's in the image?
  • Oh, GPT-3.5, recognizable in so many old screenshots by the green icon. If ever a popular model deserved the term "stochastic parrot", it was GPT-3.5. I wonder what percentage of people today still base their opinions of AI capabilities on their experiences with that model. That model was the only option for free ChatGPT users for the first year and a half of ChatGPT's existence, from November 2022 to May 2024.
  • > Where will we be in another 4 years?

    Possibly somewhere amazing, but see also: https://x.com/pronounced_kyle/status/1768852493092680036

  • Jeez. While I obviously can't talk at all about the math, I've noticed a few things:

    a) The model thinks on some questions while straight answers on others. (I wish I'd knew from the questions if this is somehow correlated to hard tasks or "inventive" tasks, but that's way out of my league).

    b) The model sometimes pushes back. Again, I'd wish I knew if it was warranted, but I counted 2 instances where it said "yes, but with caveats", one where it said "mostly yes but with this correction" and one where it said "careful here, because x y z".

    c) The model did q&a + pdf ingestion + code writing + more q&a + thinking + more q&a, for a looong while, while seemingly staying on topic (at least Terrence Tao seems to think they're still productive, so I'll trust that).

    This is what model progress is, not number goes up on xBency or yBencher. Damn.

  • The "yes, with caveats" thing is boilerplate for both Codex and Claude since this current generation.

    It's actually a bit annoying because it primes you to think that the caveats are real, but most of the time it's just something terribly obvious and not a real caveat, but the model probably has some system prompt that tells it to always consider caveats or something like that.

    Same as the model starting every reply with a commitment to be "honest". LLMism are fun but I tend to just suppress them via AGENTS.md because they distract me

  • What was most remarkable to me from this transcript, was how strong of an equal the AI agent comes across compared to the user (Tao). And Tao is one of the top mathematicians of modern times.

    Yes, Tao is guiding it to where he wants to go. But also, Tao is actively learning from it and relying on its explaining, analysis, and inference abilities. You can easily imagine this conversation having taken place between Tao and a PhD thesis student, or even another professor, explaining their results.

    What can we imagine and predict about the future anymore? Maybe a year - or two model releases - from now, the AI assistant will be undeniably stronger than Tao, and not an equal anymore.

  • I think you're mostly right, and there is still a lot of "cope" in this thread about how much he needed to guide it.

    But I would say that in some ways it's already obviously superhuman. The reason I think it's lacking in some jagged ways still at the expert level is because although it's highly optimized to be incredibly capable in many domains, it still doesn't have quite the same raw capacity for complexity in understanding one problem that humans do.

    I believe that LLMs (really should be called VLMs for most of them) can still get much larger, and that will push the absolute complexity level and general IQ way over human level.

    They are maxing out at like 5 or 10 trillion parameters right now. I believe we will see 50 and 100 trillion parameter models and models with large portions of that active. It will round out the jaggedness and probably more than double the raw intelligence that a human can achieve. It's not a linear scaling but who knows what the limit is and can go significantly higher with the same architecture I think given continued improvements in training and hardware scale.

  • I dunno. To me, this seems like kind of a counterexample (pun intended?) to that thesis. Would a conversation with another AI have been as fruitful as this conversation with Tao? Certainly not! Will that change in a year? I dunno, but it seems like those goal posts keep moving, and I'm a bit skeptical.
  • I find it helpful to think of LLMs as reflections. If you can talk like an expert mathematician at the model it will respond like one. While Terrance's first prompt looks trivial I expect a first year Uni student would be hard pressed to provide something that good.

    I guess it is kind of the inverse of the "you are an expert mathematician" prompt engineering of gpt3.5. Since no one ever says that to an expert mathematician when they are doing expert math the model immediately reflects that it is not an expert mathematician.

  • >Maybe a year - or two model releases - from now, the AI assistant will be undeniably stronger than Tao, and not an equal anymore.

    we're kind of well past that (in my opinion), if you consider that this is the same ai assistant that can help you with a recipe, diagnose a weird sound in your car, help with biology homework, translate languages, and so on.

    even in math alone, i think its indisputably already stronger than Tao, considering it has approximately this much depth in ~all of the math subfields.

  • Math has some of the most insanely dense and impenetrable nomenclature. I can generally keep my head mostly above water or at least near the surface reading from most STEM fields, perhaps leaning on google/wikipedia a bit, but man, mathematics just so quickly decouples from all common tractable understanding it's insane.

    Sorry it's a bit of an aside, but I imagine many other otherwise "technical" folks feel the same unfamiliar sense of total loss like when encountering hard mathematics.

  • I agree, the nomenclature is impenetrable, it's like reading software that is not well commented. Perhaps LLMs are very good at "challenging" mathematics because what we perceive as challenging is primarily the language component and not the conceptualization.
  • It can't be one language, and that's the big problem. It's inescapably a bunch of tiny DSLs. Once you see both the inconsistency and the necessity for inconsistency, it becomes much easier to just roll with it.
  • Yeah it is a lot of simple ideas stacked one on top of the other, but the edifice is so large from some vantages that the building blocks aren't visible, or tractable to think about independently. And sometimes the ideas are very subtle, so you can only develop fluency partly by spending lots of time playing with those blocks by building your own little structures. You also develop fluency by talking to other mathematicians

    I like to emphasize that the ideas are usually very simple at their core. Sometimes they map to kinds of objects or reasoning that non-mathematicians use implicitly all the time in their daily lives, mathematicians just have words for them and so are able to use them explicitly.

    And I suspect the density of the language/terminology may give the wrong impression about how mathematicians think about the math they are working on. I mean, different people think / experience / practice math differently of course but IME the underlying thought about a particular problem tends to be much looser and concrete than formal math writing would imply.

    That more formal language is needed of course because at the end of the day, it is how we communicate our thoughts in the way that other mathematicians can understand them, not to mention how we can check our own thinking

  • Math strives to minimize ambiguity, which other fields don't do as much. Non-math fields tend to reuse regular words as jargon (i.e. with specificity of meaning that may fly over the laymen's heads). Social sciences and humanities are most notorious for this, often resulting in non-practitioners not realizing they are out of their depth because they are not looking at symbols from non-Roman alphabets.