Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- Maybe LLMs can't, but another form of AI will. I hope nobody is interpreting this as "nothing will never be as good as us".
I see similar thinking in stories of how humanity got here. Religion has thousands of years adapting to this problem, every time we explain something, the goal post moves. Catholics today accept evolution (or least the church does), but it is the "jump" from monkeys to humans where God is the only explanation.
Just 5 years ago we didn't have a technology that knows more about everything than even most experts. We keep coming up with benchmark after benchmark and LLM/AI keeps destroying them. Now we've moved the benchmark to "the jump". Again, maybe it's LLMs or the way we currently do them that can't do this, but eventually something will.
by inerte - > Just 5 years ago we didn't have a technology that knows more about everything than even most experts.
And we still don’t. What we have are simply very advanced search results aggregators with delusions of personality. Just because your fridge says “I” doesn’t mean it is a person.
by kskdkwkdjw - > Just 5 years ago we didn't have a technology that knows more about everything than even most experts.
We still don't have that. LLMs have shown time and time again that they don't know a single thing and are incapable of reasoning.
by bigstrat2003 - Best comment on this from 6 months ago: https://news.ycombinator.com/item?id=46870575by kfarr
- "Best" because of the irony that it was written by Claude? https://news.ycombinator.com/item?id=46901199by yorwba
- True understanding requires not knowing, and LLMs cannot "not know". LLMs have to come up with an answer, this is their nature. They are search engines. We do have a similar mechanism; one can notice it by reflecting. The mechanism is an opposite of true thinking, as it merely looks up what is already "known". We "jump" when we temporarily turn this mechanism off.
That said, here's an experiment conducted by some Soviet psychologist, I forgot the name. The man wanted to study intuition. So he invented an experiment that was supposed to trigger it in laboratory conditions. (Take a moment to marvel at that; how would you approach such a task?) He gave people a few puzzles. One was to place some sticks according to some rules. Yet another was to find a path in a maze. The secret was that the path in the maze was the same figure as the solution to the stick puzzle.
And he observed interesting results. People who solved the maze after the sticks found the path much faster than the control group. If a subject was asked to comment how he was solving the maze, at the start or halfway through, the speed dropped to typical. Subjects normally didn't notice the similarities.
So there is something to study here, although it is obviously a case of pattern matching, only subconscious. This is a jump of sorts, but not the one I mean. What I mean is a Zen jump.
- I think this is more of a function of the harness and the environment than the LLM. I've seen some LLM interactions over complex environments like Godot and Unity that challenges the notion that there is no "jumping" going on at all.
An LLM in isolation from its environment might as well be a brain in a vat in some dark cave. You need an external environment to sample from and act upon to make forward progress.
by bob1029 - Why does the cave need to be dark if its just a brain in a vat?by altmanaltman
- I'm also amazed at the degree an LLM can get the drift with a vague or incomplete prompt. The ability to perceive and operate based on patterns that go beyond the language in the text makes it seem like they would be unusually good at taking leaps that haven't occurred to us.by doginasuit
- The paper is from the 30th of April this year, openAi announced the counter example to the unit distance problem on the 20th of May. That is to say this paper seems to have aged not much but quite poorly.by yk
- I'm skeptical a counterexample is good evidence of creative intuition.by Hammershaft
- Why everybody is obsessed with replacing humans with LLMs when it seems like the most profitable use cases (like coding agents) rely on enhancing human capabilities?
Until LLMs have some 0% error humans will have to be in the loop (even if they only serve to take responsibility of the process).
by yomismoaqui - Would have been better if AI was known as Augmented Intelligence as it's a tool to to boost our own intelligence. Instead the term that promises science fiction futures predominated, and now it's being used to raise a ton of money.
Saying you want to make workers more productive and provide better tools for people just isn't that sexy.
by goatlover - The theory is that creative leaps in theoretical physics require a grounding in sensory experience, but the obvious counter-argument is that humans can make creative leaps in abstract fields without such sensory grounding. They do address this at the end, saying
"In abstract domains such as Mathematics or Computer Science, the Sense Experience (E) may be grounded in high-dimensional topology or have other goals such as generality or minimality."
But if such sense experience is possible in abstract domains via some high-dimensional topology, why could a sufficiently advanced LLM not develop an equivalent high-dimensional topology for domains like physics and use it to make creative leaps?
by sobiolite - > ...but the obvious counter-argument is that humans can make creative leaps in abstract fields without such sensory grounding...
But we have no idea at how good humans are at that. Given the appalling failures of humans to handle even basic statistical situations like identifying that the same thing happens over and over, it might be that they are hilariously bad at creative leaps in abstract fields, it is just we have had nothing better available to measure against. We've spent about as long as decision theory existed trying to convince people to use it instead of flailing. Limited success, usually in exceptional cases.
And the paper seems a bit dodgy, we have models created with sensory data available. No reason a LLM can't be trained on more sensory data than a human can accumulate in one lifetime. There is a lot of visual data on YouTube.
by roenxi - Isnt the sensory grounding even in abstract cases some (limited) intuition that simulates in a mental world model?
- That's an interesting analogy. My gut sense is that theoretical mathematics requires a high level of intelligence versus more grounded domains. That may imply that deficiency in grounding can be made up for with intelligence and basically reverse engineering the gaps in grounding from first principles/limited grounding. The ultimate question would then be what is the tradeoffs between grounding and raw intelligence for the same outcome.by bob001
- The paper also fails to show that their central example, Einstein, relied on sensory experience for his intuition leaps rather than general reasoning. They just kind of claim that thought experiments require sensory experience. But you can easily ask an LLM to perform a thought experiment and simulate an outcome, and the SOTA LLMs generally seem to do about as well as a human. An LLM would certainly know that freefall feels the same as zero gravity, even if they haven't felt the sensation, which was the key intuition the paper talks about for Einstein's General Relativity. The paper's author would probably say any examples of this don't count, but without clear criteria for what would count, their claim is unfalsifiable.by nearbuy
- If there is enough cross over between real world knowledge engrams and abstract knowledge engram, would this allow for the jump?
One interesting (albeit sad) area which might be related are humans who are never raised with a first language. They seem to never developer abstract reasoning and even seem to lose the ability to develop it later in life. This might indicate there is some 'real world senses' -> 'direct language' -> 'indirect language' -> 'abstract abduction' hierarchy that develops, perhaps related to more real world abductions as a necessary side chain to developing abstract ones.
One of the obvious problems with this is just how difficult we find it to study intelligence purely in humans. We are measure a LLMs by a yardstick that is already known broken, but maybe this is still the right path.
by SkyBelow - This is literally an opinion of one dude which is not backed by any kind of quantitative evidence.
It's actually possible to answer this question rigorously:
1. Define a scientific result which qualifies as a "jump". They should be frequent enough that they happen every year - otherwise one might say humans can't jump either.
2. Identify all such "jumps" in articles published in 2026, and use LLM with 2025 knowledge cut-off to re-derive these results with minimal amount of information.
It really irks me that people boost these low-effort articles just because they confirm pre-conceived notion that LLMs are limited
by killerstorm - I would argue that the whole concept of a "jump" is contrary to the need of quantitative evidence.
The idea that everything needs to be proven with evidence is what an LLM excels at. If you give it a task such as "analyze this massive dataset and demonstrate that is proves or disproves this concept" or "find a counter argument to this theory", it will dutifully run through the information and give an answer. If you give it the task of "invent a new style of music", it will do poorly.
It will not do poorly from of lack of ability, nor lack of originality, it will do poorly as "style of music" is a human concept, "new" is a human concept and "invent" is a human concept. Being human concepts, these are things without quantitative value. There is no truth or certainty to what is "invented", as others may call it "discovered" or "developed" (developed in the sense of progressing from a previous work). While a trained ear my call a song "Breakcore" and not "Gabber", another would clump it into "EDM".
The only solution to this is to have human origin. As humans will tell the tale of how an idea came about, inspiring similar thought patterns in peers and earning the title of "new". This to me seems to happen more often with art than science or engineering. Although, it is hard to tell if it is due to increased creativity, or ability to convey thoughts.
All of this is impossible for a non-human to do, as they will give evidence to support the conclusion rather than the tale. This will turn into a possible addition to the knowledge of humanity, but can never earn acceptance as a "jump", as humans care little for evidence. Humans are inherently social beings and, as a side effect, will never acknowledge a non-human contribution.
An easy way to put this is a literal jump. The humans at Merriam-Webster define it in part as "to spring into the air". However, a quick peruse of the comments on a video about the "World's Highest Jumping Robot" will find disagreements about the definition of "jumping" and "robot". The device does spring into the air, but it does so without human legs, human movement, or human consideration for safety. Should you instead look at a video on the highest jumping bipedal robot, you will instead find nearly full acceptance that what happened was a "jump" (it should be noted that there is much disagreement in comments about if such videos are AI generated). By my understanding of the world, this is due to human ability to excite brain paths of an action when viewing another entity preforming the action. So if the viewer's brain does not fire the signal of a jump motion happening, the viewer will seek any rationale to explain the action as not being a jump.
TL;DR: It is all based on vibes, and LLMs give the wrong vibe.
by JK-Swizzle - You don't need any reproduction. Assuming researchers are using LLMs, you should just see the number of jumps increase as models get better.by NewEntryHN
- I boost it so someone else will see it and prove it wrong.by andai
- LLMs most definitely are limited. What’s your position, that LLMs have no limits? That’s obviously wrong, and the fact you hold such an unreasonable position may be why you react so strongly to pieces like this one.by emp17344
- Came for: "A computer once beat me at chess, but it was no match for me at kick boxing."
TFA was actually about leaps of intuition, sadly.
One of the experiments I've heard proposed around here is to somehow create an LLM from all text up to 1980 or 1990 and see if it can get back to making itself.
by jvanderbot - Article was plenty interesting to me.
- Is that how chessboxing was invented? Genuinely asking.by Tade0
- If would be interesting to see 5 billion LLM's working together, each with random mutations (temperature ig). Would we essentially be looking at a society through a petri dish? Ofc 5 billion is quite a lot of compute.by unfitted2545
- The curious case here is how much of a description do we give it of itself? That would almost certainly dominate success rates.
My feeling is that a prompt would have to provide a vague description of a program that meaningfully passes something like a Turing test, an API to conform to, an expectation of novel construction (no 'ifs all the way down'), and then a requirement to search broadly and pursue promising ideas and not get hung up on the philosophy. Anything more precise feels like it would corrupt the test, but as it is that description feels doomed to loop before even trying the interesting parts.
- I think this can't work because an LLM needs too much data, and before the internet there probably just wasn't enough to get close to what we have nowby elar_verole