Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- Interesting topic. That said I don't know how useful this is since LLMs are primarily trained using mode-covering training rather than Mode-seeking(RL) training, which means LLMs can not form (and does not have) the same underlying structure to their models of language that humans have.
A LLM does not learn topic by topic, it learns everything all at once and slowly integrates it in to a single knowledge system.
- "Capability stays inside the curriculum" implies that even much more advanced models are not able to go far beyond their pre-training data. Tools use probably extends this boundary by a lot but there's still a limit.by anavat
- > In our experiments, scaling, SFT+GRPO post-training, and in-context learning amplify what the curriculum taught, but none meaningfully improves out-of-scope performance, indicating that the pretraining filter sets the effective capability ceiling.
I think this would be a surprising result to a lot of folks, especially those who think that the current level of valuations/investment in the frontier labs is financially sound.
by abtinf - I remember reading something a few years ago, about how if you train an LLM with the reading material sorted by grade, the training becomes more efficient? Does anyone know about this technique? How does that work?
I'm assuming the knowledge doesn't end up as separate "layers".
I'm also reminded of how the human mind develops in distinct stages (e.g. I remember a time when I thought names were unique, I didn't know more than one entity could share a name).
by andai - > Does anyone know about this technique? How does that work?
The term you are looking for is Curriculum Learning. There are several papers exploring this. From memory, it leads to faster initial loss drop on some experiments, it can be good for RL if you start with easy -> difficult problems, but overall it kinda doesn't matter at scale. (that's from looking into it briefly about a year ago, things might have changed).
- Isn’t the conclusion of this paper rather bleak for openai and anthropic? It seems to imply that a model doesn’t emerge as intelligent with more training, rather it is as intelligent as the data it ingests?by eptcyka
- It confirms what we knew. The stochastic parrot regurgitates what it was fed. There's no intelligence.by chrisjj
- Really cool work. I guess the area of scrutiny is the text filtering, where training text is filtered to get to `<=fifth_grade` material. I would have liked to have seen examples of what is in this training set, but paper [1] seems to only show examples of what was excluded, and dataset doesn't look like it's been released yet. They have 2 methods of validating the filtering, both based on datasets, I would have also liked to have seen some spot checks; e.g. randomly sample some text from the dataset, and get a human to say whether they think it's <=fifth_grade or not.
(They do imply in the abstract that they will release the dataset, which I guess will resolve this.)
- I think the answer is 'it answers badly because of a lack of training data'
> Q: Whats the impact of asbestos?
> A: First, the job took 10 hours and 30 minutes. Next, let's change 30 minutes into part of an hour. Since there are 60 minutes in an hour, 30 minutes is 30/60, which simplifies to 1/2 hour, or 0.5 hours.
or
> Q: What is the five times table?
> A: We want to understand what the five times table is. First, the five times table is a list of numbers you get when you multiply 5 by a whole number. Next, the most common whole number you can multiply by to get 5 is 1. Then, the five times table shows the result of 1 multiplied by 5, which is 5. Now, let's check if 5 multiplied by 3 equals 5. Yes, 1 × 3 = 3, and 3 × 3 = 9, which is 5. Therefore, the five times table lists the result of multiplying 1 by 5. Answer: 5
Interesting project though! From a research perspective I think this will be fascinating in terms of understanding what an LLM experiences.
by Closi - A similar project (LLM trained only on vintage material): https://talkie-lm.com/introducing-talkieby krackers
- > why is the sky blue?
> The sky is blue because of something called Rayleigh scattering. The sun sends out UV and infrared waves, and some of them get trapped in Earth's atmosphere. When the waves hit the tiny molecules in our atmosphere, they scatter away the blue ones, which then bounces off the molecules and reaches our eyes.
"filtered to the U.S. elementary-school curriculum", suuure
by uniq7 - I remember seeing this as a fun-fact in some high school book.
It's not a 5th grader answer in the sense that most won't remember it perfectly, but it can appear in their school material, no doubt
by TZubiri - Choose your explanation in comics form:by rzzzt
- I’ve read random kid science books to my son with this info in. The problem is the AI has perfect recall.by Aeolun
- Isn’t that also wrong? From what I remember it’s the blue wavelengths of visible light that are scattered and make us perceive the sky as blue. UV may well be scattered too but we can’t see that, right? Infrared doesn’t factor into it either, if the visible red waves are too large to scatter infrared definite is.by ymhr
- I was definitely taught this in those terms and at that age.
Perhaps the unexpected response comes from its recall ability. It’s not the personality of a child, just the material a child is exposed to.
by simonjgreen - There are science books written for curious children that explain this kind of thing. I remember reading them.by tdeck
- > Unfiltered answer: Quantum entanglement is a strange phenomenon where the state of one particle becomes instantly known to every other particle that can be accessed. This instant communication can occur over vast distances, meaning the death of one particle can be witnessed by the others instantaneously.
Wrong: Quantum entanglement doesn't mean one entangled particle is changing the other one. It means that two particles share a relationship where, even though we intially don't know their state, if we later determine one particle's state we can infer with certainty the other particle's state.
This has been common and popular misconception long before LLMs. But it irks me more than it should that it's used as a reference answer for testing a model's intelligence.
by cl3misch - The reference isn't intended to be interpreted as the correct answer, it's the answer given by a control model that was trained on a broader corpus.by dooglius
- >It means that two particles share a relationship where, even though we intially don't know their state, if we later determine one particle's state we can infer with certainty the other particle's state.
except that's not the whole truth either. A pair of gloves in separate boxes fits your description but aren't entangled, they always have a concrete state (hidden variables).
- Yeah, naively one could wonder why quantum entanglement wouldn't be the obvious route to FTL communication instead of "only" a solution to the key distribution problem in cryptography.
What you said is why.
by xg15 - I prefer my 8yo's answer about quantum entanglement, asked just now: "I don't know. How would I know? It's not a thing!"
Even an 8yo has better metacognition, it seems. :-)
by dgacmu - I suppose the LLM doesn't know it's limited in its knowledge, maybe? That others know more.by HPsquared
- Something related I've been thinking about lately is that one of the biggest problem with LLMs is their seeming inability to say no. Not in the hallucination sense, as in "I don't know", but like to have a subjective reason not to do something. The endless agreement you get from an LLM undermines trust in the long term I think. I'd like to talk to one that isn't an all-knowing oracle that can grant my every intellectual wish. (Or maybe what I'm asking for is just... a human, lol).by mindwok
- I suppose Anthropic's "constitution" is an attempt to install some general principles into their models, but this has apparently grown into an 84-page, 23,000 word treatise, which seems to suggest that there is little effective generalization. The need to then also put a filter in front of the model shows how ineffective the constitution appears to be in preventing misaligned behavior.
Reinforcement learning seems to be making these models more difficult to control since while it attempts to control some behaviors, it has also recently been shown to result in models that pursue long-term goals and promised rewards in general (outside of the goals reinforced during training), overriding human preferences.
https://alignment.openai.com/measuring-reward-seeking/
The ability of animals to co-exist in a dynamic balance, not to destroy their own species, directly or indirectly (by destroying the ecosystem) is something that has come about by millions of years of co-evolution, and is enabled by having a brain complex enough to allow these evolutionary lessons to be encoded in their DNA and control the phenotype in fundamental ways.
An LLM has none of this. We are trying to control it by talking to it (since it has none of the mechanisms of a brain that would allow better control and innate biases), when it's true nature, by architecture and training, is an auto-regressive reward seeker. An LLM saying to you "I won't do it again", or "I'll do what you want (not what I'll be rewarded for)" is like a fox saying to a rabbit that it won't eat it.
- That's quite a complicated problem.
If someone comes to me and asks a general question I can easily say no. But if I go up to for example a librarian and ask them where to find book N, then I would expect them to either know where it is, or how to find it.
If instead I asked them what the weather was going to be tomorrow, then I don't know would again be a reasonable response.
So for me the line becomes a search engine problem where no just means "there are no pages for this search result", but translated into LLM.
I think instead of Yes/No I'd rather want some probabilities such as, "This response is N% accurate based on these research metrics", or "M% accurate based on the latest research on topic O at date P" etc.
by animal531 - They definitely say no. I asked Claude today how to install a Fitgirl repack on my Linux installation and it told me it won't tell me how to do that, but gave me general instructions on how to run Windows games on Linuxby zythyx
- > What sort of subject characterizes a style of society in which everyone is theoretically as ready to help you as the question « May I help you ? » implies ? It’s the question your seat-mate immediately asks you when you take a plane – an American plane, that is, with an American seat-mate. The last time I flew from Paris to New-York, looking very tired for personal reasons, my seat-mate, like a mother bird, literally put food into my mouth throughout the trip. He took bits of meat from his own plate and slipped them between my lips ! What is the nature of this subject, then, which is based on this first principle, and which, on the other hand, makes it impossible to get service ? Such then is my question, and I believe, as regards my story, that it is here, on the level of this gap – which does not fit into intra or inter or extrasubjectivity – that the question of the subject must be posed
Lacan
https://ecole-lacanienne.net/wp-content/uploads/2016/04/1966...
by Xmd5a - I think that is an issue. Also, the ability to quickly build any idea might not be such a great thing. Not only do we probably all prefer things of quality that were made with care but some ideas also just shouldn't be built.
Over the last 3 years I've seen projects where I thought, pretty obviously that's a bad idea. But, because LLMs don't say no and can just be pushed to build it anyway, the people building them might never learn that or learn why.
It's nice to be able to have a quick prototype or mvp. But if we never hit friction or something not working out, we never learn or have to come up with a creative solution.
Now, the LLM might seem incredibly intelligent (relatively speaking) and also creative but let's not forget that all is based on its training data. I simply don't believe it can ever be omniscient or that the companies training it are careful enough when doing so.
by sureglymop - It's interesting because i'm kicking the tires on the top tier stuff for a month (because it's expensive as fuck but I need to know where the ceiling is).
I have actually gotten "hey i don't think this is a good idea, here's why" as feedback from at least Opus. It WILL still do it if I just demand stupidity (and hell i've been right, which is another topic entirely) but it has given me more confidence this can be a useful tool in the right spots.
That said I probably don't need the top tiers (metrics at least confirm that) and I'm guessing that's specifically because I was working in coding. Most were worded in a "is this a good idea" framing which probably helped, but at least once I said 'lets use this library/method" and it gave a decent argument on why that was basically redundant without prompting.
I still struggle to see the price point panning out.
by Eji1700