Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • > Also, I’m not sure whether “general reasoning” even exists in the first place? Maybe humans are specialised too

    I have been wondering the same. We are now exposed to so many stimuli, we are tricked into thinking this is the norm - to have a reasonable understanding about everything, unless specialization is called for.

  • > Ban offline training/pretraining. Models must train from scratch after submission Previously this was considered impossible so rule. My model shows this is possible Guarantees no synthetic data can be used It makes the comparison fair across differet models. Otherwise some models like LLMs can benchmaxx ARC by using ungodly amounts of offline training. (Since the benchmark has been around a long time, many ARC-like datasets have been created)

    I'm not an ML researcher, so YMMV, but... how could a model learn to answer these ARC-AGI questions without training beforehand?

  • > I agree that its rare to see to face problem sets in real life where every problem is given at once. Even if it is (like an exam), humans can usually only attempt one at a time

    Just one small snippet that I thought was interesting. I would always read through ~the entire exam before starting. Both so that I could find the problems most approachable to me, but also because sometimes it helps me figure out the rest of the questions :-)

  • Even cooler is his about me mention of saving his own life https://mvakde.github.io/ > Saved myself in a medical emergency (doctors didn't know what rhabdomyolysis was)
  • Sounds like a good day to be you, top 5 on Kaggle with a publication like this. It seems like you will be on a plane to SF shortly
  • I think(?) you’ve already probably done a good job of explaining this criticism for semi-informed people. But can you dumb it down even more for those of us who are almost entirely out-of-the-loop?

    > Training on the eval puzzles is cheating / “training on test”

    > No this is false. “Training on test” specifically means training on the labels of test data. The labels were not trained on.

    > Also, ARC is a metalearning benchmark, so you’re supposed to learn from the eval puzzles.

    > Jargon: ARC has a set of train puzzles and a set of eval puzzles. Each puzzle has example pairs and test pairs. A pair consists of an input grid + output grid.

    > The ARC, the label is only the test pair’s output grid in an eval puzzle.

    > These labels were not trained on. They are hidden. You can delete it beforehand if you wish

    I think what I gather here is that the test comes with one batch of training problems, which everyone agrees you can train on. But maybe the eval problems also come with input/output examples (to help define the problem) and training on those is controversial? I can’t see why it would be controversial but is that the criticism?

  • I think you’re asking the right questions, sample inefficiency is horrible in modern LLMs. Despite this, I saw your analysis:

    > The biggest increases in scores were due to

    Modern architecture (SwiGlu instead of GELU, RMSnorm not layernorm, etc.) More data diversity, better shuffling of data scaling up: 8 layers instead of 4

    This is commonly called squeezing the lemon and is usually a bit of a last resort. You should be able to achieve near SoTa with your new method, before you squeeze any lemons. This is, because the old SoTa is typically not using new optimisers and thus your results will be distorted by a large margin.

    In terms of sample efficiency I want to add two things:

    Runtime per-puzzle fine tuning is a very good target that provides a LOT of information. People have not looked at evolutionary methods to harness induction since the 90ies - if I was to work on ARC ever again I’m fairly certain this is where I’d look.

    Best of luck, padawan

  • Hi! Author here. Surprised to see this on HN now. Happy to answer any questions!

    Some context about this:

    - This is NOT an LLM. its a small ar transformer trained from scratch. One of the points was that extremely complex problems can be tackled without LLMs

    - Till the v1 of this result, this benchmark was only scaled by LLMs or their finetunes (ofc w enormous training costs). Other attempts performed okayish but used v complex architectures or extremely high amounts of training compute. No one expected a simple AR transformer to perform this well, at this low cost and w these few training samples.

    - Sample Efficiency is one of the most important unsolved problems today in AI. That's what I was targetting with this work. We know it is easy to increase SE by increasing compute/params, so it was important to constrain cost as much as possible (also why OpenAI's Parameter Golf had fixed compute and why Modded NanoGPT is considered very sample efficient)

    - Can the perf be improved? Yes but the competition is ongoing so can't talk about it

    - Personally I think today's frontier models can be beat by training from scratch. Haven't proved this yet tho

    - Fun: I was new to ML when I posted this first (dec '25). I basically used ARC as a way to learn ML

Explore Birbla archives