Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • I used to think that way about SVGBench, after all labs can just train on the test set, right? It turns out the task was highly generalizable. Try designing a logo and you quickly see the gap between models visually. Even though there is still a gap between "shiny demo SVG" and actual real-world use.

    Same thing happened with MineBench, basically SVGBench+3D, until that got "saturated".

    Remember spinning hexagon bench[1]? Or the AI World Clocks[2]? Yeah, that used to be hard for frontier models.

    Creating games is the next iteration that still has some signal left. Assets + Game logic + UI + Sound, it let's you assess a model's "taste" very quickly.

    What else is left, once all these benchmarks get saturated?

    [1] https://x.com/flavioAd/status/1885449107436679394

    [2] https://news.ycombinator.com/item?id=45930151

  • I use Minecraft (and Warcraft and an arcade flying simulator) as a silly but directionally correct indication of the models' capabilitites.

    Compare Astra[0] with GPT 5.4[1] which was OpenAI's state of the art just six months ago.

    (all tests on more models with code and prompts available here: https://senko.net/vibecode-bench )

    Yes, it's not a scientific benchmark but it's a good heuristic.

    For a better eval, create a one-page prompt / mini spec related to whatever you're using the LLMs for, and see how well a particular one works for what's important to you.

    0: https://senko.net/vibecode-bench/2026/rts-gpt-6-astra.html

    1: https://senko.net/vibecode-bench/2026/rts-gpt-5.4.html

  • Very low quality content. Unsurprisingly pangram 100%.
  • Right answer wrong thesis.

    The example of 'labs are optimizing for this' is the wrong thesis.

    AI arbitrarily generating something of 'apparent sophistication' is not that hard - being able to produce it to spec that has invariable vague elements - and then being able to rationally modify it is the problem.

    Look at image gen: you press the magic button and get a 'Pelican on a Bike' - but you can't just change the 'hat' of the Pelican. You have to press the magic button again, and you get a whole different Pelican on a Bike.

    This is the fundamental conceit.

    It's akin to the conceit that 'writing the code is the work' - but it's not - it's the research, the design, docs, integration and all that 'know-how' that has to 'sit somewhere' so the shape of the thing output can be adapted and moulded.

    Arbitrary code output is not quite worthless but almost, it's the the '80% that means another 80% and then anther 80% to go'. It's like a nicer stating point.

    The research capabilities of the AI, which don't make for nice demos, are arguably more powerful.

  • The benchmark I want to see people adopt is:

    Build a flowsheet based steady state chemical process simulator, then use it to simulate and optimize a full scale oil refinery.

    1) Building a solver engine that works at this scale is not a trivial problem, and the successful ones rely more on heuristics than some categorically different solution approach.

    2) Defining the engineering equations relevant to this task is relies on understanding what level of fidelity is required to answer the questions people ask of steady state process models.

    3) Knowing the thermophysical properties of chemicals and crude oils is possible from the open literature, but the information is diffuse and different correlations are applicable in different situations.

    4) Creating a GUI which converts a flowsheet into matrix math is non-trivial, although a sequential modular approach is a bit easier.

    5) Defining large scale models in such a way that they solve robustly is as much art as science. For example, completely closed recycle loops like refrigeration systems are a nightmare for solvers, so it is often better to define them in an open-loop way.

    6) Optimization involves knowing the relevant commodity prices, but more importantly how to define the constraints on the model so it doesn't just say to produce infinite gasoline.

    7) Troubleshooting the inevitable convergence failures is also as much art as science. There are a large number of diagnostic techniques, but fundamentally you need to be able to relate what is happening during the solver iterations with the intent of your model because more often than not the problem is that you've asserted something impossible, redundant, or irrelevant.

  • This sounds unconvincing, because a) pelican test is subjective, there's simply nothing to leak as it has no available direct answers and maybe an extremely faint preference signal, and b) the same small models actually do perform well when you change the subject. Some models are genuinely trained to be better at some domain, such as 2D layouts or vector graphics in this case. It all depends on particular recipes and datasets. Which is the actual reason these tests are poor as vibe checks: they don't do anything to disentangle generalization, memorization, and training preference. One-shotting popular software in particular is definitely not a good test of anything as memorization is going to dominate it.

    AAII is also not very useful, neither is any generic score/benchmark. If you want a weather forecast you aren't looking at the average temperature of Earth.

    (actually when did the term "one-shot" get hijacked to mean something other than "one example"?..)

  • > That’s the problem, these tests can’t tell you how good a model is anymore because it’s trivial for labs to optimise for exactly these tests by the next release.

    But they don't prove the claim. Are the models amazing at recreating Minecraft, but the second you swap the word Minecraft out with another game or a custom game, it shits the bed?

    That's not what I see. My feed is full of people using Astra to recreate all sorts of games from Diablo to some random idea they came up with, in ridiculously polished detail like animations that would have taken me weeks of iteration in gpt-5.6-sol but it was a single shot by Astra.

  • I keep seeing Astra make beautiful 3d stuff online, yet when I feed it some old school RuneScape assets (even tried with some very detailed guidelines) and asked it to generate some new plausible assets it failed horribly.

    I think there’s still something really off with current (frontier) models when it comes to creating “novel” stuff? Even 2004 style graphics..

    Or am promoting it wrong?

Explore Birbla archives

Recreating Minecraft Is Not a Benchmark · Birbla