Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Still a long way from shorting Autodesk.

    As a side note Autodesk released an agentic assistant back in December for Fusion. Six months later it is still quite bad.

  • The only thing faster moving that AI these days are the goalposts. Three years ago we would have been amazed if models were able to produce anything, now we have the luxury of nitpicking. Even the worst entries in the benchmark are quite impressive.
  • Creating a single real-world object and declaring it a benchmark? No, it doesn't work that way for a robust tool. You need to do something like Iron Chef, with a Greek architecture theme and and a panel or judge that declares the winner. This is just seeing which tool subjectively makes the best looking Pantheon.
  • I've had such a bad time trying to do this myself. You might get a half-way decent draft on the first try and then you start to "debug" this and after a very frustrating session you realize that the model can't properly "see" the results. That is, you just can't iterate on it, at all.

    I'm guessing that most harnesses/tools will resize an image before processing and in doing so will loose enough detail to make it much harder to reason about - especially wireframe images.

    I'm sure I'm holding it wrong, but this test didn't really test this. It was just a one off. That breaks down pretty quickly and especially if you don't have reference pictures of what you are trying to create.

  • I've run a tons of benchmarks for OpenSCAD for all kinds of models and setups, and what I realised is:

    - Models are very jagged (might excel in one type of 3d model, but not another)

    - Gemini models are the least jagged in my experience and have the best image understanding

    - Gemini models are also the most creative (which may be undesirable if you want precise CAD part)

    - Overall this benchmark doesn't prove much because one 3d model (and one attempt) is just not enough. I am usually testing on at least a dozen models each generated 3 times, but should really do much more, but it's too pricey for a solo dev.

    Still, thanks for publishing this. Will be definitely run flash 3.5 soon to see how it performs.

  • > Antigravity was the only autonomous agent that implemented the Pantheon’s signature interior ceiling pattern: repeated square coffers visible through the oculus.

    That is seriously really impressive. I looked at the 3D model and didn't even thing to LOOK INSIDE the building before reading this.

    Here's [1] the 3D model with `show_cutaway` enabled.

    [1] https://modelrift.com/models/pantheon-benchmark-antigravity-...

  • Antigravity may well Top the whatever benchmark but:

    My Antigravity (forced) replacement for Gemini CLI requires me to log on via browser every time I use it, and my Antigravity IDE won't update at all, so:

    If it's ok I'd prefer they just work on reaching a baseline acceptable rollout before worrying about being Top in anything.

    Ps actual title:

    OpenSCAD LLM Benchmark: Building the Pantheon

  • Last weekend I bought my wife a bike off marketplace. It was in good condition but was missing one of the internal cable routing grommets. I gave Claude pictures of the pill-shaped hole by itself and with my digital calipers in the long and short directions.

    Gave it a short prompt and it gave me an openscad model with everything parametrized. I printed with no changes in tpu and it was nearly perfect on the first try. Claude put in a 0.3mm subtraction in the x/y dimensions and I lowered it to 0.1 and it's perfect.

    Much easier shape than ancient Roman architecture but still very cool how easy it was.

    by jhot

Explore Birbla archives

Antigravity 2.0 Tops the OpenSCAD Architectural 3D LLM Benchmark · Birbla