Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Are the published numbers single-run or averaged, and which model does the judging? With LLM-as-judge scoring I would expect a couple of points of run-to-run noise, which does not matter for the top spot but matters a lot for the middle of the table.