

Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- Hi HN, I'm the first author, ask me anything. I've done AI research before but this is our first time building an agent benchmark so lots of lessons learned. If you're curious about AI docs agents or curious about the benchmark itself, ask away!by frances-liu
- The winning model (Cloud Agent) is Promptless, and in very small grey text, you might be able to spot a "Built by Promptless" on the page if you look closely.
While I'm sure this benchmark is one worth optimizing against, and some real thought was put into it - it also seems possible that the rubric was specifically designed to be one that helps promote its creator. If one runs a contest, it's a dangerous situation to also have an entrant. Either the contest can be favored for the entrant, or the entrant can have special training privileges. To put it another way, it can't be assumed you'll have a fair trial when your judge is also your uncle.
Not saying that this has happened here, but third party, unaffiliated benchmarks are always preferred for this reason.
by yathern