Discussion summary

A discussion about the Senior SWE-Bench open-source benchmark highlights concerns over subjectivity and industry standards in assessing engineering talent.

What the discussion says

  • Some users criticize the benchmark for being too subjective and narrow.
  • Others emphasize the importance of taste, elegance, and practical judgment in engineering.
  • Several comments point out the difficulty in objectively measuring engineering skill and level.
“Benchmarks like this are too subjective and narrow to be useful.”
— fiso64
“Bad code is bad code regardless of the scope of the feature.”
— iLoveOncall

Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • If something like this works wouldn't that imply technical interviews can be automated?
  • The value of a senior situation is to apply known solutions and strategies, to novel problems. I can not see how any benchmark, without ever changing, can provide a novel challenge for long.

    Any decent benchmark would use the whole of TRIZ to generate a giant ball of a problem first and watch a AI deduce a optimal solution.

  • It's nice to see a new public benchmark from Snorkel. They're doing some pretty sophisticated stuff over there.
  • Top solve rate is currently 24% with Opus 4.8... What's a competent human supposed to score?
  • This makes so much sense as to why I've always felt that Opus 4.8 was leagues ahead of GPT 5.5. It's so good at taking underspecified requirements and filling in the gaps with sensible approaches for your project
    by _345
  • I wonder how they're planning for the benchmark to stay relevant over time.

    If the benchmark is to implement features that are part of an open source project, and LLMs have those changes as part of their training dataset, it seems that they could just give a verbatim or slightly modified version of the change in their training data.

    And if one updates the benchmark to only incorporate code changes that are past the models knowledge cutoff, then the benchmark is less comparable over time, since the changes in the benchmark at time T and T+1 aren't the same.

    by jfim
  • Staff SWE Bench: LLM doubts whether we should do any of this, calls the entire project into question, refuses to merge code, but is happy to delete it.
  • I saw on Twitter that in an ML course at Tsinghua University, one of the tests asks students to write quizzes that fail the most LLM models as possible.

    What if we create a benchmark that works like this and assigns ELO scores? Models fight head-to-head by writing a question, a bug, or an incomplete implementation, which the opponent has to answer, fix, or finish.

Explore Birbla archives