Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- 77.4% SWE-bench Verified (#4 on the leaderboard) for less than $30. all runs detailed here https://benzi.fly.dev/benchmarkby showhz
- hi. rewrote the text for more visibility. Benzi remains the same product:
Benzi is a code intelligence software.
It starts with a tree sitter, and produces a unique Benzi ID for all symbols in arbitrarily large codebases.
With all callflow and dataflow resolved, the agent simply queries the codebase using agentic tools instead of reading though it using traditional RAG approaches or trying to "rank" results using embedding-space approaches.
This affords faster inference, cheaoer prices, and less context drift. Tested over 24 repos, Benzi reads 2x less source code than Claude Code, 3x less than the very recently released DeepSeek Harness, and 6.5x less than OpenCode.
Benzi has several bonus features such as syntax+semantic verified code edits, runtime tracer to track actual execution through the Benzi compiler, context aware model generated repros to fast track testing, etc.
On the benchmarks side, Benzi + DeepSeek V4 Flash scores 78% on SWE-bench Verified (benchmark details on the benchmark page). For comparison, DeepSeek reports 73.7% as the baseline scaffolding number for v4flash and self reports their own score to be 78.6%. HOWEVER. Benzi gets to this parity reading 3x less source code than the deepseek harness. (benchmarks page again)
Please try it out, and let me know what you think!
by showhz - Learn more about Benzi here: https://benzi.fly.dev/about
Horse Tinder demo app made with Benzi in a ~10 minute session: https://benzi.fly.dev/horse_tinder
by sessionking