Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • > he trigger was watching deepseek-flash fail on the simplest /review run, every shellCommand and readFile call bouncing back with a raw zod issues blob, the model unable to recover because the error wasn't in a form it could read. by the end deepseek v4 pro was beating opus 4.7 6/10 times on our internal evals.

    I think this is why Xiaomi forked OpenCode and created their own agent hardness, to reduce friction between model and harness:

    https://github.com/XiaomiMiMo/MiMo-Code

    by bel8
  • hey HN, sharing harness engineering deep dive on tool calling repairs for open models. i've been thinking about why "open model bad at tool calling" is almost always a harness problem, not a model problem. spent time looking at billions of tokens from DeepSeek (and other open models) in our coding agent. ended up building a tool-input repair layer on top of Zod. by the end, DeepSeek V4 Pro was beating Opus 4.7 in 6/10 of our internal evals. the main things that helped:

    - most failures came from a small set of recurring schema mistakes - switched from preprocess-then-validate to validate-then-repair - handled some weird cases like markdown auto-links leaking into file paths

    full writeup: https://x.com/MrAhmadAwais/status/2050956678502420612

    video version (more detailed): https://www.youtube.com/watch?v=f61DCDwvFis

Explore Birbla archives