hey HN, sharing harness engineering deep dive on tool calling repairs for open models. i've been thinking about why "open model bad at tool calling" is almost always a harness problem, not a model problem. spent time looking at billions of tokens from DeepSeek (and other open models) in our coding agent. ended up building a tool-input repair layer on top of Zod. by the end, DeepSeek V4 Pro was beating Opus 4.7 in 6/10 of our internal evals. the main things that helped:
- most failures came from a small set of recurring schema mistakes
- switched from preprocess-then-validate to validate-then-repair
- handled some weird cases like markdown auto-links leaking into file paths
Is the output, that others can use, the research and findings, or is there a tool or something that came out of this that I can plug in to one of the common harnesses?
I guess that brief note in the twit is that it'll be part of that harness you're opening up in the future?
Thanks! I've seen many developers use the entire tweet as a prompt to improve their harness. If you trace errors, you can find tool call failures per billion tokens. I've seen like many versions of this on pi, oc, or whatnot. This is already live in our hardness command code for almost two months now and we are now fixing one million tool repairs per one trillion tokens.
- most failures came from a small set of recurring schema mistakes - switched from preprocess-then-validate to validate-then-repair - handled some weird cases like markdown auto-links leaking into file paths
full writeup: https://x.com/MrAhmadAwais/status/2050956678502420612
video version (more detailed): https://www.youtube.com/watch?v=f61DCDwvFis