Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

hey HN, sharing harness engineering deep dive on tool calling repairs for open models. i've been thinking about why "open model bad at tool calling" is almost always a harness problem, not a model problem. spent time looking at billions of tokens from DeepSeek (and other open models) in our coding agent. ended up building a tool-input repair layer on top of Zod. by the end, DeepSeek V4 Pro was beating Opus 4.7 in 6/10 of our internal evals. the main things that helped:

- most failures came from a small set of recurring schema mistakes - switched from preprocess-then-validate to validate-then-repair - handled some weird cases like markdown auto-links leaking into file paths

full writeup: https://x.com/MrAhmadAwais/status/2050956678502420612

video version (more detailed): https://www.youtube.com/watch?v=f61DCDwvFis



Thanks! Nice findings.

Is the output, that others can use, the research and findings, or is there a tool or something that came out of this that I can plug in to one of the common harnesses?

I guess that brief note in the twit is that it'll be part of that harness you're opening up in the future?


Thanks! I've seen many developers use the entire tweet as a prompt to improve their harness. If you trace errors, you can find tool call failures per billion tokens. I've seen like many versions of this on pi, oc, or whatnot. This is already live in our hardness command code for almost two months now and we are now fixing one million tool repairs per one trillion tokens.

also decided to open source Command Code by v1 release, so hopefully end of the month. https://www.youtube.com/watch?v=laEzOCgtK6c (live stream on this)


here's a lil pi extension I made yesterday based on it https://gist.github.com/maxjustus/3dd4c83865b1c439e26fdd925b...


see that's what i was saying.

in v1 we'll have a super strong extension api — week out probably.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: