Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

The checklist half automates cleanly, you can lint that every route calls authorize. It won't catch authorize being handed the wrong policy, which is the one that ships. And when I put a second model on review duty, the common failure is it agreeing with the first model's misreading of the spec, almost word for word. So I'd expect the forgetting class of bugs to mostly go away and the misread-the-spec class to sit exactly where it is.


LLMs trivially catch these things in review, and if you find agents agreeing incorrectly with each other (weak models?), maybe you need to be doing adversarial review.

It's pretty much solved, though I only use sota models.

I think for these convos, we need to see concrete fail cases so we can see what you're talking about and whether the truth matches up with the claim.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: