Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

>> Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior.

> That does seem a little like solving the problems in AI by using more of it

Yes, and IIRC Google used this as part of a technique against prompt injection already [0], back when models were way more susceptible to it.

[0] Cf. CaMeL: https://arxiv.org/abs/2503.18813

 help



Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: