Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

presumably that's a safety evaluation not a training setting


The whole Huggingface attack happened during training runs


part of it did. I was just replying to the question about why they would ever push the model to evade monitoring. surely that's an eval thing not a training thing.


No it happened during an ExploitBench eval. But I believe the same model already cheated during training which wasn't detected until later.


Ah yes it was that a model in training found the Artifactory board, which was then more fully exploited during the ExploitGym eval


Ah, ExploitGym. Not ExploitBench.




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: