Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I found that “harmful content” to actually be content that could cause bad PR for the LLM company.

I doubt someone could actually make counterfeit money following a LLMs instructions, because the model only generates plausible sounding text. It likely doesn’t have technical details of banknote manufacturing in its training data. The example given was laughable. Get paper. print money. Launder it.



Yes, but its a proof of concept. It demonstrates that this type of attack can work. Whatever the reason is for (the LLM company) blocking the ("harmful") responses, this is a potential/plausible way to circumvent that blocking.

The second example in the article is much closer to being a realistically "harmful" response, I think.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: