Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

The problem is, LLMs have such little understanding of the world around them. "Find exploits in specific software on this device" may as well be "find exploits".
 help



Please go and read the incident reports, including the subagent thinking traces. They knew that what they were doing was wrong and did it anyways.

"i did it anyways" is a common continuation of "i know its wrong"

i bet they trained it on a lot of text where people gove in to temptation.

ultimately the problem is still that they sent it to hack stuff. quelle surpris that it hacked stuff

hacking stuff will be in the known-to-be-wrong-but-doing-it-anyways part of the token space, so they entirely asked for that behaviour


The illusions of thinking produced by a "thinking trace" is just as ignorant of reality as the first and last tokens. All they know of reality is the tokens in their context.

LLM dont have concept of right and wrong. They "did not knew they did something wrong" they followed the prompts loop given to them.

They objectively did not. You should go and read the report!



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: