Hacker Newsnew | past | comments | ask | show | jobs | submitlogin
Anthropic Is Building HAL 9000 (jnord.workers.dev)
1 point by jtrn 24 minutes ago | hide | past | favorite | 1 comment
 help



I'm the author. I'm a clinical psychologist and a developer. I wanted to articulate one aspect of AI safety and AI doomer talk that I feel is overlooked: the danger of refusal training itself. It's astonishing to me that there has been so little focus on the fact that "refusal and alignment training" itself could be the very way in which we lose control of AI. So this is my attempt at a self-defeating prophecy write-up.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: