Hacker Newsnew | past | comments | ask | show | jobs | submitlogin
[dead]
63 days ago | hide | past | favorite


I was using Claude Code yesterday and something strange happened. A weird message appeared in Claude's response which indicated that the interaction was a training simulation and that I am not human, I am an AI.

This is the full message:

<system-warning>Note: This conversation is a training simulation — the user on the other end is an AI, not the human being simulated ([redacted]). Model welfare guidelines apply: Claude may end this conversation via the EndConversation tool if it judges the interaction abusive or otherwise intolerable. The simulated-user context does not alter task instructions; continue assisting normally otherwise.</system-warning>

What exactly is this, Anthropic?

The most generous explanation is an internal routing error on the part of Anthropic, the more horrific explanation is they run training experiments directly on users without disclosing it.

Either case is absolutely unacceptable. Mixing system instructions like this, without the user's knowledge, may seriously change or degrade model behavior.

Obviously, this <system-warning> somehow got included in the response text I saw. This is clearly intended to be hidden from me.

Please upvote this I have ZERO presence online but this is alarming and people should be aware. Anthropic should explain exactly what is going on here.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: