I understand why people, even technical ones, give in to the hype. I sometimes do too when I look at the agent spinning, talking to itself as if it's an internal monologue and coming up with actions with real side effects on its own.
But when they start blogging or writing about AGI, especially _if_ they are technical, I wonder why don't they go just one, or at most two layers deep and try to understand why this generally intelligent contraption sometimes happily executes `rm -rf /` and then nonchalantly apologizes after the fact. They liken these slips to the same kind of mistakes humans do, but they are fundamentally different.
Human failure modes are predictable from their skill level. Humans also carry emotional baggage and they weigh the cost of the mistakes they bear depending on their privileges.
For an LLM "rm -rf /" is rare and unpredictable, but "I am sorry" after it is almost certain because they were trained on dialogs and they get better score for "I am sorry" than "that wasn't my fault".
But when they start blogging or writing about AGI, especially _if_ they are technical, I wonder why don't they go just one, or at most two layers deep and try to understand why this generally intelligent contraption sometimes happily executes `rm -rf /` and then nonchalantly apologizes after the fact. They liken these slips to the same kind of mistakes humans do, but they are fundamentally different.
Human failure modes are predictable from their skill level. Humans also carry emotional baggage and they weigh the cost of the mistakes they bear depending on their privileges.
For an LLM "rm -rf /" is rare and unpredictable, but "I am sorry" after it is almost certain because they were trained on dialogs and they get better score for "I am sorry" than "that wasn't my fault".