Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

This seems like a problem with the level of abstraction the user interface is working at. It is highly unlikely that in any real world task lasting less than one day there were really hundreds of distinct decisions that needed to be made by the user about appropriate actions to be taken by the agent/harness. It is also highly unlikely that the problem of decision fatigue seen here is somehow magically different to the same problem that countless UI designers had encountered and designed around in other systems long before harnesses running LLMs came along.

This is unfortunately the kind of result you get when you eliminate skilled and experienced people with real understanding of their field and replace them with repeated automatically-generated attempts to solve the same problem until something meeting some basic standard of correctness is found. It's as if the story of agentic AI as it exists today had been compressed into one perfect example of what it can do that is good but also why it's still fundamentally flawed.



I think we agree that there's an interface problem. But there's an intelligence-complete problem hiding underneath; which is why the first instinct was to recruit the human-in-the-loop in the first place. Turns out the human has one of those unintuitive failure modes that occurs when the system gets past a certain level of reliability.

Meanwhile, let's leave the hobby-horses in the closet for now. I won't comment on people's programming tool preferences.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: