Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Average people cannot tell reliably isn’t an appropriate or interesting test tho, otherwise Eliza and markov models etc. the framing that matters is explicitly adversarial. Play like your life depends on it instead of rooting for the machine, and you can’t win?

One way to play the game is causals and counterfactuals where Humans perform at like 90%+. Models can get close to that, but want some causal cot harness, and until the routing problem is completely solved, then that will necessarily degrade performance elsewhere, say in understanding jokes or poetry.

Check out cladder benches and related, lookup roughly equivalent psych research on children, etc. Even too-good performance is a signal as well!

Certainly if you think about this stuff a bit, accept the adversarial by default framing, and play to actually win.. it’s crazy that we are going around saying this is not only solved but solved 10 years ago.

 help



I found a cladder example that hits 97.7% passing on that benchmark? And it's like an insanely small dumb model that hit it.

I don't think philosophy has any real value in assesment here. I'm bias but even before AI I thought it wasn't accurate model of how thought works and I think AI has reinforced that.

Reminds me of the 4 humors of medicine in medieval europe. It has some truth but it's not really accurate.




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: