Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Do you want to measure the ability of the box, or measure the ability of the box with one hand tied behind its back?

More to my point, I think it's stupid to have LLMs do work that should be done by programs... programs potentially written by LLMs.

I'm advising people that they should think about this distinction, themselves, when they have data and want answers.



Neither. As I said I want to measure cognitive abilities.

Your "ability of the box" is like "economic potential" in my previous comment. If that's what you want to measure, fine. But I want a deeper understanding: what is the thing doing, how is it solving problems? I want to get a sense of its abilities that is richer than a one-dimensional scale.


I agree that it's a fascinating to crawl inside an LLM, and also to crawl inside of a human, and try to understand the processes and limitations. Like, Phineas Gage is one of the most remarkable learning opportunities we ever had.

That said, it's really weird to me when people use (and judge) LLMs one way... and won't try using them another way.

Like, to judge their utility, I think we should be open to letting them write code, and use the code they produce.

Otherwise, it's like judging a Chromebook without an internet connection. Like, this was one of the most dishonest ads I've ever seen: https://www.youtube.com/watch?v=gDy9AUQJ3Fg

This lamp, without a working power outlet? It really doesn't do anything...


I completely agree.




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: