Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Well, we can't even get biases out of humans in any reliable way. Maybe we will do better with machines.


We won't, because the machines are trained on data created by humans.

We won't, because the machine results people want will be those that appeal to their biases.


Exactly. And it’s not even a new phenomenon, we know our algorithms are biased.

https://www.nature.com/articles/d41586-019-03228-6

https://www.technologyreview.com/2020/07/17/1005396/predicti...


> We won't, because the machine results people want will be those that appeal to their biases.

When you see an idea pushed/accelerated to an absurd conclusion, you might more easily see what's wrong with it.


You won’t. That’s why Poe’s law is a thing.

https://en.wikipedia.org/wiki/Poe%27s_law


Possibly, possibly. But those extremes are not what you get from an effective bias appeal.


I believe GPT-4 pre-rlhf was much more accurate overall.

If you finetune it in formats that humans likes, it actually gains similar biases.

I believe the ‘sparks of AGI’ talks about this, where models can much more accurately predict the probability of events than humans.

After RLHF, it mimicks human bias. So it might be that we can create these models, but we just don’t like them/like to use them.


RLHF = nerfing and censorship

What, did you think the plebs would get access to the real deal?


If so, we irrational humans will just reject the validity of machines.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: