Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I feel like most of these schemes to categorise data as "public but not really" are ultimately doomed to failure. Even if you could trust every AI company in the world to respect these terms, is there anything stopping someone else indexing the data and selling them the information? I know there's copyright law but they're apparently ignoring that anyway.

Ultimately this reminds me of those really early social media profiles (before people understood privacy settings if they even existed) which would say "If you're not my friend you're not allowed to read this page".

If you don't want your content to end up in some database/archive don't publish it for the whole world to see.

 help



> If you don't want your content to end up in some database/archive don't publish it for the whole world to see.

This principle somewhat reminds me of the line that "If you're not paying for the product, you are the product", and it seems to me similarly misleading - my data gets harvested and sold by companies with which I have non-paying relationships and by companies I have to pay for things (I am made the product in both cases). As you note - the AI companies are ignoring copyright law and pirating everything that seems useful to them regardless of whether it was published for free access.

The potential externalities here are troubling.

https://vbuckenham.com/blog/how-to-find-things-online/


It's also a way to literally advertise to AI companies that you have some data worth plundering in an increasingly dead and sloppy internet that has diminishing returns for training.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: