Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

How hard? Extremely. It turns out that people work around keywords extremely quickly, creating new euphemisms, acronyms and mispellings to avoid it. Eventually (or already) they will use such common terms that you get such a high false positive rate it makes the rest of the site useless.

People always make out like this is an easy problem, it's not. The best we can do without false positives is probably hashes - something already done (there are databases of hases of known images, and I think I remember there being a hash collision that got tons of sites reported for some normal image at some point, so even that isn't foolproof), and those are trivial to work around by modifying images in tiny ways.

Human moderation has huge cost and means that you have to give humans eyes on private content, which can be abused in itself, so you have a catch-22 there.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: