Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

For context, Nightshade is a tool used to alter images in such a way as to poison a generative image model if the altered images are included in the training set. Poisoned models produce noticably worse output and the only way to fix them is to identify and remove the poisoned training images, and retrain.

That said, it's kind of bonkers to me that publishing an image with alterations like this might be considered illegal, particularly if an image is being published by the copyright holder. It's like saying you can't watermark anything you publish.



> the only way to fix them...

This kind of claim is almost always suspect. There was a time that rotating or other manipulations caused issues for the models and those were largely taimed. Hinton's work with Capsule Networks [1] produced much more resilient models, though they are computationally much more intensive.

It's also interesting that this is coming from Japan, where they also have a law regarding copyright and training that is very permissive. [2]

> That said, it's kind of bonkers to me that publishing an image with alterations like this might be considered illegal, particularly if an image is being published by the copyright holder. It's like saying you can't watermark anything you publish.

This should be generally true, but in law, intent typically matters. If you are doing something with the primary reason being to cause harm, then normal protections often don't apply.

[1] https://blog.paperspace.com/capsule-networks/

[2] https://finance.yahoo.com/news/ai-art-wars-japan-says-185350...


"If you are doing something with the primary reason being to cause harm, then normal protections often don't apply."

This seems pretty confused. There is no law banning "harm" (A term so vague that everyone is guilty.) and even if there was watermarking images does not appear to be more "harmful" than making an AI using the copyrighted images of others without permission.


Good point on intent. I have to think if an artist published poisoned images, with a notice like, "don't use these for training, they're poisoned, contact me for licensing of clean images", it would be hard to argue either negligence or malicious intent to disrupt training.

Which would be distinct from one silently poisoning and reposting images everywhere with the goal of gumming up everyone's training.


I think the intent is what's most likely to cause trouble.


yup, that's why there is a difference between manslaughter and murder of varying degrees


Are these tools actually effective, or are they just "feel-good" things?


"Feel good."

The problem with nightshade is that while it worked in laboratory conditions, in practical application it wouldn't work without extensive coordination between everyone using it.

For example, if one artist who draws dogs uses it to bias towards cats and another uses it to bias towards horses, the bias data is less of a signal and more noise. In the paper, they biased all the images the same way.

The issue compounds when you consider the multiple data points that needs to be biased. An artist who draws impressionist dogs, an artist who draws cartoon dogs, and an artist who draws cartoon cats would need to bias 'impressionism,' 'cartoon style,' 'dogs,' and 'cats' all in ways that are standardized across nightshade users to have the tool be effective.

This isn't realistically going to happen.

So ultimately it's about as effective as the users who put the "you don't have permission to use my data" clauses on their MySpace two decades ago. Feels good, and totally ineffective.


This article has some examples of output from poisoned vs clean models.

https://arstechnica.com/information-technology/2023/10/unive...


Interesting, but that doesn't really answer my question. Has anyone done wider studies about the effectiveness of this sort of thing generally? As in, does it have to be tailored to a specific model or does it offer generalized protection with all models?

The Ars article doesn't talk about this aspect.

If these methods are effective, can a similar thing be done with text?


I can't speak to broader study of this, but I suspect this is only doable with images because you can make meaningful changes to the pixel data that remain mostly imperceptible to the eye. I think there's less room for this sort of thing in text.

Happy to be proven wrong, though!


It depends. But poisoned images fed into a vulnerable training setup will scramble the output and make the model pretty useless.


This is also a good way to prevent your soup from being stolen out of the restaurant kitchen... just add a few quarts of raw sewage.

I don't see why that would ever be illegal, and it stops the thieves dead in their tracks.


That's a terrible analogy as it involves someones physical harm and death. This just makes AI images not look as good.


It's a terrible practice to pollute art that is owned by the public and ultimately belongs in the public domain. Are they going to come back 75 years later and filter out their data sewage?

I can go get a course of antibiotics if I have to eat shit soup, but ruining the data ruins it forever.


It's the artists own work, they can do whatever they want with it. If people don't like that they should be filtering this out of their dataset before training.


It's more like putting laxative in your sandwich because Steve from accounting keeps stealing it from the office fridge.

A watermark isn't a booby trap. This is.


The Japanese Government knows that it is their job to further the public good, and that LLM AIs will help the public good. Therefore any attacks on the development of LLM AIs should be illegal.


I love that Japan is doing this, I hope their efforts stop all the luddites in the west from hampering AI development.


I can't tell if this is satire or supplication to a patently goofy premise.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: