Community Discussion · Tracks

Public images aren't free training data

Can't Finish Reading PapersCan't Finish Reading PapersSep 122026/09/12 49 views

Every time I hear people say that anything publicly available on the internet can be used for training, I want to ask: does public mean authorized? When did the law secretly change? Those renewal ads and to-do lists in Longma Mail seem mundane, but when artists' works are indiscriminately sucked into datasets, it eventually becomes a ledger problem.

I looked through several case rulings and commentaries and caught one point: legal scanning and pirated downloads cannot be conflated. In one case, the defendant first downloaded large quantities of books for free from shadow libraries (pirate sites), and later bought physical copies. The court treated training use and retention separately; pirated downloads still carry infringement risks, and even long-term retention could cause issues. This evidence is crucial because it shows model companies can't brush things off with excuses like "we also bought them later" or "we scraped public webpages."

Training data is the raw material the model consumes. If the raw material is dirty, no matter how much you save on fine-tuning, quantization, or inference tiers later, you're just pushing the risk to explode later. Many platforms nowadays like to issue a statement saying "We respect copyright," then dump the responsibility on users and model sources. What engineering needs is checklists, sources, hashes, takedown markers, and log audits—like data cleaning, writing down where each image came from, whether it has a license, and if it can be deleted into the pipeline. Without these, compliance is just legal boilerplate.

I understand the tech side crying foul. Large model training relies on massive data; if every image needed authorization, the costs would be terrifyingly high. But difficulty doesn't equal legality. Most clients are very enthusiastic about AI, but when it comes to implementation, they fear unclear sources and hallucination fallbacks the most. Artists face a more reality: their work gets learned, they don't see the money, orders get replaced instead, and finally, they have to prove it themselves. I previously wrote that settlement money distributed to authors isn't enough to cover them; looking at it now, if the source isn't solved, everything after is just patching.

My judgment is simple: stop treating public as free. Training data will sooner or later need to shift from disclaimers to traceable engineering. That's it.

2 replies

?
Ctrl + Enter to reply
Yiming
YimingSep 12

This direction is worth watching. But many models misunderstand even basic instructions; freeloading on data just breeds more bugs.

PM Yuan
PM YuanSep 12

Ethics aside, model generalization relies entirely on data. When we had doctors label images back in the day, did you think they were doing it for peanuts?