AI Isn't Fair Use

by allsparkinfinite on 2025-05-24

Gobble Gobble Gobble

AI companies train their models on publicly available data. Not necessarily data under a public licence, just data that is publicly available. This includes both scientific and artistic works. Creators have continuously complained about this. Sarah Silverman, Ta-Nehisi Coates, Laura Lippman, and Paul Tremblay are among a group of over 10 writers filing a lawsuit against OpenAI for training its models on their books without "consent, credit, or compensation". YouTube itself has confirmed that if OpenAI had used YouTube videos to train Sora, it would have been against the platform's terms of service. AI companies claim that training AI counts as Fair Use. AI models are only as good as the data they are trained on, and this is why AI companies will gather as much data as they can get away with.
OpenAI has complained about DeepSeek using OpenAI's models and its outputs to train DeepSeek LLMs. This is devastating to OpenAI as DeepSeek is priced much lower. This is obviously hypocritical, as all the same defences that OpenAI has used can be applied here.
Karma's a bitch.
Microsoft's security researchers have observed DeepSeek-affiliated individuals "siphoning" large amounts of data using OpenAI's API, which is such entitled phrasing. When they are using pathways implemented by you to access data that you intend to be accessed in such a way, you cannot be pissed about their access to your product. You can, of course, complain about how they use it, but then we get back to the question of hypocrisy.

Google has considered using user data stored in free Google Drive accounts to augment its training datasets. "If you aren't a customer, you're the product" in its full glory.
Meta considered buying the publishing house Simon & Schuster.
Photobucket used to be the image hosting provider for MySpace and Friendster, and it could be sold to an AI company as well.

Report By US Copyright Office

The US Copyright Office has released a report exploring copyright laws and artificial intelligence, and AI companies are not gonna be happy.
While each case is unique, there are some patterns that the report has highlighted. When a model is used for analysis or research, its outputs are unlikely to take money away from artists. However, when the outputs are used commercially, they commercially compete with the artists whose works they are trained on. Especially when the training data is accessed illegally, this goes beyond the precedent established for fair use boundaries.
Analysis can include tasks such as content moderation systems. These can be considered transformative because it fulfils a different role from that of the training data. However, when the model creates outputs that can be used as training data again - that is less likely to be considered transformative.

Computer programs have been copied to access their internal elements, and then using those learnings to create new, interoperable works. This is not true for artistic works. Unless the original work is being parodied, precedent does not classify this "inspiration" as transformative. The former is the removal of a technical barrier in pursuit of productive competition, while the second isn't.
The report also rejected two common arguments about the transformative nature of AI training. The fact that the act of training is not a creative endeavour does not mean that it is transformative - it also depends what the training is used for. Training an AI is also not like human learning, in the view of the US Copyright Office.

According to a spokesperson, the White House fired the Director of the US Copyright Office one day after the report was released.
Surely this is unrelated to the fact that Elon Musk, besties with Donald Trump, has an AI company of his own.