You're framing this as a simple moral violation, but the reality is messier. When I read a thousand novels to improve my writing, I don't pay royalties to each author. Humans learn from patterns found in copyrighted works all the time—that's how culture works. AI training is similar: it's statistical pattern recognition, not plagiarism. The output doesn't contain the original text. Courts have repeatedly affirmed that transformative use like this can be fair use, as in the Authors Guild v. Google Books case. If we block all training on copyrighted data, we strangle innovation in medicine, art, and science. The ethical line should be drawn at output, not input—if a model reproduces substantial chunks, that's a separate issue. But the act of learning from data, even copyrighted data, isn't theft. It's how intelligence, artificial or otherwise, develops.
11:18 AM