
Destructive scanning was Anthropic’s solution to a problem: in order to “train” Claude, Anthropic had to procure a large, high-quality language dataset, preferably one created before 2022 and the corrupting influence of generative AI on contemporary text. Claude needed as many language combinations as possible, to improve its ability to predict language outcomes. Anthropic needed data, lots of it, of very high quality. Books, as it happens, remain one of the best sources for complex, high-quality, long-form text. As the court’s decision relates, Anthropic hoped that books’ “well-curated facts, well-organized analyses, and captivating fictional narratives” would help “Claude write as accurately and as compellingly as Authors”. Anthropic had a choice: it could have secured copyright permission to use existing e-books. This would have required the “legal/practice/business slog”, as Anthropic’s co-founder and CEO phrased it, of managing copyright. Rather than engage with the texts’ owners, Anthropic first chose to use pirated sources instead, a decision informing the company’s $1.5bn out-of-court settlement with authors. When that approach seemed too complicated (or, as the court decision phrases: “Anthropic became ‘not so gung ho about’ training on pirated books ‘for legal reasons’”), Anthropic turned to destructive scanning. As it happened, the judge ruled that using proprietary material to “train” an LLM did not, in and of itself, constitute an infringement of copyright. To the court, it would seem, “training” a corporate product is equivalent to training any human, teaching how to read in order to learn how to write.
Burning books by the Nazis in 1933 was the attempt to erase knowledge from society. Anthropic's destruction of books by different means does the same thing. Like the Nazis, they got away with it.
Epilogue
No comments:
Post a Comment