AI companies buy up and scan physical books to circumvent copyright when training models
AI companies buy older physical books, scan them to train models and then destroy the original. Unlike piracy, destroying a legally purchased book in this way is permissible according to the source.
According to the source, AI companies buy older physical books, scan them and use the resulting text to train AI models. They cite the fact that these books were written by humans as the reason, giving them an advantage over much of today's web content as a source of training data.
The key point is the legal distinction between different ways of obtaining content: while downloading books through piracy is illegal, destroying a book that a company legally purchased during the scanning process is legally permissible. According to the source, companies thus use this approach to sidestep the copyright risks associated with digitally copying protected texts.
Details can be found in the source article.
Why it matters
The approach offers AI companies a source of high-quality, human-written training data with lower legal risk than pirated digital copies of books. For authors and publishers, this means that even older printed works can be used to train models without their knowledge or consent, because the source describes destroying a physical copy as a legally distinct act from digital copying.
Relevant practical impact
What this means
For a business
According to the source, buying and scanning physical books and subsequently destroying them is a legal route to training data for AI companies, unlike digital piracy, which is illegal — companies can thus reduce the legal risk associated with obtaining content for training models.
Risks and compliance More business impacts →Check the original
Event sources
only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.