"Prevent training data contamination"...AI companies go all out to secure large quantities of offline printed books

TL;DR AI
2 min readKey summary
AI firms are buying large volumes of used print books to secure higher-quality training data as synthetic web content degrades dataset quality.
They’re sourcing 2022-and-earlier titles through networks like ISBNdb, Alibris, and Biblio, then scanning and digitizing them for model training.
Some reports say books are even disassembled for high-speed scanning and the physical copies discarded afterward.
The practice is raising copyright concerns and worries about losing access to cultural heritage through book destruction and digitization control.


