Switch language한국어
Back to the list

"Prevent training data contamination"...AI companies go all out to secure large quantities of offline printed books

TL;DR AI

Key summary

2 min read
  1. AI firms are buying large volumes of used print books to secure higher-quality training data as synthetic web content degrades dataset quality.

  2. They’re sourcing 2022-and-earlier titles through networks like ISBNdb, Alibris, and Biblio, then scanning and digitizing them for model training.

  3. Some reports say books are even disassembled for high-speed scanning and the physical copies discarded afterward.

  4. The practice is raising copyright concerns and worries about losing access to cultural heritage through book destruction and digitization control.

Read the original