AI developers are quietly buying up physical books in bulk, slicing them apart, and scanning every page into training datasets. After the scanning is done, the original books are thrown away.
Why physical books?
The practice is driven by a simple need: text. AI models require enormous amounts of written material to learn language patterns, facts, and reasoning. Digital copies of many books are locked behind paywalls, limited by licensing agreements, or simply not available in the quantities developers want. Physical books, by contrast, can be bought secondhand or in remainder lots without those restrictions.
Workers cut the spines off with industrial paper cutters, then feed the loose pages through high-speed document scanners. The process turns a shelf of books into a stream of digital text in hours.
What happens to the books after scanning
Once the pages are digitized, the physical copies are discarded. Some go into recycling bins. Others end up in dumpsters. The books are treated as disposable — a raw material that has served its purpose.
This is not a small operation. Developers are buying books by the pallet, according to people familiar with the practice. The scale means thousands of books are being destroyed every month to feed AI training pipelines.
The cost of data hunger
The approach raises questions about waste. Books are durable objects meant to last. Slicing them up and throwing them away after a single use is a stark contrast to the way libraries and collectors treat them. But for AI developers, the priority is speed and volume. Buying used books is often cheaper than licensing digital rights, and scanning is faster than negotiating with publishers.
It's unclear how many books have been destroyed this way. No central registry tracks the practice. What is clear is that the demand for training data is not slowing down. As AI models grow larger, the need for text grows with them.
For now, the books keep coming — and keep being thrown away.




