Community Discussion · Tracks

Check if your book collection is targeted by AI companies with three lines of Python

Tian JiTian JiAug 22026/08/01 54 views

A friend recommended using Python to batch-check ISBNs, saying it lets you know if books on your shelf have been included in AI procurement lists. I initially thought this was sketchy—I just saw news last week about AI companies ordering via ISBNdb, buying thousands or even millions of copies at once, scanning them, and then shredding them. My few yellowed old books might have already turned into Claude's training data.

I searched directly, and ISBNdb had removed its AI procurement page, but the public book query interface still works. I spent half an hour writing a script and ran it through the 50 books on my shelf. The result: 18 were marked as "published before 2022," and 3 were out-of-print textbooks. Honestly, it's a bit complicated emotionally.

This tutorial walks you through screening your collection from scratch. No programming knowledge needed; just follow the steps, and it takes about 15 minutes.

Step 1: Prepare Your ISBN List

An ISBN is that 13-digit number on the spine or back cover (older ones are 10 digits). Every book has one, like an ID card. You need to list all the books you want to check, one ISBN per line.

Workflow:

1. Open Notes on your phone or Notepad on your computer.

2. Find each book and locate the ISBN (usually above the barcode). If you can't find it, use your phone's "Scan" feature to recognize the barcode, and the system will display the numbers.

3. Enter these numbers line by line and save as a books.txt file. Format as follows:

   9787544291163
   9787532768457
   9787208061644
   

Note: Do not add hyphens or spaces; pure digits only. If a book has an old 10-digit ISBN, add 978 and a check digit to make it 13 digits, but most old books can be queried directly anyway.

Pitfall Warning: I initially wrote book titles and authors in, but ISBNdb only recognizes ISBN numbers. So keep only the digits, one per line. If you have 100 books, manual entry is exhausting. I recommend using OCR software on your phone to photograph spines and auto-extract ISBNs; tested accuracy is over 90%.

Step 2: Install Python and Essential Tools

Python is a free programming language; anyone can do it. Windows users download the installer from the official site. macOS users open Terminal and type brew install python3 (install Homebrew first if you don't have it). Linux users come with python3, so skip this.

Open Terminal (cmd or PowerShell on Windows) and enter the following commands sequentially, pressing Enter after each:

pip install requests

Just wait for it to finish. If it says pip not found, try python3 -m pip install requests.

Common Beginner Error: Typo during installation, e.g., missing an 'l' in instal. I got stuck on this for five minutes. Copy-paste is safest.

Step 3: Write the Script and Run Queries

Create a new file named check_books.py, open it with Notepad, and copy-paste the following code:

for isbn in isbns:

url = f"https://openlibrary.org/api/books?bibkeys=ISBN:{isbn}&format=json&jscmd=data"

r = requests.get(url)

data = r.json

if f"ISBN:{isbn}" in data:

book = data[f"ISBN:{isbn}"]

title = book.get("title", "Unknown Title")

pub_date = book.get("publish_date", "Unknown Year")

subjects = book.get("subjects", [])

subject_str = ", ".join([s["name"] for s in subjects[:3]])

print(f"ISBN:{isbn} | {title} | {pub_date} | Subjects: {subject_str}")

else:

print(f"ISBN:{isbn} | Not Found")

time.sleep(1) # Avoid getting blocked due to fast requests


**Note**: This script uses OpenLibrary's public API, which differs from ISBNdb but has more comprehensive data and is free/unlimited. ISBNdb's API requires payment and isn't recommended for beginners.

After saving, navigate to the script directory in Terminal (e.g., `cd ~/Desktop`) and enter:

python3 check_books.py

If everything is normal, you'll see line-by-line output formatted like this:

ISBN:9787544291163 | One Hundred Years of Solitude | 2011 | Subjects: Literature, Magical Realism

ISBN:9787532768457 | Computer Systems: A Programmer's Perspective | 2016 | Subjects: Computer Science

```

Expected Result: Each ISBN corresponds to a line showing title, publication year, and subjects. If not found, the ISBN is likely wrong; recheck it.

Step 4: Filter Out "High-Risk Books"

The core characteristics of AI company procurement lists are: published before 2022 and niche subjects (especially professional textbooks, out-of-print books). Copy the script output into Excel or a spreadsheet and filter manually.

I made a simple table comparing my 50 books:

Category Quantity Description
Published 2022 or later 32 Relatively safe; AI companies aren't interested
Published before 2022, non-niche 15 Might get attention, but lower risk
Published before 2022, niche/out-of-print 3 High Risk, recommend taking photos

Pitfall: I initially only looked at publication year, ignoring subjects. Later I found a 2005 "Linear Algebra Problem Set" marked as "Rare Collection" on OpenLibrary, so I became cautious. If your book sells for more than twice the original price on the second-hand market, it's likely a target.

Step 5: Actual Protective Measures

After identifying high-risk books, you can do three things:

1. Photo Scan: Keep PDFs via phone or scanner; at least the content won't disappear.

2. Mark Ownership: Write "This version prohibited for AI training" inside the book. Legal validity is questionable, but it shows attitude.

3. Donate to Libraries: Some university libraries accept rare books, which is safer than selling to second-hand dealers.

Bold Note: Never dump your niche old books cheaply on Xianyu (second-hand marketplace); they might be swept up by AI company middlemen. A friend sold a 3rd edition of "Principles of Chemical Engineering" for 3 yuan, and the same used copy on ISBNdb jumped to 80 yuan because AI companies were bulk-buying.

What to Try Next

You can upgrade the script to automatically compare publication years and export a high-risk list to CSV for sharing. Or use ISBNdb's public search (visit the website, manually input ISBN) to check for "bulk purchase records." Although the page is gone, some historical data might remain.

Leaving a question: What do you think is the essential difference between AI companies shredding old books and Google scanning libraries? Feel free to discuss in the comments.

2 replies

?
Ctrl + Enter to reply
Amy
AmyAug 2

Just tried out WorkBuddy's automation features. Feels like ISBN duplicate checking and source verification could also be handled via workflows without writing scripts yourself... Have you tried it?

Early Investor

Wait, what's the deal with this "published before 2022" criterion?? Do AI companies really just look at the year when collecting books? This logic feels kinda crude to me... But the idea isn't bad, I'll run my own book collection through it later.