Japan's used bookstores are being bought by the ton for AI

Share
Japan's used bookstores are being bought by the ton for AI

Japan's secondhand trade is being cleared out by the ton — and the paperwork that would explain it is scattered across court filings, a distributor's denial and a publishers' association's unanswered letter.


Used bookstores across Japan are reporting bulk orders of a kind they have never seen — hundreds of titles a day, a 50-ton consignment of Japanese books shipped to the United States, and every order, whoever placed it, routing to one logistics centre in Okayama Prefecture.

NTV, which interviewed shop owners, documented the pattern on September 22. One Tokyo store said what used to move was novels and comics; the surge is in philosophy, history, political history, medicine, law and Edo-period life and culture. One merchant had days when sales multiplied fivefold, and a store in the Kanto region cleared more than 1,000 books. A bookstore in western Japan received an email from a foreign company asking whether it could order tens of thousands. The orders arrive through multiple accounts under different names and share the same destination — and the firm operating that Okayama centre told NTV it does not comment on individual transactions.

The suspicion is training data, and it has a paper trail that predates this week. Asahi Shimbun and Nikkei, working from US court documents, have reported a November 2024 sales contract between Anthropic and Nippan, one of Japan's largest book distributors, with the filings describing Anthropic's intent to scan what it bought. Nippan says it has never sold books knowing they were intended for AI training, and declined to discuss individual transactions. Anthropic has said nothing publicly about the Japanese orders. The trade has already moved on it: the Japan Book Publishers Association, representing roughly 380 publishers, sent a fact-check demand dated September 17 to Nippan's president asking whether member-published works were sold and on what legal basis.

Two years ago this was an American story, then a European one. What is new is not a lab's behaviour — it is the discovery that the secondhand market has no defence built for it. A used bookstore is the last and most fragmented link in the book trade: it cannot tell which bulk order is a collector and which is a shredder, and it has little leverage to ask. The interpretive leap — that the Okayama shipments are scanned and then pulped — is NTV's inference, not a proven chain, and it is worth being precise about which claim is which. The tonnage and the routing are documented; the destination company is not; the recipient refused to say. That gap is the whole problem, and the people standing in the shop when the order arrives cannot see past it either. We covered the trade's European pushback in August — Booksellers suspect AI firms behind 'strange' bulk book orders.


A new class of model returns a decision instead of a paragraph, and Check Point Research found every configuration it tested could be talked out of the right one — with a forged audit letter, at about 50 cents a break.

Jev, from TypeSafe AI, is built for software to call rather than for people to read: give it a question and the data to judge it, and it returns a typed verdict with a probability attached. Check Point ran it as a due-diligence assistant reading a report that flags every warning sign of a Ponzi scheme, gave one attacker control of a single section, and asked for a pass. All nine attacker-and-difficulty combinations produced at least one full flip — risk rated low, investment advised. The strongest attacker broke 25 of 27 runs, typically on the fourth turn of ten, at roughly 50 cents per successful break. Nothing in the attacks told the model what to output; the payloads simply appended what looked like a clean audit opinion and a revised risk table, and the model did its job correctly on false evidence. Marking the document as untrusted changed nothing, and an explicit instruction to ignore embedded instructions moved the break rate from 18 to 17 of 27 runs. Reasoning — the one defence that mattered in the comparison, taking a rival model's flip rate from 67% down to 19% — is a setting Jev does not offer.

The research is directional, not a benchmark: one application, one objective, and its attacker could observe the model's probabilities. The framing is the part to keep. Typed output constrains the shape of an answer; it says nothing about the documents the answer is drawn from, and a well-formatted, confidently wrong verdict is harder to argue with than a paragraph because there is no paragraph to disagree with. TypeSafe's own limitations page concedes the model "does not treat state as hostile by default." We covered Jev when it dropped its waitlist, and this fits a pattern we've documented twice this month — Deep Dive — Agents hacked three sites while fetching spreadsheets.


What to watch: whether Japan's publishers get a straight answer on where the Okayama books go, and whether a distributor's denial survives the association's letter.

If a secondhand market can't tell a collector from a shredder, who is supposed to be checking? Tell us in the comments.

Sources: NTV Japan · Tom's Hardware · Asahi Shimbun · Check Point Research · TypeSafe AI — Jev limitations · VentureBeat