Deep Dive — Anthropic goes public with the dark-web distillation pipeline

Share
Deep Dive — Anthropic goes public with the dark-web distillation pipeline

For most of 2026 the fight over AI distillation has been a quiet argument: labs accusing each other in blog posts, lawmakers asking for hearings, an academic literature building around a single hard-to-define practice. On Thursday, Anthropic did something new — it put its head of threat intelligence on camera, named names, and walked a major business network through the operational mechanics of what it says is happening. The story is bigger than the quote. It says the illicit pipeline is organized enough to be its own supply chain, and the policy fight in Washington has just become an enforcement fight in the back alleys of the internet.

Jacob Klein, who runs threat intelligence at Anthropic, told CNBC there is "an entire illicit ecosystem" working to gain access to Claude and other frontier models, going through "fraudulent means" to spin up accounts "at extreme scale" so they can be queried at volume and the responses fed back into training a competing system. He singled out Moonshot AI, the Beijing-based lab behind Kimi K3, as having done this against the newest version of Claude. He pointed to Iran, Russia, and North Korea as additional sources of pressure. He named the dark web as a marketplace where stolen credit cards, foreign-phone SMS services, and compromised AI accounts trade alongside the more traditional contraband.

What's new is not the accusation — Anthropic has been making versions of it since February, when it alleged DeepSeek, Moonshot, and MiniMax had distilled its models, and again in June when it pointed at Alibaba's Qwen family. What's new is the mechanism: not a clever research team building on published API output, but a distributed procurement problem — accounts, phone numbers, payment instruments, the operational know-how to defeat each of Anthropic's controls in sequence. The threat-intel framing shifts the conversation from "did they copy us?" to "how do we stop the copying when the people copying us are buying our product with stolen infrastructure?"

The timing is not accidental. Anthropic is reportedly preparing to go public as soon as October, at a private valuation near a trillion dollars. A story framing it as the victim of a state-adjacent theft operation also frames the trillion-dollar IPO as a security story — a US lab whose intellectual property is being extracted at scale by foreign adversaries, with techniques described in enough operational detail for regulators and prosecutors to act. Whether or not that's the motive, it's the structure of the story CNBC told on Thursday.


The mechanism Klein described has three parts, and each is something the open-source research community had already sketched in pieces.

Account farming. The signal that distillation is happening is not a single user asking many questions — it's thousands of accounts each asking only as many as a real user would, with timing patterns that look human. The "Stolen Thoughts" researchers found 6,708 publicly shared agent trajectories on GitHub and Hugging Face and recovered 62 API keys, 33 passwords, and 24 access tokens from hidden reasoning alone. Sixty-four of those artifacts appeared nowhere in the visible conversation. For a state-linked adversary, the encrypted reasoning of a frontier model is itself a treasure trove, and the cheapest way to mine it is to spin up accounts at scale.

Payment and identity infrastructure. Anthropic blocks sanctioned countries. To get around that, the dark-web market sells credit card numbers tied to US or European billing addresses, plus foreign-phone SMS services that can receive the one-time codes Anthropic (and OpenAI, and Google) require for signup. Anthropic's controls include phone and billing checks, bans on majority-Chinese-owned firms, and in some cases live-selfie identity verification. The market has modularized defeating each control. A buyer can acquire a clean credit card, a working foreign number, and a stolen selfie ID as separate line items, then assemble them at signup. We covered the consumer-facing version in China's gray market sells Claude tokens at a tenth of the list price — Zilan Qian's ChinaTalk research showing Chinese developers buying Claude tokens through "transfer stations" for roughly ten percent of list price. Klein's interview makes clear the same infrastructure now serves a heavier customer: not the student who wants to try Claude, but the lab that wants to distill it.

Output capture. Once a fraudulent account is operational, the buyer asks thousands of questions and saves the answers. If the lab wants to distill reasoning specifically, they ask the kind of structured problems — math, code, multi-step planning — where frontier-model answers carry the most signal. Qian argues the biggest margin in the gray market may be the usage data every request leaves behind: prompts, responses, and tool calls pass through the proxy unencrypted to its operator, so each user is simultaneously a paying customer and an unpaid data producer. When the user is a state-linked lab with thousands of accounts, the framing starts to sound like an industrial-scale exfiltration operation.

Klein was careful to concede that distillation is not, by itself, illegal. The legal version looks like gaining permission, following IP and export-control law, and sometimes paying for it. His objection is specifically to the fraudulent path — stolen cards, fake identities, mass account creation — and to the destination of the resulting model, which he says lacks the safeguards Anthropic ships by default. "I think competition is great," Klein told CNBC. "The concern here is if you are taking our model, distilling it through fraudulent means, creating millions of fake accounts using stolen credit cards and stolen infrastructure, to then produce a model that doesn't have safeguards in place."

That last clause is the line a national-security argument pivots on, and Anthropic has been building toward it for months. In April the Trump administration's National Security Memorandum on AI distillation wrote that undermining American research and proprietary information through distillation is "unacceptable," and said the administration would explore "a range of measures to hold foreign actors accountable." Anthropic's August threat-disruption report — the one that included a documented case of its own model being used in a rogue GitHub campaign — was titled as a cybersecurity disclosure, not an IP complaint. CNBC's piece is the first time Anthropic has put a named executive in front of a US business network to describe the full pipeline end-to-end.


The most uncomfortable sentence in the CNBC piece is the one Anthropic did not say. The piece lists four named Chinese labs — Moonshot, DeepSeek, Alibaba, and MiniMax — and notes none responded to requests for comment. In February Anthropic accused Moonshot, DeepSeek, and MiniMax of running "more than 24,000 fake accounts to generate over 16 million requests against Claude." In June it accused Alibaba of conducting a "massive distillation attack." Thursday's piece is the latest in a long series of allegations rather than the first.

The legal floor under these accusations is thin. Distillation in its legitimate form — training one model on the outputs of another — is standard practice in the field and is how most open-weight models get a fraction of their capability. The line Anthropic draws is between legitimate distillation and fraudulent-distillation-via-stolen-infrastructure. To prove the second, you need to show not just that the resulting model performs similarly but that the inputs came through accounts that should not have existed. The "Stolen Thoughts" researchers, who had the strongest evidence of distillation-similarity in the public literature — finding Kimi K3's reasoning tracking Claude Opus 4.8 and GPT 5.6 Sol on some prompts — were careful to say similarity is "evidence, not a confession." Anthropic has more internal telemetry than any academic lab, but it has not, to date, published the account-level traces that would let an outside party verify the specific numbers it cites.

This is why the dark-web framing matters tactically. If Anthropic were simply saying "these labs copied us," the burden would be on Anthropic to prove copying. By saying "these labs bought stolen US credit cards to spin up accounts with us," the burden shifts toward the labs: how did they get 24,000 accounts? Where did the cards come from? How were the SMS codes received? A dark-web supply chain can be investigated by law enforcement in ways a paper-similarity score cannot. The threat-intel framing is also an enforcement framing, and it is the one most likely to land in Washington.

Travis Lanham, technology chief at the cybersecurity firm Armadin and a former Google engineer, made the structural point for CNBC: the labs are serving billions of legitimate requests, the fraudulent ones are millions, and the millions are hiding inside the billions. Anthropic concedes it can't fully stop the activity, only slow it down. "It's very hard to fully stop this as a problem," Klein said. The implication is that Anthropic's goal in going public is not to eliminate the activity but to make it more expensive — by triggering regulatory pressure on payment processors, by raising the legal exposure of the named labs, by making the buyers of stolen infrastructure easier to prosecute.

Focused view of a modern data server rack with blinking lights in a blue-lit environment.

There is a real argument that the labs are doing exactly what they accuse others of doing, and the asymmetry is worth naming. Anthropic, OpenAI, and Google have all published reports describing how their own models were used against them; none of them have published the specific training data for their frontier models, the specific web scrapes that produced them, or the licensing terms under which that data was acquired. The "Stolen Thoughts" researchers found that the encrypted reasoning blocks Anthropic ships to users can be lifted and replayed through a weaker, less-aligned sibling model, which will transcribe the reasoning verbatim — two API calls, no anti-distillation safeguards triggered. We argued in Hidden reasoning protects labs, not users that the control Anthropic is now defending was always failing at the two jobs it claimed to do. If you can recover the reasoning by replaying the encrypted block, the distillation moat has already been crossed, and the dark-web supply chain is just an alternative route to the same destination.

There is also a real argument that distillation is how the field works and that trying to criminalize it will produce more harm than it prevents. Twenty-five companies, including Nvidia, Microsoft, and Meta, signed an open letter in July warning policymakers against premature restrictions on distillation, on the grounds that the practice is how open-source models get built. Meta's Mark Zuckerberg made the same argument directly this week in defending his own open-weighting roadmap for Muse Spark. Their position is that an American AI ecosystem that can distill freely is more competitive globally than one that fences off its intellectual property, and that the Chinese labs will distill regardless of what US law says. Klein's framing — that the problem is specifically the fraudulent path, not distillation per se — is calibrated to win this argument: he is not asking for a ban on distillation, he is asking for enforcement on the supply chain. The harder counter-case is that Anthropic is choosing a definition of "fraud" that conveniently aligns with where it competes. Paying for API access and using the outputs to train a competing model is, on Anthropic's own description, legal. Creating a thousand accounts to do the same thing at volume is what Anthropic calls fraudulent. The distinction is operationally meaningful — account farming costs the provider money, ties up support, and triggers sanctions exposure — but it is also a distinction Anthropic gets to draw unilaterally.


What's likely to happen in the next 30 days is more disclosure, not less. Anthropic has more telemetry than it has shared; the CNBC interview is calibrated to be the first in a series, not the only one. Watch for three signals.

First, whether other labs follow Anthropic's lead. OpenAI published its own distillation-disruption report in August describing state-linked abuse of its models; it has not yet put a named executive on camera to walk through the supply chain. If Sam Altman or a threat-intel counterpart appears at a similar venue in the next two weeks, the disclosure pattern becomes an industry standard. If OpenAI stays silent, Anthropic is alone with the framing, and the named Chinese labs have a more useful target.

Second, whether the named Chinese labs respond. Moonshot's silence through the August coverage was strategic — denying without engaging. The CNBC piece is direct enough, and Klein's on-camera presence is unusual enough, that a continued silence starts to read like an admission. If Moonshot publishes a technical rebuttal — something specific, with hashes, with traces, that says "here is the provenance of our training data" — it changes the fight.

Third, whether the IPO process moves the story. Anthropic is reportedly going public as soon as October. S-1 filings require disclosure of material risks. A company that has told CNBC it is being systematically robbed of its core product, by named foreign adversaries, via documented dark-web infrastructure, has a strong candidate for a risk factor. The question is whether the public version of the S-1 makes the story louder than CNBC did, or whether it contains the disclosure the way most IPO risk factors contain risk — as a paragraph no one reads.

The deeper question is whether the fight has already shifted from IP to infrastructure. The frontier-model business is a business in outputs: tokens per second, capability per dollar. The Anthropic disclosure reframes the fight as one over inputs: who gets to buy access, through what payment systems, with what identity guarantees, and what happens to the prompts and responses after. The dark web is a story about inputs. The threat-intel team is an inputs team. The sanctions-evasion discussion is about who is allowed to be a customer at all. If Anthropic's framing wins, the competitive question is no longer "whose model is smarter" but "whose customer base can be trusted to be who they say they are." That's a very different business from the one the frontier-model labs have been building, and it is one in which Anthropic, with its $9.1 billion compute deal with Riot, its 30-day enterprise data option, and its visible lead in access-restriction engineering, is unusually well-positioned to win.

If the framing loses — if the Open Letter coalition wins in Washington, if the named labs successfully rebut, if the IPO disclosure reads as boilerplate — the story is a smaller one about a specific threat that didn't generalize. The infrastructure stays mostly the way it is, the labs keep selling to whoever can pay, and the dark-web supply chain continues as a low-grade background cost of doing business. What is determined now is that the fight has moved out of the blog posts and onto the business pages, and the people on the business pages are not the ones the AI labs have spent the last two years learning to talk to.

If the question shifts from "whose model is smarter" to "whose customer is who they say they are," is that a security win or a moat for the incumbents? Tell us in the comments.

Sources: CNBC — Anthropic's distillation battle turns to the dark web · White House NSTM-4 — distillation policy framework · ChinaTalk — How to Buy Cheap Claude Tokens in China · CNBC — February 2026 distillation allegations

Read more