Anthropic locks Claude's thinking blocks to kill API distillation
Two stories this morning that look unrelated and aren't: Anthropic just closed the last easy door for anyone training a model on Claude's reasoning, and researchers are close enough to decoding animal speech that bioethicists are already writing the warning label.
Anthropic has rewritten how its Messages API handles thinking blocks, and the change is aimed squarely at distillation. With Claude Fable 5.1, new API accounts can no longer edit the system prompt, tools, or earlier messages that sit around a prior thinking block in a multi-turn conversation — if the surrounding context has been modified, the API returns an error. A "non-strict" mode is available as an escape hatch: the request goes through, but the affected thinking blocks are silently dropped, so the model answers without seeing its own prior reasoning.
The loophole being closed is more specific than "people were copying outputs." Distillers found that by rewriting the context around a thinking block mid-conversation — swapping an early system prompt, editing what the model had supposedly been told before — they could push Claude into reproducing reasoning it would otherwise keep encrypted. Anthropic says the abuse runs on thousands of fake accounts running automated scripts around the clock, and its stated motivation isn't lost revenue but what it calls the decoupling of capability and safety: a distilled model inherits the reasoning, not the alignment work that took tens of millions of dollars to install.
Our read: this is the most consequential API change Anthropic has shipped in a while, and the mechanism matters more than the policy. Anthropic is not detecting distillation — it's removing the technical precondition for it, by making surrounding context immutable. The phase-in is unusually narrow, applying only to new Claude Platform, Bedrock, Vertex, and Azure Foundry accounts created on or after August 31, which reads less like a product rollout and more like a controlled experiment before a wider enforcement that Anthropic has already said is coming for future models. The tell is the framing: Anthropic is calling this "preserved thinking" rather than an anti-distillation lock, because consistent context has a second benefit — prompt caching hits far more often, cutting latency and cost for every developer who never touched the loophole. Legitimate teams doing context compaction or mid-conversation system reminders now have to choose between rewriting their harnesses or silently losing reasoning continuity. We went deep on the industrial side of this trade earlier this week — Deep Dive — Anthropic goes public with the dark-web distillation pipeline — and this change is what happens when a lab decides detection isn't working.
Bloomberg's Morgan Meaker reports that AI-driven animal communication research has advanced to the point where bioethicists are warning about a new class of harm. The proof of concept is crows: two biologists who have spent 25 and 30 years on crow calls have collected roughly 150,000 vocalizations, cross-referenced them against video of what the birds actually do when they call, and are finding distinctions that decades of human listening missed. Billionaire Jeremy Coller, who funds work in the area, predicts two-way animal communication by 2030.
The concern is not that the science fails but that it works. Researchers quoted in the piece raise the prospect of a crow being told a location is safe and then trapped, or of sonar-style tactics historically used to drive whales to the surface being repurposed with better targeting. Others point at the privacy and ecological cost of blanketing habitats with always-on microphones, and one biologist offered a bleaker inversion: selectively bred farm animals — 40-day-old chicks that only vocalize about food — may have so little to say that measuring them widens the gap between humans and animals rather than closing it.
Our read: the asymmetry here is what deserves attention. Every tool built to interpret animal communication is also a tool for manipulating it, and there's no equivalent of informed consent on one side of the conversation. The fieldwork is slow and careful — decades of recordings, video grounding, real skepticism about whether small acoustic differences carry meaning — while the deployment incentives, from agriculture to pest control, are fast and commercial. That gap is where the harm lands, and nobody is building the governance for it yet.
What to watch: whether Anthropic extends preserved thinking to existing accounts on the next model, and whether any animal-communication funder publishes a use policy before someone ships a product.
If you had to pick one — should labs be able to lock down their reasoning traces by default, or does that break legitimate research? Tell us in the comments.
Sources: Claude Help Center — Preserved thinking · 36Kr · Anthropic — Fable 5.1 migration guide · StartupHub.ai · Techmeme · BriefRay