Deep Dive — The four-token blind spot inside DeepSeek V4

ByteDance's Seed research team says it has found the cause of one of the stranger recurring complaints about DeepSeek's models: the same question, asked with nothing changed except a few junk characters bolted onto the front, can flip the model from right to wrong. Their paper, posted to arXiv on September 28, traces the wobble to a memory-saving trick used during long-context inference, and reports that DeepSeek-V4-Flash-Base's retrieval accuracy swings by as much as 40.2 percentage points depending on nothing more than where the answer sits relative to an internal window boundary. QbitAI reported the findings this week. The phenomenon has a name now — phase sensitivity — and it lands at an awkward moment, just as long-context capability has become the industry's headline spec.
A code snippet, a string of equals signs
The Seed researchers started with a small, almost petty experiment. They took an FP8 quantization function out of DeepSeek's official V4 inference code and asked DeepSeek-V4-Flash-Base to complete the final token. The right answer is 8 — the function casts to an FP8 type — but the model often insisted it should be 32.
Then they inserted a purely decorative docstring containing a run of repeated equals signs ahead of the code, and started changing how many equals signs it held. The code was untouched. The completion point was untouched. The correct answer never moved. What moved was the count of meaningless filler tokens before it — and the model's verdict oscillated: wrong answer, right answer, wrong answer, on a cycle of exactly four tokens. Across the sixteen filler lengths they tested, padding lengths leaving a remainder of 0 or 1 when divided by four pushed the model toward the wrong answer; remainders of 2 and 3 pushed it toward the right one. At the bad positions the wrong answer carried an average probability of 71.3 percent against 26.4 percent for the right one; at the good positions the numbers inverted — 91.5 percent correct versus 7.2 percent wrong. Two characters of padding, and the model's confidence in the same fact swings by sixty points.
That would be a curiosity piece on its own. The next test made it a story about the whole long-context stack.
40 points of accuracy, hiding in the padding
The team built a needle-in-a-haystack task the way this genre expects: a 128,000-token context containing roughly 16,000 key-value pairs (K1→V1, K2→V2, and so on), one lookup question, context length held constant, key-value mapping held constant. The only variable was where the target information sat relative to the boundary of the compression window the model uses internally — its "phase," in the paper's term.
The accuracy curves came back periodic. DeepSeek-V4-Flash-Base differed by up to 40.2 points between its best and worst positions; DeepSeek-V4-Pro-Base by 34.8 points. Post-training tames it but does not remove it: DeepSeek-V4-Flash-0731 narrows to 19.1 points, DeepSeek-V4-Pro-0813 to 14.8, and DeepSeek-V4.1-Flash-0910 to 6.1 — though the rhythm survives even there, with V4 wobbling on a four-token cycle and V4.1 on a two-token cycle. That ratio is the tell: the period matches each generation's KV-cache compression stride exactly.
What is chunked KV-cache compression, and why does it do this? When a model reads a very long prompt, it caches a key and value vector for every token it has seen, because the attention computation needs them again at every later step. The cache is the dominant memory cost of long-context inference, and memory is money. DeepSeek's approach, per the paper, compresses consecutive runs of tokens into fewer cache entries at a fixed stride — every four tokens become one compressed entry in V4. The saving is real. The side effect is that a token's identity now depends on its slot: a fact landing at position one of a compression window gets written into the cache under different conditions than the identical fact at position three. The authors call the difference between those slots the phase, and show it silently reweights what the model can retrieve later.
They did not stop at observing DeepSeek. Suspicious that the effect was an artifact of some quirk of one model, they pretrained their own family of transformers on a Qwen3-0.6B architecture across multiple compression designs, with a full-attention control. Every chunked-compression variant reproduced the periodicity, with the cycle tracking the stride they configured — set it to four, six, or eight, and the wobble followed. The full-attention control did not. The effect survived removing RoPE position embeddings and swapping learned compression weights for plain averaging. The conclusion is blunt: chunked compression by itself creates periodic retrieval weak spots; it is a property of the design, not a DeepSeek bug.
Why the industry should care
Start with what this costs DeepSeek's users. A 40-point swing is the difference between a model that reliably reads your contract and one that loses the clause depending on how the document's formatting happens to fall. Anyone running long documents through an API in production — code reviews over large repositories, legal review, agent memory over multi-hour sessions — has almost certainly hit a phase edge and blamed the model for being flaky. As we explained in our AI 101 on what a context window actually is, the market has spent two years selling context length as the spec that matters. The paper's implicit correction is that nominal context length and retrievable context length are different numbers, and only one of them is on the pricing page.
The economics explain why the problem exists in the first place. Compression at a stride of four exists because serving a million-token context at full cache is ruinously expensive — the pressure we described in Tokens halved in price. The bill to make them didn't has pushed every inference vendor toward cheaper per-token serving. Phase sensitivity is what that optimization charges on the other side of the ledger. Nobody chose flakiness; they chose memory savings, and the flakiness came bundled, unmeasured.
The authors tie this to a simple training-dynamics story: gradient flow during pretraining naturally sharpens position preferences, because a component that specializes in one phase can optimize for it without interference. Division of labor is efficient for the network and expensive for the user, since the reliability cost lands wherever the lookup happens to fall.
That measurement gap is the second-order story. The paper's sharpest claim is not about DeepSeek at all: average benchmark scores conceal systematic positional failures. A model can be near-perfect at some phases and mediocre at others, post an excellent aggregate, and ship. This is the same structural critique we raised when an independent run confirmed DeepSeek V4 Flash's 82.7% score — a headline number is only as meaningful as the distribution underneath it. The Seed team's recommendation follows directly: evaluate any model using chunked compression by moving the same fact across phases and testing each one, rather than trusting the mean. A diagnostic benchmark built for exactly this class of failure, KVDiagnosis, made the same argument in August — aggregate task scores reveal neither which correct executions fail nor why.
What skeptics say
Three objections are worth taking seriously.
First, the source: ByteDance Seed is DeepSeek's direct competitor, and the paper volunteers a rival's flaw with a reproducible recipe. Competitor research is not automatically wrong — the from-scratch reproductions are precisely what makes this hard to dismiss — but it is not disinterested either. DeepSeek has not, as far as we have seen, responded publicly.
Second, the severity is already shrinking with each release. The headline 40.2-point figure comes from a base model. On the current DeepSeek-V4.1-Flash-0910 the gap is 6.1 points, and DeepSeek's own iterative work — without knowing it was being graded on phases, apparently — was already hammering the effect down. A problem that post-training reduces by 85 percent is an engineering defect with a known repair path, not a wall.
Third, the tradeoff is deliberate and probably worth making. Full attention with an uncompressed cache is the control that passed — and it is also the configuration nobody can afford at a million tokens. Asking DeepSeek to ship the clean version is asking it to triple serving costs so that answers stop depending on token position modulo four. The realistic ask is phase-aware training and phase-stratified evaluation, not abandoning compression.
What to watch
Whether DeepSeek addresses stride or phase balance in the next V4.1 or V4.2 refresh — if the periodicity vanishes in a future release, this paper is the reason. Whether any eval suite adopts phase-stratified reporting as a standard column, because right now nothing on any public leaderboard forces it. Whether competitors quote "40 points" in launch materials — the most likely near-term use of this research is marketing, not methodology. And whether other long-context providers, whose compression schemes the paper did not test, quietly discover the same rhythm when someone finally moves the needle and holds everything else still.
If your team has seen long-context retrieval fail on one run and succeed on an identical rerun, was it the padding or something else? Tell us in the comments.




