DeepSeek's Accuracy Varies by Token Position: Phase Sensitivity in KV Cache Compression
ByteDance's Seed team discovered that DeepSeek-V4 models exhibit periodic accuracy fluctuations every 4 tokens in long-context retrieval, caused by chunked KV cache compression creating phase sensitivity where retrieval performance depends on information position relative to compression window boundaries, a phenomenon mitigated but not eliminated by post-training.
ByteDance's Seed research team investigated why DeepSeek-V4 models show inconsistent performance — sometimes correct, sometimes wrong — on identical tasks when only irrelevant prefix tokens are added. They traced this "god-ghost duality" to the model's chunked KV cache compression mechanism.
Code Completion Experiment Reveals 4-Token Periodicity
Using DeepSeek-V4-Flash-Base, researchers took an FP8 quantization function from the official inference code and asked the model to complete the last token (correct answer: 8). They prepended a decorative docstring with varying numbers of equals signs, changing only the prefix length. The model's prediction flipped periodically: when the prefix length modulo 4 was 0 or 1, it favored the wrong answer 32 (71.3% probability vs 26.4% for 8); when modulo 4 was 2 or 3, it favored the correct answer 8 (91.5% vs 7.2%). The entire cycle repeated every 4 tokens.
Needle-in-Haystack Test Confirms Large Accuracy Gaps
To test generality, the team constructed a 128K-token context with ~16,000 key-value pairs and queried for specific values. Keeping the key-value mapping, question, and total context length fixed, they varied the target information's position relative to the compression window boundary. DeepSeek-V4-Flash-Base showed a maximum accuracy gap of 40.2 percentage points across positions; DeepSeek-V4-Pro-Base showed 34.8 points. Post-trained versions reduced but did not eliminate the gap: Flash-0731 (19.1 pts), Pro-0813 (14.8 pts), V4.1-Flash-0910 (6.1 pts). Notably, V4's period was 4 tokens, while V4.1's period shortened to 2 tokens.
Root Cause: Chunked KV Cache Compression and Phase Sensitivity
DeepSeek-V4 uses chunked KV cache compression to reduce memory and compute for long contexts. Consecutive tokens are grouped into windows and compressed into fewer cache entries. The researchers defined "Phase" as a token's position within its compression window (e.g., 1st, 2nd, 3rd, 4th for stride 4). They discovered Phase Sensitivity : retrieval capability varies systematically across phases. Even when both key and value fall in the same window, accuracy differs significantly by phase, indicating the issue stems from how information is written to and read from compressed cache, not merely from window-boundary splits.
Controlled Experiments Isolate the Compression Design
To confirm causality, the team trained models from scratch using the Qwen3-0.6B architecture with various chunked compression schemes, plus a full-attention baseline. All chunked compression models exhibited periodic fluctuations matching their compression stride (stride 4 → ~4-token period, stride 6 → ~6-token period, stride 8 → ~8-token period). The full-attention baseline showed no comparable periodicity. The effect persisted without RoPE position encodings and with simple averaging instead of learned compression weights, proving it is inherent to the chunked compression design itself.
Phase Specialization in Attention Heads
Intervening on attention heads revealed Phase Specialization : different heads contribute preferentially to different phases. Some heads excel at retrieving information from certain window positions, others from different positions. This division of labor aids overall retrieval but creates weaker phases. Theoretical analysis of a simplified training model suggests gradient flow naturally drives the compression module to develop stable positional preferences during training, explaining why the periodicity is systematic rather than random noise.
Implications for Evaluation and Deployment
Standard benchmarks that average across positions can mask phase sensitivity — a model may score well overall while failing catastrophically at specific phases. The authors argue that evaluating chunked-compression models requires testing the same information at every phase. Despite this flaw, chunked KV cache compression remains valuable for its memory and compute savings; post-training and architectural iteration (as seen in V4.1) can substantially narrow the phase gap, though the fundamental trade-off persists.
Paper: https://arxiv.org/pdf/2609.36322
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Machine Learning Algorithms & Natural Language Processing
Focused on frontier AI technologies, empowering AI researchers' progress.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
