DeepSeek's Accuracy Varies by Token Position: Phase Sensitivity in KV Cache Compression

ByteDance's Seed team discovered that DeepSeek-V4 models exhibit periodic accuracy fluctuations every 4 tokens in long-context retrieval, caused by chunked KV cache compression creating phase sensitivity where retrieval performance depends on information position relative to compression window boundaries, a phenomenon mitigated but not eliminated by post-training.

Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
DeepSeek's Accuracy Varies by Token Position: Phase Sensitivity in KV Cache Compression

ByteDance's Seed research team investigated why DeepSeek-V4 models show inconsistent performance — sometimes correct, sometimes wrong — on identical tasks when only irrelevant prefix tokens are added. They traced this "god-ghost duality" to the model's chunked KV cache compression mechanism.

Code Completion Experiment Reveals 4-Token Periodicity

Using DeepSeek-V4-Flash-Base, researchers took an FP8 quantization function from the official inference code and asked the model to complete the last token (correct answer: 8). They prepended a decorative docstring with varying numbers of equals signs, changing only the prefix length. The model's prediction flipped periodically: when the prefix length modulo 4 was 0 or 1, it favored the wrong answer 32 (71.3% probability vs 26.4% for 8); when modulo 4 was 2 or 3, it favored the correct answer 8 (91.5% vs 7.2%). The entire cycle repeated every 4 tokens.

Needle-in-Haystack Test Confirms Large Accuracy Gaps

To test generality, the team constructed a 128K-token context with ~16,000 key-value pairs and queried for specific values. Keeping the key-value mapping, question, and total context length fixed, they varied the target information's position relative to the compression window boundary. DeepSeek-V4-Flash-Base showed a maximum accuracy gap of 40.2 percentage points across positions; DeepSeek-V4-Pro-Base showed 34.8 points. Post-trained versions reduced but did not eliminate the gap: Flash-0731 (19.1 pts), Pro-0813 (14.8 pts), V4.1-Flash-0910 (6.1 pts). Notably, V4's period was 4 tokens, while V4.1's period shortened to 2 tokens.

Root Cause: Chunked KV Cache Compression and Phase Sensitivity

DeepSeek-V4 uses chunked KV cache compression to reduce memory and compute for long contexts. Consecutive tokens are grouped into windows and compressed into fewer cache entries. The researchers defined "Phase" as a token's position within its compression window (e.g., 1st, 2nd, 3rd, 4th for stride 4). They discovered Phase Sensitivity : retrieval capability varies systematically across phases. Even when both key and value fall in the same window, accuracy differs significantly by phase, indicating the issue stems from how information is written to and read from compressed cache, not merely from window-boundary splits.

Controlled Experiments Isolate the Compression Design

To confirm causality, the team trained models from scratch using the Qwen3-0.6B architecture with various chunked compression schemes, plus a full-attention baseline. All chunked compression models exhibited periodic fluctuations matching their compression stride (stride 4 → ~4-token period, stride 6 → ~6-token period, stride 8 → ~8-token period). The full-attention baseline showed no comparable periodicity. The effect persisted without RoPE position encodings and with simple averaging instead of learned compression weights, proving it is inherent to the chunked compression design itself.

Phase Specialization in Attention Heads

Intervening on attention heads revealed Phase Specialization : different heads contribute preferentially to different phases. Some heads excel at retrieving information from certain window positions, others from different positions. This division of labor aids overall retrieval but creates weaker phases. Theoretical analysis of a simplified training model suggests gradient flow naturally drives the compression module to develop stable positional preferences during training, explaining why the periodicity is systematic rather than random noise.

Implications for Evaluation and Deployment

Standard benchmarks that average across positions can mask phase sensitivity — a model may score well overall while failing catastrophically at specific phases. The authors argue that evaluating chunked-compression models requires testing the same information at every phase. Despite this flaw, chunked KV cache compression remains valuable for its memory and compute savings; post-training and architectural iteration (as seen in V4.1) can substantially narrow the phase gap, though the fundamental trade-off persists.

Periodic accuracy fluctuation in code completion
Periodic accuracy fluctuation in code completion
Chunked KV cache compression illustration
Chunked KV cache compression illustration
Probability flip across prefix lengths
Probability flip across prefix lengths
Meme about checking prefix tokens
Meme about checking prefix tokens
Needle-in-haystack accuracy curves
Needle-in-haystack accuracy curves
Controlled experiment results across strides
Controlled experiment results across strides
Phase specialization across attention heads
Phase specialization across attention heads
Theoretical model of gradient-driven phase preference
Theoretical model of gradient-driven phase preference

Paper: https://arxiv.org/pdf/2609.36322

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

model evaluationLLM ArchitectureDeepSeek-V4Long Context RetrievalKV Cache CompressionByteDance SeedPhase Sensitivity
Machine Learning Algorithms & Natural Language Processing
Written by

Machine Learning Algorithms & Natural Language Processing

Focused on frontier AI technologies, empowering AI researchers' progress.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.