Physix Frontier · News Briefing Card (QbitAI · Oct 9, 2026)

ByteDance Finds DeepSeek-V4 Retrieval Wobbles Every 4 Tokens

KEY FACTS

  • ByteDance's Seed team found that DeepSeek-V4's performance varies periodically with input position every 4 tokens.
  • In a 128K long-context retrieval test, moving the same information to a different position produced a maximum accuracy gap of 40.2 percentage points.
  • The phenomenon is tied to the block-wise KV cache compression technique used by DeepSeek-V4.
  • The team named this phenomenon phase sensitivity and found that different attention heads exhibit phase specialization.
  • Post-training and architecture iterations can narrow the gap, but the periodic variation persists.

KEY DATA

40.2 percentage pointsMaximum accuracy gap
128K tokensContext length
about 16,000Number of key-value pairs
4 tokensFluctuation period

PHYSIX OBSERVATION

Block-wise KV cache compression is a key means of cutting the cost of long context, but ByteDance's finding reveals that it may introduce a systematic weakness of positional preference. Conventional evaluations that take average scores easily mask this fluctuation, and users in real-world use may encounter cases where the same question phrased differently gets answered incorrectly. This reminds the industry that long-context capability evaluation needs to introduce a phase dimension, and model iteration also needs to pay attention to the hidden bias brought by compression mechanisms.

Source: QbitAI report

ByteDance Finds DeepSeek-V4 Retrieval Wobbles Every 4 Tokens | Physix Frontier Briefing