Physix Frontier · News Briefing Card (IT Home · Oct 9, 2026)

ByteDance Seed Paper Exposes DeepSeek Long-Context Phase Sensitivity

KEY FACTS

  • ByteDance's Seed team submitted a paper to arXiv in late September studying the phase sensitivity of blockwise KV cache compression.
  • The study evaluated models including DeepSeek-V4-Flash, V4-Pro, and the post-trained V4.1-Flash.
  • Blockwise KV cache compression compresses contiguous token windows at a fixed stride to reduce long-context inference costs.
  • The compression introduces a new positional coordinate called token phase, meaning a token's position relative to the compression window boundary.
  • The same information is harder to retrieve at different phases, and long-context retrieval accuracy can differ by 40 percentage points.

KEY DATA

40 percentage pointsLong-context retrieval accuracy gap across phases

PHYSIX OBSERVATION

This study punctures an industry blind spot: good average benchmark scores do not mean long context is stably usable. Phase sensitivity means users may suddenly hit retrieval failures at specific positions, and vendor benchmarks may not expose this. For teams building long-document and RAG applications, model selection should add a phase robustness test, not just look at average scores.

Source: IT Home report