Physix Frontier · News Briefing Card (Leiphone · Sep 10, 2026)

DeepSeek V4.1 Flash Cuts Long-Context VRAM by 75%

KEY FACTS

  • DeepSeek released the V4.1 Flash model, featuring new CED and CSA2 architectures.
  • Runtime VRAM usage dropped by 75%, while persistent cache shrank to one-eighth of the previous generation.
  • The model achieved a 74.2% pass rate on the DeepSWE benchmark, outperforming several competitors.
  • With 552B total parameters, only 8B are activated during the prefill stage.

KEY DATA

75%VRAM Reduction
1/8Persistent Cache Ratio
74.2%DeepSWE Pass Rate
552BTotal Parameters

PHYSIX OBSERVATION

DeepSeek has broken the high-cost curse of long-context agents through architectural innovation rather than brute-force scaling. The cliff-like drop in VRAM and storage costs signals that million-token applications are moving from labs to large-scale commercial deployment. This 'small bet big win' sparse pathway may force the industry to reevaluate the cost-performance logic of dense LLMs, accelerating the adoption of low-cost intelligent agents.

Source: Leiphone report