Physix Frontier · News Briefing Card (TMTPost · Sep 11, 2026)

DeepSeek V4.1 Flash Launches: Outperforms Pro, Prices Drop

KEY FACTS

  • DeepSeek released the V4.1 Flash model, featuring an asymmetric architecture with native multimodal support.
  • The model's memory requirements dropped to one-quarter of the previous generation, achieving measured output speeds of 300 to 500 tokens per second.
  • V4.1 Flash scored 90.9 on the GPQA Diamond benchmark, surpassing V4 Pro across the board.
  • Official pricing was cut, with cached input costs falling to 0.02 yuan per million tokens, doubling during peak hours.
  • Lead Cui Tianyi announced that requests for V4 Pro will be routed to V4.1 Flash and billed at the new lower rates.

KEY DATA

552BTotal Parameter Scale
90.9GPQA Diamond Score
0.02 yuan/million tokensCached Input Price
300-500 tokens/secOutput Speed

PHYSIX OBSERVATION

DeepSeek is reshaping the competitive landscape with a 'more for less' strategy, unifying high performance and low cost through architectural innovation. This move not only directly squeezes the survival space of same-size models but also signals a shift in the AI industry from a pure parameter race to an extreme cost-performance game. For developers, low-cost, high-throughput APIs are becoming the preferred infrastructure for building agent applications.

Source: TMTPost report