Physix Frontier · News Briefing Card (enbrief · Sep 12, 2026)

DeepSeek V4.1 Flash Launches: Outperforms Pro, Cuts Prices

KEY FACTS

  • DeepSeek released the V4.1 Flash model, featuring an asymmetric architecture and native multimodal support.
  • The model reduces memory requirements to one-quarter of the previous generation, with measured output speeds reaching 300 to 500 tokens per second.
  • V4.1 Flash scored 90.9 on the GPQA Diamond benchmark, comprehensively surpassing V4 Pro.
  • Official pricing was lowered, with cache-hit input costs dropping to 0.02 yuan per million tokens, doubling during peak hours.
  • Lead Cui Tianyi announced that V4 Pro requests will be routed to V4.1 Flash and billed at the new lower rates.

KEY DATA

552BTotal Parameter Scale
90.9GPQA Diamond Score
0.02 yuan/million tokensCache-Hit Input Price
300-500 tokens/secOutput Speed

PHYSIX OBSERVATION

DeepSeek is reshaping the competitive landscape with a 'more for less' strategy, achieving high performance and low cost through architectural innovation. This move not only directly squeezes the survival space of same-size models but also signals a shift in the AI industry from a pure parameter race to an extreme cost-performance game. For developers, low-cost, high-throughput APIs are becoming the preferred infrastructure for building Agent applications.

Source: enbrief original report ↗