Physix Frontier · News Briefing Card (enbrief · Sep 17, 2026)

iFlytek Launches Spark-Audio-1.0 on Domestic Compute

KEY FACTS

  • iFlytek has released the Spark-Audio-1.0-Preview, trained entirely on a domestic compute cluster.
  • The model features a 0.65B audio encoder and a 30B-A3B MoE architecture, utilizing 13 million hours of audio data.
  • It supports recognition for 99 languages and 202 dialects, with end-to-end capabilities for understanding semantics, emotions, and ambient sounds.
  • The model achieved SOTA on the Fleurs Chinese test set, scoring higher than Gemini-3.1 Pro in multiple speech evaluations.
  • It outperforms same-size competitors in multilingual recognition without showing significant performance drops in general knowledge tasks.

KEY DATA

0.65BAudio Encoder Params
30B-A3BLLM Structure Params
13 million hoursTraining Audio Duration
99Supported Languages

PHYSIX OBSERVATION

This move marks a shift for China's AI infrastructure from 'usable' to 'optimal.' Full-stack domestic training not only mitigates supply chain risks but also addresses information loss inherent in traditional cascaded systems via an end-to-end architecture. Its leading performance in real-world scenarios like high-noise environments demonstrates that domestic models now possess the practical potential to replace top-tier international closed-source models, offering a cost-effective new option for vertical industry deployment.

Source: enbrief original report ↗