Physix Frontier · News Briefing Card (IT Home · Sep 16, 2026)
iFlytek Launches Spark-Audio-1.0 on Domestic Compute
KEY FACTS
- iFlytek has released the Spark-Audio-1.0-Preview, trained entirely on a domestic computing cluster.
- The model features a 0.65B audio encoder and a 30B-A3B MoE architecture, utilizing 13 million hours of audio data.
- It supports recognition in 99 languages and 202 dialects, with end-to-end capabilities for understanding semantics, emotions, and environmental sounds.
- The model achieved SOTA on the Fleurs Chinese test set, outperforming Gemini-3.1 Pro in multiple speech evaluation metrics.
- Compared to competitors of similar size, it excels in multilingual recognition without showing significant performance drops in general knowledge tasks.
KEY DATA
0.65BAudio Encoder Params
30B-A3BLLM Structure Params
13 million hoursTraining Audio Duration
99Supported Languages
PHYSIX OBSERVATION
iFlytek's move signals that China's AI infrastructure is transitioning from merely 'usable' to truly 'effective.' Full-stack domestic training not only mitigates risks associated with compute supply disruptions but also addresses information loss inherent in traditional cascaded solutions through an end-to-end architecture. Its leading performance in real-world scenarios, such as high-noise environments, proves that domestic models possess the practical potential to replace top-tier international closed-source models, offering a cost-effective new option for vertical industry deployment.
Source: IT Home report
Physix Frontier