
DeepSeek launches voice features: look beyond the four voices
My judgment is put here first: DeepSeek is gray-testing voice chat, and text models are entering the voice interaction scenario. The four voice tones are just surface-level differences; latency, interruption handling, noise processing, post-transcription tagging, summarization, archiving workflows, and cost structures determine commercial value. This shouldn't be viewed merely as a new feature but also as a migration of interaction entry points.
According to IT Home feedback, tested users see a small speaker button in the top-right corner of the App, supporting four voice tones.
I've used DeepSeek for two months; the text side is already smooth enough for asking about industries, breaking down frameworks, and writing meeting minutes. But voice is a different story. Behind voice chat lies a string of engineering involving ASR, LLM, TTS, RTC, and VAD. From public feedback, DeepSeek's current gray test might be wrapping the voice processing pipeline within the App, or it might still rely on third-party capabilities at the underlying level. Benchmarking against overseas cases, similar paths usually fall into two categories.
One category is model vendors building direct entry points. The advantage is a unified experience where voice tone, interruption, and context management can be tuned together. The disadvantage is heaviness; voice processing isn't what model companies are best at, and they have to handle operations and compliance themselves. The other category is middleware like Agora aggregating services. A quote in the materials is quite straightforward: 2 lines of code, less than one cent per minute, interruption response as low as 340ms, capable of blocking 95% of environmental noise. This path is light, connecting large models to voice, suitable for quick launches. But lightness has its costs; supplier bargaining power, quality attribution, and cost fluctuations are all externalized.
Looking at the relationship between suppliers and buyers, in the middleware route, supplier bargaining power concentrates in ASR, TTS, and RTC, posing a replacement risk for model factories. Buyer switching costs depend on whether historical corpus data can retain usable records. The direct route is the reverse; once user data and voice tones run smoothly, migration costs become higher. The former looks like product development, while the latter looks like supply chain management. If DeepSeek pursues voice long-term, it will eventually have to pick a position between these two routes.
I wrote about customer interview automation a couple of days ago, starting with getting the directory structure right. When I did digital transformation projects at McKinsey, the most common issue was undefined processes leading to premature tool adoption. The pitfalls for landing voice dialogue are actually the same. Tools let you start speaking quickly, but enterprise processes won't automatically clean up just because you have voice.
Scenarios need to be narrowed down first. Meeting notes, pre-sales FAQs, outbound customer service calls, and elderly companionship all look like "talking," but their acceptance metrics are completely different. Customer service requires low misrecognition and compliant scripts; meetings require multi-speaker separation and minutes; outbound calls require concurrency costs and emotion detection. Post-transcription tagging, summarization, and archiving workflows matter more than voice tone. Voice comes in, transcription goes out, then tagging, summarizing, and archiving. If the directory isn't standardized, cross-departmental use becomes chaotic. Compliance boundaries must be defined upfront. Recording authorization, voice cloning, minors, financial/medical scripts—these aren't things to patch after launch. Costs must be calculated based on duration and recalculation. If a user interrupts once, the system may re-recognize, re-generate, and re-synthesize; costs cannot be directly analogized to text chat.
The moat for Voice AI isn't whether the model can speak, but whether it can stably take over a segment of business process. Previously, I thought that behind cheap crawlers, costs had just shifted positions. Applied to voice, costs might shift from the development side to the operations side; model calls become cheaper, but noise processing, interruption strategies, manual review, and voice licensing become ongoing expenses.
Don't initiate a project just because of four voice tones. Run a two-week pilot with a low-risk scenario, such as internal meeting minutes or organizing pre-sales questions.
Define acceptable deliverables first: transcription accuracy, speaker diarization, manual correction time, and cost per conversation. Then integrate with small traffic; don't go full-scale on customer service immediately.
If an enterprise already has a DeepSeek text assistant, treat voice as an experimental entry point rather than a strategic one. What is scarcer in enterprise digitization is acceptance criteria. If you can build directories, tags, SLAs, and review workflows, the tools can be swapped out and still work.
DeepSeek opening its mouth is a signal that large model competition is moving from answering text to voice chatting. Next, we'll see who can turn every voice interaction into traceable, reusable, and assessable data assets.
Physix Frontier