Valuation Surges $11B in Six Months: Why AI Audio Monetizes Faster Than Video?
Source: Letter AI (Authorized release by TMTPost) | Toutiao
Some say AI programming makes money, some say AI Agents make money, and some say AI video generation makes money. But did you know that AI audio is actually quite profitable too?
Foreign media reports state that ElevenLabs is internally discussing a secondary market share sale transaction, allowing employees to sell their stocks. This deal is expected to close before September, with a valuation of approximately $22 billion. Just five months ago, the company completed a $500 million Series D funding round at a $11 billion valuation. In less than half a year, the valuation doubled.
The key question is: How does an AI voice company dare to increase its valuation by $10 billion in six months?
From AI Dubbing Tool to Voice Infrastructure
In 2022, ElevenLabs' two founders took the TTS (Text-to-Speech) route. Traditional TTS is essentially "concatenative synthesis," losing tone, rhythm, and emotion. They used deep learning to help models understand text meaning, directly generating speech with emotion, rhythm, and pauses.
In 2024, they launched the Conversational AI platform (later renamed ElevenLabs Agents). The logic is: User voice → Speech-to-text → LLM understanding → Text-to-speech → Emotionally rich response, all in just a few hundred milliseconds.
ElevenLabs' advantage lies in using its own models for both "listening" and "speaking":
- eleven_flash_v2_5: ~75ms latency, specialized for real-time conversation
- eleven_v3: Covers 70+ languages, more expressive, suitable for content production
Based on this, they expanded into Dubbing (multilingual dubbing) and Music (music generation). In July 2026, Netflix used ElevenLabs AI to reconstruct the voice of late actor Gene Wilder for program narration, authorized by the estate committee—celebrity voices have become licensable IP.
By the end of 2025, ARR was nearly $350 million; by April 2026, it exceeded $500 million. Over 2 million Agents have been created on the Eleven Agents platform, processing more than 33 million real production conversations in the first half of 2026.
Why Does Voice Make Money Before Video?
First, lighter cost structure. Speech generation processes time series, while video generation must maintain consistency in characters, backgrounds, actions, camera angles, lighting, and frame continuity. Single-output costs for video are extremely high, and results aren't always "usable." Audio product forms are mature, with single-instance costs far lower than video.
Second, more certain scenarios. Dubbing, audiobooks, short video narration, localization translation, customer service calls, sales outreach, employee training, online education, game NPCs... AI audio isn't creating new demand; it's replacing existing dubbing methods and expanding capacity.
Third, much lower barriers than video. Speech only needs to meet four metrics: clear quality, natural emotion, low latency, and good stability. Once these thresholds are met, it can go directly into production workflows. The distance from "demo-ready" to "production-ready" is short, enabling faster monetization.
Fourth, voice is the natural entry point for Agents. Text interaction was the interface of the internet era; voice is the interface of the Agent era. Voice can also extend beyond screens—phone lines, headphones, car systems, and offline stores.
Why Is There No ElevenLabs in China?
Take Doubao as an example. Doubao's voice is part of ByteDance's AI entry point. Short video dubbing uses CapCut, web novel listening uses Tomato Novel, AI assistant voice uses Doubao, and enterprise voice APIs use Volcano Engine—the most lucrative voice scenarios have long been captured within the ByteDance ecosystem.
ElevenLabs' Concerns
Aggregation platforms like Vapi sit at a higher orchestration layer, turning TTS into replaceable components. Today they connect to ElevenLabs; tomorrow they can switch to OpenAI, Cartesia, or PlayAI. ElevenLabs relies solely on itself; if subsequent computing power fails to keep up with updates, it may inadvertently boost aggregation models like Vapi.
Original Link: Valuation Soars $11 Billion in Half a Year: Why Does AI Audio Make Money Before Video?
Physix Frontier