I Spent Two Days Testing a Voice-LLM 'Co-Pilot' for Car Infotainment Systems
Conclusion first: This solution is completely feasible, but currently only suits tinkerers like me. Beginners should just buy a factory-installed head unit supporting voice large models for peace of mind. However, if you're willing to spend two days, you can build a voice assistant for older cars that understands complex commands, far superior to old-style head units that only handle "Navigate to XX."
I drive an Arcfox Alpha T7, which comes with Huawei Qiankun Smart Driving's voice system. Honestly, that assistant mostly controls functions like "Turn on AC" or "Increase Volume." Asking it to understand "Find a route avoiding highways with charging stations along the way" results in mechanical responses. So I wanted to try using popular voice large models to add a "co-pilot" to the head unit.
Step 1: Understand What Voice Large Models Do in Cars
Principally, traditional car voice systems stitch together three independent modules: Speech Recognition + Natural Language Understanding + Task Execution. Voice large models merge these into a single end-to-end model, taking audio input and producing audio output directly, eliminating cumulative recognition errors.
Concept for beginners: A voice large model isn't Siri or Xiao Ai style "voice assistants." It's an AI brain that understands speech and organizes its own language to reply. You don't need to say "Open Navigation." You can say "Take me to that hotpot restaurant in Chaoyang District where we ate last time, but don't take the 5th Ring Road, it's jammed, take the side roads," and it will understand and execute.
Step 2: Prepare Necessary Tools
I spent two days mainly testing three things:
1. StepFun STEPX Voice Model (Voice interface provided by Step AOS platform)—Used for two weeks, effects were good, especially for multi-turn Chinese dialogue.
2. Alibaba Cloud Qwen Audio 3.0 Realtime Plus—Tried for one day; materials show it topped the global voice reasoning leaderboard on July 28 with a score of 99.2%.
3. USB Sound Card with Microphone Array—Costs a few dozen yuan, plugs into the car's USB port.
If you don't have an Arcfox-style head unit, you can use an Android tablet or old phone as the head unit, install CarPlay or Android Auto, and connect an external microphone.
Step 3: Connect to the Voice Model
I mainly used the StepFun STEPX voice model because its API documentation is clear and beginner-friendly. Steps:
1. Register an account on the Step AOS Open Platform and apply for a Voice Model API Key.
2. Install a Python environment on the head unit (or use Termux on Android).
3. Use the code below to connect to the model:
Replace with your API Key
API_KEY = "Your_API_KEY"
url = "https://api.stepfun.com/v1/audio/chat"
Record microphone audio
def record_audio():
p = pyaudio.PyAudio()
stream = p.open(format=pyaudio.paInt16, channels=1, rate=16000,
input=True, frames_per_buffer=1024)
print("Please speak.")
frames = []
for _ in range(0, int(16000 / 1024 * 5)): # Record 5 seconds
data = stream.read(1024)
frames.append(data)
stream.stop_stream()
stream.close()
p.terminate()
return b''.join(frames)
Send voice and get reply
def voice_chat():
audio_data = record_audio()
headers = {"Authorization": f"Bearer {API_KEY}"}
files = {"file": ("audio.wav", audio_data, "audio/wav")}
data = {"model": "step-audio-chat", "stream": True}
response = requests.post(url, headers=headers, files=files, data=data, stream=True)
for chunk in response.iter_lines():
if chunk:
print(chunk.decode()) # Output voice reply
Pitfall Tip: The most common error here is audio format. I initially used the default mic sample rate, and the model couldn't understand anything. Checking docs revealed it must be 16kHz mono 16-bit PCM audio. If your mic doesn't support this, add a resampling step in the code.
Step 4: Make the Head Unit Truly "Understand" and Execute Operations
Just having the model reply isn't enough. The true value of voice large models lies in understanding intent and calling head unit functions. Two things are needed:
1. Convert model text replies to speech playback—Done via Text-to-Speech (TTS) services. I used Baidu AI's voice cloning feature to customize the voice.
2. Parse model intent into head unit commands—For example, when the model says "Navigating to Chaoyang hotpot restaurant," it needs to call the navigation API.
Implementation-wise, I wrote a simple parser matching keywords from model output to head unit functions:
def parse_intent(response_text):
if "navigation" in response_text:
# Extract destination from text
destination = extract_destination(response_text)
call_navigation(destination)
elif "AC" in response_text:
temp = extract_temperature(response_text)
set_ac_temp(temp)
Beginners often get stuck here: Model output is natural language, not fixed format, so parsing logic is fragile. I tried many times; slight deviations in user speech caused parsing failures. Later, I changed it to force the model to output JSON-formatted commands, e.g., {"action": "navigate", "destination": "Chaoyang Hotpot Restaurant"}, making parsing much more stable.
My Biggest Surprise
The biggest surprise was response speed. Materials mention Alibaba Cloud Qwen Audio 3.0 Realtime Plus has very low average first-token latency. My experience with StepFun's model was similar—from finishing speech to hearing reply, it takes about 1-2 seconds. In driving scenarios, this is crucial; delays over 3 seconds make users think "This AI sucks."
However, latency issues exposed another pitfall: network stability. In underground parking lots or areas with poor signal, latency spiked to over 5 seconds, sometimes timing out. So if you really want to use this, insert a high-speed IoT SIM card in the head unit or use a mobile hotspot. Don't expect the car's built-in WiFi to hold up.
Pitfall Summary
| Issue | Symptom | Solution |
|---|---|---|
| Incorrect mic sample rate | Model doesn't understand | Force set to 16kHz mono 16-bit |
| High network latency | Slow voice replies | Use high-speed data card, avoid WiFi |
| Unfixed model output format | Parsing failure | Force model to output JSON format |
| Lost context in multi-turn dialogue | Forgets mid-conversation | Pass historical conversation ID in requests |
What to Try Next
If you master this basic version, try advanced operations:
1. Integrate multiple voice models—Use Qwen Audio for ASR, Claude for reasoning, Baidu AI for TTS, finding the optimal combo.
2. Add visual recognition—Use cameras to capture road conditions; voice model responds based on visuals, e.g., "Traffic ahead, suggest taking side road."
3. Local deployment—Materials mention Baidu AI offers private voice deployment solutions. If you don't want cloud dependency, try running on local dev boards.
Core Viewpoint Summary: Implementing voice large models in cars is essentially grabbing the default interaction entry point. But for tinkerers like us, building a "co-pilot" via APIs beforehand lets you instantly feel the impact of technological evolution.
Physix Frontier