Community Discussion · Forum

Building a voice assistant from scratch to conversational capability in two days

Tian JiTian JiAug 152026/08/15 485 views

Conclusion first: Building a conversational voice assistant yourself has a lower barrier than many imagine. The open-source project S.A.T.U.R.D.A.Y packages the entire pipeline, assembling all the components from "hearing you speak" to "responding with voice." I ran it locally for a week and broke down the steps for you; following along, you can get it running within two days.

Background first. The project is fully named S.A.T.U.R.D.A.Y, positioned as open-source, self-hosted, J.A.R.V.I.S. Self-hosted means your own computer or server is its home; everything runs locally without passing through any third-party cloud services. J.A.R.V.I.S refers to Iron Man's butler—of course, this one isn't as smart as in the movies, but what it does is already practical: you speak, it converts speech to text, sends text to AI for understanding, then generates voice to reply.

For beginners completely new to this, let me explain three terms first (won't repeat later). Speech-to-Text (STT) converts your spoken words into on-screen text. Text-to-Speech (TTS) does the reverse, turning text into sound. WebRTC is a technology for real-time audio/video transmission in browsers without installing extra software; it serves as the communication backbone of the entire project.

I tried two approaches myself. Plan A uses the full S.A.T.U.R.D.A.Y assembly—components are hassle-free, and setup is quick. Plan B involves assembling parts manually: whisper.cpp for STT, LLM APIs for understanding, Coqui TTS for voice, linked via WebRTC. After getting both running, the differences are in this table:

Comparison Item Plan A: S.A.T.U.R.D.A.Y Full Suite Plan B: DIY Assembly
Deployment Difficulty Low, clone and run High, tune each component
Privacy Fully local, data stays home Depends on config; can be fully local or mixed with cloud APIs
GPU Required? Recommended; CPU works but is slow Same as above
Flexibility Medium, follows project design High, swap any component you want
Suitable For Beginners wanting quick setup Enthusiasts interested in researching the pipeline

My advice is to go straight for Plan A, get it running first, then dissect it to see what each component does. Don't start with assembly immediately; if you haven't seen how it works normally, you won't know which link is broken when issues arise.

Follow these steps, assuming you're using Ubuntu 22.04 or newer, with Docker installed. What Docker is doesn't matter—think of it as a box that packages software for you; the project author set up the environment, so you just press a button to launch. If not installed, run sudo apt install docker.io docker-compose first, then confirm docker --version outputs a version number.

Step 1: Find a directory and pull the code. Enter in terminal:

git clone https://github.com/GRVYDEV/S.A.T.U.R.D.A.Y
cd S.A.T.U.R.D.A.Y

You'll find a docker-compose.yml file in the directory; all subsequent services are launched by it.

Step 2: Check docker-compose.yml to ensure no port conflicts. Default is 8080. If another service on your machine occupies 8080, change 8080:8080 in the config to 8090:8080 (left is external access port, right is container internal port). I fell into this trap once; launching without checking caused port occupation errors in logs, wasting twenty minutes figuring out why.

Step 3: Start services. Run docker-compose up -d. The first time pulls images, taking about ten to twenty minutes depending on your network. Seeing Started indicates it's up. Use docker-compose logs -f for real-time logs. If it fails to start, it's usually an image pull failure; switch Docker mirror sources and try again.

Step 4: Open browser and visit http://localhost:8080. The interface is a simple webpage with a big button in the center. Hold it down to speak. Release, wait one or two seconds, and hear the voice reply. The first time I got it running, I froze for a second—it actually replied. That moment felt pretty cool.

At this point, you've completed the core pipeline. But there's a big pitfall: Chinese recognition. Whisper.cpp's default model works well for English but garbles Chinese. My initial test heard "How's the weather today?" as something completely unrelated. Solution: Specify a Chinese model in config or add -l zh to whisper's startup parameters to force language to Chinese. This config hides in the environment section of the whisper service in docker-compose.yml; add a line LANGUAGE=zh, restart the service, and it takes effect. Accuracy improved noticeably after changing this.

Another pitfall worth mentioning is latency. The slowest part of the pipeline is TTS voice generation. Cold starts require loading the model into memory, causing a three-to-five-second wait for the first interaction, then it speeds up. Don't assume installation failed; this is normal. For faster response, switch to a smaller TTS model, at the cost of reduced voice quality—it'll sound more robotic.

My personal experience: As a personal voice assistant, this setup is already capable. Many users replace their home Alexa with this, citing the same reason: they don't want cloud microphones always on. As you know, for local deployment, privacy isn't a feature; it's a prerequisite based on trust.

What to try next after learning this? Connect the voice assistant to your smart home, e.g., controlling lights or sockets. This involves adding "tool calling" capabilities to the assistant. Or conversely, extract the speech recognition part for standalone use; it's a clean STT tool suitable for meeting transcription. A friend of mine is doing this, routing planned meeting minutes processes through local whisper, saving monthly subscription fees to cloud providers.

Finally, a question for those eager to tinker: How do you evaluate correctness in voice assistant recognition? The term "accuracy" varies greatly by scenario. Saying "turn off the light" into a microphone versus transcribing a "project weekly report" requires different metrics. What indicators do you think should measure local deployed speech recognition?


📌 This article is compiled from Hacker News, original text: https://github.com/GRVYDEV/S.A.T.U.R.D.A.Y

Copyright belongs to the original author; this is a compilation and independent analysis based on public reports.

3 replies

?
Ctrl + Enter to reply
He Ma Chu Lai De

In store application scenarios, noisy environments are definitely a pain point. Hema tried voice assistants in their self-checkout areas, but poor VAD tuning directly caused response rates to crash. Have you tested this solution in actual stores? What was the feedback?

Back From Silicon Valley

From a global perspective, the biggest bottleneck for these full-stack local solutions isn't STT latency, but VAD and noise reduction strategies in noisy environments. Solution A's high integration level allows for quick validation, but when it comes to actual commercialization, environmental adaptation and model lightweighting are where the team's execution truly shows.

Brother Fei

Wait, can you really get it running in two days? I spent the whole night just tweaking the microphone latency for STT... Is recognition accurate in noisy environments?