A voice agent is a loop that hears speech, decides, and speaks back — often while the user is still talking. The hard parts are not the model names. They are turn-taking, barge-in, and what the product does when the network hiccups in a kitchen in Manisa.
We run voice on Polylingo practice and on a small internal ops line. The practice loop is forgiving: a learner repeats a phrase, we score it, we speak a correction. The ops line is not: it must not invent a deployment status. Two products, two policies, one pipeline shape — stream in, decide, stream out, never wait for a perfect transcript before we start thinking.
Budgets we write on the ticket
Time-to-first-audio under 800 ms on a median mobile network. Full turn under two seconds for practice. If we miss the budget, we shrink the model or skip a flourish — we do not add a spinner that says “thinking” for three seconds. People hang up on spinners they cannot see.
Barge-in is a feature, not a bug. When the learner starts speaking over the prompt, we cancel TTS and listen. Implementing that on the web means a playback graph we can stop, not a blob we fire and forget. On mobile it means the audio session is ours, not the OS default that ducks us into silence.
Transcripts are data, not the product
We show the recognised line so the learner can correct a name. We do not fine-tune on raw child audio. Retention follows the same inventory as the AI Act piece. A voice feature that cannot explain where the waveform went does not ship.
Fallback is a designed state
If the socket dies, we finish the current sentence, switch to text, and say so. If confidence is low, we ask one clarifying question instead of guessing. Guessing in voice feels like the product is arguing. Clarifying feels like a teacher.
Questions we keep getting
Do you build your own speech model? No. We buy ASR and TTS, own the dialogue policy, and keep the product logic in .NET. The moat is the lesson graph, not the spectrogram.
Why not a single speech-to-speech model? We tried it in a spike. Great latency, weak control. We could not reliably pin a CEFR band or a forbidden topic. Until we can test those, we keep the pipeline split.
Is this SignalR? The transport is often WebSockets. The product concern is the state machine. We wrote about SignalR separately — this article is about the voice loop that rides on top.