Conversational latency is cumulative. Improving one component is useful only when the remaining stages do not dominate the user experience.

Instrument every stage

Timestamp the end of user speech, final or usable ASR result, LLM request, first generated token, first TTS audio and first played audio. Those measurements make optimization empirical rather than subjective.

Streaming changes the latency equation

Streaming ASR, token generation and TTS can overlap work that would otherwise happen serially. The benefit is largest when the orchestration layer can begin useful work before the entire utterance or model response is complete.

Protect natural turn-taking

Aggressive VAD thresholds can reduce latency but increase interruption and truncation errors. The right configuration depends on language, acoustic conditions and caller behavior. Tune latency and conversational correctness together.