Video summary

How I Built an End-to-End Local Voice Agent, and Made It Fast

Main summary

Key takeaways

Technology

End-to-end local voice agent (local “AI stack”)

  • Built a fully local voice agent where the entire pipeline runs on a single 12GB GPU.
  • Uses three core model components:
    • LLM: ~35B Mixture-of-Experts (Qwen3.5 / Qwen 3.6 35B MoE) for “real work,” not just chat.
    • Speech-to-text: Whisper for listening.
    • Text-to-speech / voice: Breeze TTS2 (3.5B) as the talking voice.
  • Adds turn detection for hands-free operation using Cilero VAD (detects when the user starts/stops talking).

Motivation: prior experiments had multi-second latency (5–6s) and sometimes required two machines; this build targets “feels like talking to a real person.”


Key technical goal: reduce latency (and the “silence” problem)

  • The hard part isn’t wiring models—it’s making the system fast enough and responsive enough to avoid awkward waiting.
  • Measured baseline: about 10–11 seconds from user stop talking to first response audio.
  • It gets worse for longer answers because:
    • The system waits for the model to finish writing the whole response before the voice begins speaking.

Performance & responsiveness optimizations (multi-part strategy)

1) Cut long initial silence (streaming response by sentence/phrase)

  • Fix: don’t wait for the entire answer.
  • As the model generates text, the text is chunked into short phrases (e.g., ~3 words) and sent to the TTS engine progressively.
  • This reduces “long silence while writing.”

2) Reduce pauses between sentences (parallelize TTS generation + playback)

  • Fix: while a sentence is being spoken, the system starts generating the next sentence’s audio in advance.
  • Uses a queue so the speaker always has audio ready.
  • Remaining issue: if the next sentence is very long, the queue can sometimes empty briefly.

3) Improve initial “time-to-first-sound”

  • Baseline voice generation speed was near the edge: voice audio generation ~1.1× real time, which can cause occasional gaps.
  • Further fix targets other pipeline waits.

Pipeline-level latency reductions (transcription + thinking behavior)

4) Transcribe while the user is still talking (not after)

  • Previously: waited for speech end → then transcribed the entire recording.
  • Fix: start transcription in parallel every couple seconds while the user speaks.
  • After stop, it still keeps about ~1 second of silence before responding to avoid triggering mid-pause interruptions.

5) “Thinking model” adjustment (faster first response)

  • The Qwen model can “think” before speaking.
  • Fix: disable thinking only for the first response after the user speaks.
  • Later parts (tool calls, actual work) can still allow thinking for quality.

Fit everything into 12GB (practical VRAM accounting + quantization)

  • Relies on careful memory budgeting:
    • Context length tradeoffs: reduced from larger context sizes (e.g., 128k → 64k → 32k) to fit in VRAM.
    • Voice model dominates GPU usage, requiring context reduction.
  • Major speed/memory win:
    • Breeze TTS2 switched from BF16 to Q8 quantized (via audio.cpp), resulting in:
      • 2–3× faster generation
      • ~4GB VRAM freed
    • audio.cpp also supports streaming audio, reducing speaker waiting for full sentence completion.

Agent tooling upgrades: “real work” + visible UI

6) Agent harness: Pythagoras

  • All changes are integrated into the speaker’s agent framework called Pythagoras.
  • Includes UI “work panels” for debugging/visibility:
    • Browser panel
    • Terminal panel
    • Canvas
    • Displays when the AI uses tools (with transitions/animations)

7) Faster tool use: speed up prefill + long tool outputs

  • Tool read cost problem: web pages can be tens of thousands of tokens.
  • Observed: one YouTube page took ~1.5 minutes to read.
  • Cause: micro batch size 128 leading to slow prefill.

Fixes:

  • Free VRAM via voice quantization (Q8).
  • Increase micro batch size to ~1024.
  • Pull expert layers back onto GPU.
  • Restore context range upward (eventually up to ~100,000 tokens), while keeping spare space for spikes.
  • Added efficiency: conversation caching so it doesn’t reread everything every turn; cache persisted to disk and reloaded.

8) Talking while working (avoid silence during tool calls/thinking)

  • New behavior in voice mode:
    • User messages get a tag (e.g., indicating they came from the voice interface).
    • System prompt makes voice replies brief.
  • Before tool use, the AI must say out loud what it will do.
  • Additional fix for “silent thinking/output reading”:
    • When the system starts thinking or waiting on work, it sends pre-chosen ready-made phrases (randomized) like “Let me think about that…”
  • Also handles context overflow:
    • Announces when context is full and it will compact the conversation.

End results (measured improvements)

  • Turn latency improved from ~8 seconds into the ~2–3 second range across turns.
  • Faster agent behavior in practice:
    • Example: open YouTube → search channel → describe first image detected on the page.
    • Example: switch to local llama subreddit → generate a report summarizing what’s trending (tool+web workflow).

Build philosophy / extensibility

  • The interface and orchestration are a major part of “feels like the future,” not only raw intelligence.
  • Modular approach:
    • The model is the main swappable component, and improvements to local models can be dropped in.
  • Configuration is portable:
    • Numbers depend on your card, but the strategy (context/compute buffering/voice quantization) stays the same.
  • Code is available as open source: integrated in Pythagoras on GitHub.
  • Encourages using it and submitting PRs for gaps.

Main speakers / sources

  • Speaker: the video creator (self-referential “I” throughout; presenting measurements and building steps).
  • External sources mentioned:
    • OpenAI (reference to the Astra demo as inspiration)
    • Qwen (Qwen3.5 / Qwen MoE models referenced)
    • Whisper (speech-to-text)
    • Breeze TTS2 (voice model; referenced as top in an “Artificial analysis leaderboard”)
    • Cilero VAD (turn detection)
    • CodeRabbit (sponsor; PR reviewer tool used to catch race conditions)
    • GitHub (Pythagoras open-source repository)

Original video