Voice Loop

byo-key · cascaded voice, timed
turn-taking voice 1.05×

Say something. Then watch what it cost.

This is a cascaded voice agent: speech → text → a text model → speech. Four separate systems in a line, each one waiting on the last. The chat is not the point — the latency budget on the right is. Every turn is broken into the stage that produced it, live, with a running session median.

Hold Talk (or press and hold Space) to take the floor explicitly. Switch to continuous and the browser guesses when you have finished instead — same loop, a different endpoint, usually a different bill.

The raw recogniser output is shown verbatim alongside every turn, because mishearing is the failure mode that actually bites, and a transcript-free voice UI hides it.

What the mic heard

— nothing yet —

Verbatim recogniser output: settled text in solid, in-flight guesses in italic. Recognition confidence, where the browser reports it, is shown on the turn.

Latency budget 0 turns

stagelastmedian
speech-end → transcript
transcript → first token
first token → first audio
stop talking → start hearing

Stage 1 is endpointing plus recogniser finalisation — in continuous mode it includes the settle window the page waits out before it believes you have stopped. Stage 3 includes the wait for enough text to be worth speaking: this page flushes to the synthesiser at the first sentence boundary rather than the end of the reply.

Turn log

    No turns yet.

    idle