Spoken Replies
What it takes for Sophon to speak a reply, how sentence-by-sentence speech and progress lines work, what Markdown becomes for the ear, and why a reply is sometimes not spoken at all.
Two different things can read a reply to you, and they behave differently. Your browser's own voice reads the finished reply once the turn is over. A configured text-to-speech provider starts speaking while the reply is still being written.
Everything on this page below the first section belongs to the provider path.
What it takes
| Your setup | What speaks | When you hear it |
|---|---|---|
| No text-to-speech provider | Your browser's built-in voice, or your phone's | Once, after the turn finishes |
| An active provider, with speech output set to Configured provider first or Configured provider only | That provider | Sentence by sentence, while the reply is still being written |
| Speech output set to OS / browser only | Your browser's built-in voice | Once, after the turn finishes |
Sentence-by-sentence speech, spoken progress lines, Markdown rewriting, the audio cache and the voice metrics all require an active provider. None of them happens on the browser path. See Set Up Voice for adding a provider, and Settings for the output modes.
Sentence by sentence
As the model writes, the text is split into sentences and each one is synthesized and sent as soon as it is ready.
- A sentence ends at
.,!,?or;followed by whitespace or the end of the text. Abbreviations, decimal numbers, ellipses and dots inside a URL do not split. - A sentence longer than
Sophon:Voice:MaxSpokenSentenceChars(default 280 characters) is split again, preferring a comma, semicolon or colon, and falling back to word boundaries when there is no punctuation to use. - Pieces are synthesized strictly in order, with no lookahead. Nothing is sent ahead of the piece in front of it.
- The trailing sentence is held back until another one starts, so a one-sentence reply gets no early start — you hear it when the turn finishes, like the browser path.
- The first real audio chunk is what announces "speaking", and it is where time to first audio is recorded.
- Audio is addressed to the voice session, not to one connection, so every attached tab hears it. See Conversations.
No provider streams audio progressively within a sentence. One sentence at a time is as fine-grained as it gets, for every provider.
Progress lines during long work
During a long, tool-heavy stretch Sophon can speak a short line so a quiet minute does not sound like a dropped connection. The line reads:
Still working — {status}.All four of these must hold before one is spoken:
| Condition | Default |
|---|---|
Sophon:Voice:SpeakProgressUpdates is on | true |
| The task reports a new phase, different from the last narrated one | — |
No spoken reply text for Sophon:Voice:NarrationGapSeconds | 12 s |
At least Sophon:Voice:NarrationIntervalSeconds since the last progress line | 20 s |
Two things follow that are worth knowing before you rely on this:
- The line reads the raw status text the task reports, not a friendly paraphrase. You will hear things like "Still working — Running tool: web.search."
- A phase change is required, so a single long step stays silent for its whole run. Progress lines fill the gaps between steps, not the inside of one. A friendlier narration that can speak during a step is roadmap, not shipped.
Whatever the model left half-buffered is flushed and spoken first, so a progress line never fuses onto an unfinished sentence.
Markdown for the ear
On the streaming path, each piece of text is rewritten before it is synthesized.
| In the reply | Read aloud as |
|---|---|
| A fenced code block | "code shown on screen." |
| A table | "table shown on screen." |
| A Markdown link | Its link text |
| A bare URL | Its host name |
| Inline code | The text, without backticks |
| Headings, block quotes, list markers, bold and italic | Stripped |
| A line break between two lines of prose | A sentence break |
Splitting happens before rewriting, and the splitter does not know about code fences. A code block containing a ; or a . is cut into pieces first, and each piece is rewritten on its own, so the fence never matches and the code is read aloud with its backticks stripped. The same happens to any code block or table longer than 280 characters. Treat code and tables as things to read in the chat, not things to listen to. A fence-aware split is on the roadmap.
Rewriting also does not run on the paths that do not stream. A turn that produced no streamed text, a spoken error message, and every reply read by the browser's own voice are all read as written, Markdown and all.
What the Voice page shows while it speaks
- With an active provider, the Voice page and the Command Bridge show no reply text while it is being spoken. The reply is delivered as audio; its text lands in the chat thread.
- Without a provider, the whole reply appears after the turn finishes, and the browser voice reads it. The same happens on the rare turn that produced no streamed text at all, because that reply is delivered in one piece rather than sentence by sentence.
Either way the reply is a normal message in the chat thread, so the chat voice ribbon shows it in the thread behind the ribbon.
Failover between providers
- Every sentence is tried against your providers in priority order.
- Each provider uses its own configured voice. The voice pinned when the session started is only a fallback for a provider that has none of its own, because voice ids are provider-specific.
- A provider that returns no audio counts as a failure, exactly like one that throws.
- Every call is time-limited — synthesis at 30 seconds — so a provider that stops responding surfaces as a named error. See Health & Troubleshooting.
- If every provider fails, Sophon sends a fallback notice naming what failed, and the browser voice speaks the reply when Native speech fallback is allowed. When it is not, you get the error "All configured speech providers failed. …" instead.
The spoken-audio cache
Sentences already spoken by a provider are replayed from an in-memory cache instead of being synthesized and billed again.
- A 32 MB byte-bounded LRU, held in the Gateway's memory. Nothing is written to disk.
- The key is the whole sentence plus voice id, language, speed and provider type. A repeated phrase inside a different sentence is a different key and does not hit the cache.
- It is cleared when the Gateway restarts.
- The key carries no tenant, so two users on the same Gateway using the same provider, voice, language and speed share the entry for an identical sentence. Per-tenant caching is roadmap.
When a reply is not spoken
When provider speech was expected and cannot happen, the reply carries a red notice naming the actual reason. There are six:
| Situation | The notice says |
|---|---|
| A registered provider failed its health checks, and the browser voice will still speak | "Your TTS provider is failing health checks — replies use the browser's voice. Check Voice Center -> TTS providers." |
| The same, but nothing will speak | "Your TTS provider is failing health checks, so replies aren't spoken. Check Voice Center -> TTS providers." |
| Providers are registered but all inactive, and the browser voice will still speak | "Your TTS providers are turned off — replies use the browser's voice." |
| The same, but nothing will speak | "Your TTS providers are turned off, so replies aren't spoken." |
| Nothing is registered and speech output is Configured provider only | "Provider-only speech is on but no TTS provider is available, so replies aren't spoken." |
| Nothing is registered and Native speech fallback is off | "No TTS provider is set up and the browser voice is off, so replies aren't spoken. Turn on the device voice in Voice Center or add a TTS provider." |
Two cases deliberately produce no notice. Speech output set to OS / browser only skips provider speech by choice. And a fresh install with nothing registered and native fallback on is the intended zero-configuration path — the browser simply reads the reply, and the availability banner already says so.
In hands-free the notice clears on the next re-arm. In push-to-talk it stays on screen until you press the mic again. The reply itself is in the chat thread in every one of these cases.
Interrupting
Press the mic while Sophon is speaking and the speech stops at the next sentence boundary. The rest of that reply is never spoken, but it still appears in the chat, and the task behind it keeps running — interrupting silences the reply, it does not cancel the work. There is no automatic barge-in: Sophon does not stop because it hears you talking. See Conversations.
Where to go next
- Conversations — hands-free, interrupting, reconnects and session limits
- Approvals & Questions — which voice reads a prompt, and why it can be silent
- Set Up Voice — adding a text-to-speech provider and picking a voice
- Providers — the six provider types and what each one supports
- Health & Troubleshooting — health checks, time limits and the voice metrics
Conversations & Hands-Free
Push-to-talk and conversation mode, how a spoken turn runs, what interrupting does and does not stop, reconnects and session limits, and how long Sophon waits for a transcript.
Approvals & Questions by VoiceNEW
How Sophon reads an approval or a question out loud, exactly what it says, how your spoken answer is matched, the timeouts, and what happens when nobody answers.