Sophon 2.0 is here
Sophon Docs
Sophon Voice

Spoken Replies

What it takes for Sophon to speak a reply, how sentence-by-sentence speech and progress lines work, what Markdown becomes for the ear, and why a reply is sometimes not spoken at all.

Two different things can read a reply to you, and they behave differently. Your browser's own voice reads the finished reply once the turn is over. A configured text-to-speech provider starts speaking while the reply is still being written.

Everything on this page below the first section belongs to the provider path.

What it takes

Your setupWhat speaksWhen you hear it
No text-to-speech providerYour browser's built-in voice, or your phone'sOnce, after the turn finishes
An active provider, with speech output set to Configured provider first or Configured provider onlyThat providerSentence by sentence, while the reply is still being written
Speech output set to OS / browser onlyYour browser's built-in voiceOnce, after the turn finishes

Sentence-by-sentence speech, spoken progress lines, Markdown rewriting, the audio cache and the voice metrics all require an active provider. None of them happens on the browser path. See Set Up Voice for adding a provider, and Settings for the output modes.

Sentence by sentence

As the model writes, the text is split into sentences and each one is synthesized and sent as soon as it is ready.

  • A sentence ends at ., !, ? or ; followed by whitespace or the end of the text. Abbreviations, decimal numbers, ellipses and dots inside a URL do not split.
  • A sentence longer than Sophon:Voice:MaxSpokenSentenceChars (default 280 characters) is split again, preferring a comma, semicolon or colon, and falling back to word boundaries when there is no punctuation to use.
  • Pieces are synthesized strictly in order, with no lookahead. Nothing is sent ahead of the piece in front of it.
  • The trailing sentence is held back until another one starts, so a one-sentence reply gets no early start — you hear it when the turn finishes, like the browser path.
  • The first real audio chunk is what announces "speaking", and it is where time to first audio is recorded.
  • Audio is addressed to the voice session, not to one connection, so every attached tab hears it. See Conversations.

No provider streams audio progressively within a sentence. One sentence at a time is as fine-grained as it gets, for every provider.

Progress lines during long work

During a long, tool-heavy stretch Sophon can speak a short line so a quiet minute does not sound like a dropped connection. The line reads:

Still working — {status}.

All four of these must hold before one is spoken:

ConditionDefault
Sophon:Voice:SpeakProgressUpdates is ontrue
The task reports a new phase, different from the last narrated one—
No spoken reply text for Sophon:Voice:NarrationGapSeconds12 s
At least Sophon:Voice:NarrationIntervalSeconds since the last progress line20 s

Two things follow that are worth knowing before you rely on this:

  • The line reads the raw status text the task reports, not a friendly paraphrase. You will hear things like "Still working — Running tool: web.search."
  • A phase change is required, so a single long step stays silent for its whole run. Progress lines fill the gaps between steps, not the inside of one. A friendlier narration that can speak during a step is roadmap, not shipped.

Whatever the model left half-buffered is flushed and spoken first, so a progress line never fuses onto an unfinished sentence.

Markdown for the ear

On the streaming path, each piece of text is rewritten before it is synthesized.

In the replyRead aloud as
A fenced code block"code shown on screen."
A table"table shown on screen."
A Markdown linkIts link text
A bare URLIts host name
Inline codeThe text, without backticks
Headings, block quotes, list markers, bold and italicStripped
A line break between two lines of proseA sentence break

Splitting happens before rewriting, and the splitter does not know about code fences. A code block containing a ; or a . is cut into pieces first, and each piece is rewritten on its own, so the fence never matches and the code is read aloud with its backticks stripped. The same happens to any code block or table longer than 280 characters. Treat code and tables as things to read in the chat, not things to listen to. A fence-aware split is on the roadmap.

Rewriting also does not run on the paths that do not stream. A turn that produced no streamed text, a spoken error message, and every reply read by the browser's own voice are all read as written, Markdown and all.

What the Voice page shows while it speaks

  • With an active provider, the Voice page and the Command Bridge show no reply text while it is being spoken. The reply is delivered as audio; its text lands in the chat thread.
  • Without a provider, the whole reply appears after the turn finishes, and the browser voice reads it. The same happens on the rare turn that produced no streamed text at all, because that reply is delivered in one piece rather than sentence by sentence.

Either way the reply is a normal message in the chat thread, so the chat voice ribbon shows it in the thread behind the ribbon.

Failover between providers

  • Every sentence is tried against your providers in priority order.
  • Each provider uses its own configured voice. The voice pinned when the session started is only a fallback for a provider that has none of its own, because voice ids are provider-specific.
  • A provider that returns no audio counts as a failure, exactly like one that throws.
  • Every call is time-limited — synthesis at 30 seconds — so a provider that stops responding surfaces as a named error. See Health & Troubleshooting.
  • If every provider fails, Sophon sends a fallback notice naming what failed, and the browser voice speaks the reply when Native speech fallback is allowed. When it is not, you get the error "All configured speech providers failed. …" instead.

The spoken-audio cache

Sentences already spoken by a provider are replayed from an in-memory cache instead of being synthesized and billed again.

  • A 32 MB byte-bounded LRU, held in the Gateway's memory. Nothing is written to disk.
  • The key is the whole sentence plus voice id, language, speed and provider type. A repeated phrase inside a different sentence is a different key and does not hit the cache.
  • It is cleared when the Gateway restarts.
  • The key carries no tenant, so two users on the same Gateway using the same provider, voice, language and speed share the entry for an identical sentence. Per-tenant caching is roadmap.

When a reply is not spoken

When provider speech was expected and cannot happen, the reply carries a red notice naming the actual reason. There are six:

SituationThe notice says
A registered provider failed its health checks, and the browser voice will still speak"Your TTS provider is failing health checks — replies use the browser's voice. Check Voice Center -> TTS providers."
The same, but nothing will speak"Your TTS provider is failing health checks, so replies aren't spoken. Check Voice Center -> TTS providers."
Providers are registered but all inactive, and the browser voice will still speak"Your TTS providers are turned off — replies use the browser's voice."
The same, but nothing will speak"Your TTS providers are turned off, so replies aren't spoken."
Nothing is registered and speech output is Configured provider only"Provider-only speech is on but no TTS provider is available, so replies aren't spoken."
Nothing is registered and Native speech fallback is off"No TTS provider is set up and the browser voice is off, so replies aren't spoken. Turn on the device voice in Voice Center or add a TTS provider."

Two cases deliberately produce no notice. Speech output set to OS / browser only skips provider speech by choice. And a fresh install with nothing registered and native fallback on is the intended zero-configuration path — the browser simply reads the reply, and the availability banner already says so.

In hands-free the notice clears on the next re-arm. In push-to-talk it stays on screen until you press the mic again. The reply itself is in the chat thread in every one of these cases.

Interrupting

Press the mic while Sophon is speaking and the speech stops at the next sentence boundary. The rest of that reply is never spoken, but it still appears in the chat, and the task behind it keeps running — interrupting silences the reply, it does not cancel the work. There is no automatic barge-in: Sophon does not stop because it hears you talking. See Conversations.

Where to go next