Sophon 2.0 is here
Sophon Docs
Sophon Voice

Limits

The honest list — what Sophon Voice does not do in 2.0, stated plainly and framed against the roadmap, so you know before you rely on it.

2.0 rebuilt the voice engine. Several things people reasonably expect from voice software are not part of it yet. This page is the list, in one place, with what to do instead. Nothing here is a bug report — these are known positions, and the ones that are planned say so.

Listening

  • There is no wake word and no always-on listening. Sophon listens after you press the mic or arm the Command Bridge, and it stops on its own after the silence timeout or the inactivity limit. Nothing listens in the background.
  • There is no talk-over detection. Speaking over Sophon does not stop it. You interrupt by pressing the mic, and that silences the reply, not the task behind it — work already under way keeps running, and by default your next spoken request queues behind it. Roadmap: detecting speech over playback. A configuration key, Sophon:Voice:BargeInMinWords, is reserved for it and is read by nothing today.
  • Only Deepgram shows your words while you speak. The other five speech-to-text providers are buffered: nothing is sent until you stop talking, and the transcript arrives in one piece, one round trip per utterance. In the Dashboard that means the preview card fills in after you finish rather than during.
  • End-of-utterance detection is a silence detector, not an understanding of what you said. For the five buffered providers Sophon waits for at least 300 ms of real speech and then for the configured endpointing gap of silence. A thoughtful pause mid-sentence can end your turn. Roadmap: semantic end-of-turn detection.
  • Without a provider, voice needs a browser that has speech recognition. Chrome or Edge. Firefox has none: the mic button still appears and fails when pressed, and hands-free is not offered.
  • Sophon Mobile has no live transcription through the Gateway. It listens with the phone's own recognizer, and the optional Gateway fallback uploads the finished recording after you stop. Mobile cannot stream audio to the server.

Speaking

  • Speech needs a text-to-speech provider. With no provider, the browser's or the phone's own voice reads the finished reply once, at the end. Sentence-by-sentence speech, spoken progress lines, Markdown rewriting, the audio cache and the voice metrics all require an active provider.
  • Speech is streamed a sentence at a time, never word by word. No provider streams audio progressively within a sentence, and a one-sentence reply therefore gets no early start — you hear it when the turn finishes.
  • Code and tables can be read aloud. A reply is split into sentences before it is rewritten for the ear, and the splitter does not know about code fences. A code block containing a ; or a . , and any code block or table over 280 characters, is cut into pieces first, so the fence never matches and the code is read out with its backticks stripped. Read code and tables in the chat. Roadmap: a fence-aware split.
  • Rewriting for the ear runs on the streaming path only. A turn that produced no streamed text, a spoken error message, and every reply read by the browser's own voice are read as written, Markdown and all — even when a provider is active.
  • A spoken progress line reads the raw status the task reports, not a friendly paraphrase, and a phase change is required, so a single long step stays silent for its whole run. Progress lines fill the gaps between steps, not the inside of one.
  • While a provider speaks, the Voice page does not show the reply. You get the orb and your own words; the full reply text lands in the chat thread.
  • The spoken-audio cache is not per tenant. Two users on the same Gateway sharing a provider, voice, language and speed share the cache entry for an identical sentence. It is in memory and is cleared on restart. Roadmap: per-tenant caching.
  • Finer-grained synthesis exists but is unused. One provider implements progressive, chunk-by-chunk synthesis of a single sentence, and the capabilities endpoint reports supportsStreaming: true for it, but nothing calls it — every provider is asked for one complete sentence at a time.

Approvals and questions

  • Prompts are read in the browser's built-in voice, or the phone's on Sophon Mobile — never in the text-to-speech provider voice you picked, even when one is active and speaking the replies. Turn the browser voice off, or set speech output to Configured provider only, and prompts are not read aloud at all: you get the card, the buttons and the mic.
  • The Dashboard's approval card hint understates what High and Critical need. It reads Tap a button, or say "yes" / "no" at every risk level, while only "yes, approve" (or "yes approved") actually approves at High and Critical. The spoken prompt is correct on both surfaces, so listen to it rather than the hint. Sophon Mobile's hint is already risk-aware. Roadmap: bring the Dashboard card in line with Mobile's.
  • Spoken option numbers cover one to four, and only in an answer of four words or fewer. "I think option two is best" is treated as words, not as picking option 2. Options five and up are read out with their numbers but can only be chosen by name or by tapping. Roadmap: wider number matching.
  • Symbol matching assumes your recognizer returns words. C++ is read and matched as "C plus plus", C# as "C sharp". A recognizer that writes the literal symbol back instead will not match the option, and you are asked again — tap it instead. This has not been verified against every recognizer.
  • Negating an option on a fixed-choice question re-asks. "Email, no rush" does not pick Email, even though you named it.
  • Sophon Mobile does not read questions aloud. Approvals are spoken and can be answered out loud; questions are answered on screen.
  • The chat voice ribbon never reads prompts aloud. They stay on the chat's own cards, where you tap them.
  • Approvals are never spoken to a paired node. A task started from a node voice session raises its approval in the usual places — the Dashboard, Mobile, the CLI, or the channel the conversation came from — and waits there.

Surfaces

  • The CLI manages voice but cannot hold a conversation. sophon voice status, settings, stt-providers, providers, transcribe and runtime configure and inspect voice. There is no spoken conversation inside the CLI, and none is planned for this release.
  • The desktop app has no voice features of its own. There is no global push-to-talk shortcut, no tray voice mode and no voice overlay. It hosts the Dashboard, so the Dashboard's voice surfaces are present inside it; how well browser speech recognition works in that runtime is not something this release verifies.
  • Sophon Node takes no part in voice conversations. It cannot listen, has no wake word, and out of the box it does not speak replies either: the only path that makes a node speak is off by default and no shipped client calls it. See Privacy & Data.
  • The Voice page and the Command Bridge start a new chat session with the default agent every time, with no agent picker. Pass ?sessionId= in the Voice page's URL to attach to an existing chat instead.
  • Ending voice in one tab ends it in every tab attached to that conversation.
  • The Command Bridge is hands-free only. There is no push-to-talk on the Bridge.
  • Switching chats with the voice ribbon open restarts voice, listening, on the new chat. It does not simply stop.

Running it

  • Voice turns use a smaller, voice-focused tool set than a typed chat turn: five core tools, a base set of up to 16, and more loadable on demand through tool search, with no tool index. Do not assume a voice turn can reach everything a chat turn can — some jobs are better typed.
  • Speech providers cannot be reordered or disabled after they are added. Priority is set at add time. The only operations on an existing provider are test, change its voice (text-to-speech only), and remove. To change a priority, remove the provider and add it again.
  • Providers are configured per Gateway, not per tenant. Every user on the Gateway sees the same list, and managing it is Admin-only.
  • Provider health status trails an outage and does not survive a restart. Two consecutive failed checks at a 15-minute interval are needed before a provider is flagged, so it can take roughly 30 minutes, and a Gateway restart resets every provider until the next sweep. The banner also stays silent while any provider of that kind is still working.
  • There are no speech-to-text metrics. The three voice instruments cover time to first audio, per-provider synthesis latency and text-to-speech provider failures, and all three are recorded on the text-to-speech path only. Nothing measures transcription latency or speech-to-text failures. The metrics endpoint requires authentication.
  • Speech providers are built in, not pluggable. The six speech-to-text and six text-to-speech types are the whole list. Roadmap: speech providers from plugins.
  • Voice shares the chat connection. A dedicated realtime voice transport was designed for this release and deliberately deferred; every voice surface runs over the same /hubs/chat connection chat uses. Nothing in 2.0 assumes a separate voice hub, and integrations should not either.

Two settings are labelled in the product as not doing what their name suggests, and the labelling is accurate. Host defaults → Default language is marked "Not used yet" — recognition and synthesis always use each person's own language. Sophon STT fallback is read only by Sophon Mobile. Both are covered on Voice Settings.

Where to go next