Sophon VoiceNEW
Talk to Sophon instead of typing. A spoken request runs as a regular agent task on your chat, queued and approval-gated the same way, and the answer comes back as speech.
Voice is another way into the same agent task, not a separate assistant. You press the mic, say what you want, and Sophon submits it to your chat session exactly as it would a typed message: the same queue, the same approval gates, the same session policy. Your words and the reply are stored in the thread like typed messages, and the turn shows up live in the chat while it runs.
Sophon 2.0 rebuilt the voice engine underneath. It is not a new product: the surfaces look much as they did, but listening, speaking, interrupting and answering prompts now follow one set of rules everywhere.
What a voice turn is
A spoken request becomes an ordinary agent task on the chat session the voice conversation is attached to.
- It is queued like a typed message. The default session policy runs one task at a time and holds up to three more, so a new spoken request waits behind whatever is already running.
- It is approval-gated like a typed message. Nothing skips the gate because it was asked for out loud, and answering by voice is held to a stricter standard at High and Critical risk.
- It is stored like a typed message, in the same thread and the same history.
- It runs with a smaller, voice-focused tool set than a chat turn. The core tools are
datetime.now,memory.search,memory.write,message.sendandweb.search, with a base set of up to 16 tools; the agent can load more on demand through tool search. Do not assume a voice turn can reach everything a chat turn can.
See Conversations & Hands-Free for the mechanics.
Where you can talk
| Surface | What it is |
|---|---|
| Voice page | A full-screen page in the Dashboard: orb, live transcript, availability banner, approval and question cards |
| Chat voice ribbon | A strip inside a chat, talking on that chat. Prompts stay on the chat's own cards |
| Command Bridge | The fleet view's hands-free voice bar. Arm it and it listens; it turns itself off and shows why |
| Sophon Mobile | The Voice tab, using your phone's own speech recognition by default |
Two more places touch voice without being places you talk:
- The CLI manages voice; it cannot hold a voice conversation.
sophon voice status,sophon voice settings,sophon voice stt-providers,sophon voice providers,sophon voice transcribeandsophon voice runtimeconfigure and inspect voice. There is no spoken conversation inside the CLI, and none is planned for this release. - A paired node can speak, but only if an operator turns it on. A node reads text aloud through its operating system's own speech synthesiser. It cannot listen, has no microphone capture and no wake word, and the Gateway routes that drive it are off by default.
Per-surface rules, including what each one does and does not do with approvals and questions, are in Voice Surfaces.
What 2.0 changed
- One engine. The Voice page, the chat voice ribbon and the Command Bridge run on the same state machine, so listening, playback and interrupting follow the same rules on all three.
- Turns are queued agent tasks. A voice request is submitted the same way a typed one is, appears live in the chat thread, and a failure is spoken or shown rather than dropped.
- Sessions belong to the chat, not the connection. Reconnects re-attach to the same voice session, and every open tab attached to that conversation hears the reply.
- Speech starts before the reply is finished. With a text-to-speech provider, the reply is spoken sentence by sentence while the model is still writing it.
- Hands-free works with every speech-to-text provider. A server-side silence detector supplies the end of an utterance for the five providers that cannot detect it themselves.
- Approvals and questions can be answered out loud. High- and Critical-risk approvals need you to say "yes, approve"; questions are read out with numbered choices; an unclear answer is asked again at most twice and never guessed at.
- Providers are watched. Configured providers are health-checked in the background, the Voice page tells "not set up" apart from "failing" and names the failing provider, and every provider call is time-limited.
sophon voice statuslists every configured provider with its health and reports server and browser hands-free separately.
What you need
Nothing is preconfigured, and voice works with no speech provider at all. Adding providers changes what you get, not whether it works.
| Setup | Listening | Replies |
|---|---|---|
| No speech provider | Your browser's own speech recognition, in Chrome or Edge. It ends the utterance at a pause. Firefox has none. | Your browser's built-in voice reads the whole reply once the turn has finished |
| + a speech-to-text provider | Server-side transcription. Hands-free works with any provider. With Deepgram your words appear on screen in the Dashboard while you are still speaking; the other five transcribe after you stop | Unchanged |
| + a text-to-speech provider | Unchanged | The reply is spoken sentence by sentence while it is still being written, long multi-step turns can get a short spoken progress line, Markdown is rewritten for the ear, and repeated sentences are replayed from an in-memory cache instead of being synthesized again |
With no text-to-speech provider there is no sentence-by-sentence speech, no spoken progress line and no rewriting for the ear: the browser voice reads the reply text as it stands.
The Voice page says which of these you are on. An info-tone banner means a side isn't configured and the browser is being used; a warning-tone banner means a configured provider is failing, and it names that provider. See Health & Troubleshooting.
Availability
Voice is available on every Sophon tier. No tier or licence check gates listening, speaking, hands-free, approvals by voice or any surface.
Provider management is Admin-only. Adding, removing and testing speech providers, and changing host-wide listening defaults, require the Admin role. Everyone can change their own personal preferences — language, speed, conversation mode, their own listening overrides — and everyone can use voice. Provider configuration is host-wide, not per tenant.
What voice is not
Stated plainly, because these are the things people reasonably expect:
- It is not on-device. With a provider, your audio goes through your Gateway to that vendor. With no provider, the browser or phone recognizer decides where the audio goes, which for Chrome means Google's speech service and for Safari means Apple's. Never assume audio stays on your machine. See Privacy & Data.
- There is no wake word. Sophon listens after you press the mic or arm the Command Bridge, and stops on its own after silence or inactivity.
- There is no automatic barge-in. Talking over Sophon does not stop it. You interrupt by pressing the mic. Detecting speech over playback is on the roadmap; a configuration key is reserved for it and does nothing today.
- Interrupting silences the reply, not the task. Work already under way keeps running, and its reply still lands in the chat thread.
- The CLI cannot hold a conversation, and a node cannot listen.
The full list, with what to do instead, is on Limits.
Where to go next
- Voice Surfaces — every place you can talk, and what differs between them
- Conversations & Hands-Free — push-to-talk, conversation mode, interrupting, reconnects
- Spoken Replies — sentence streaming, progress lines, Markdown for the ear
- Approvals & Questions — the risk rule, re-asks, and answering out loud
- Set Up Voice — the zero-config path, then adding providers
- Privacy & Data — where audio goes on each path, and what is stored
Scripting & Automation
Use the Sophon CLI in scripts and CI — JSON output, exit codes, piping, and shell completions.
Voice Surfaces
Every place you can talk to Sophon — the Voice page, a chat's voice ribbon, the Command Bridge and Sophon Mobile — plus what a paired node and the desktop app do, and the rules that differ between them.