Sophon 2.0 is here
Sophon Docs
Sophon Voice

Sophon VoiceNEW

Talk to Sophon instead of typing. A spoken request runs as a regular agent task on your chat, queued and approval-gated the same way, and the answer comes back as speech.

Voice is another way into the same agent task, not a separate assistant. You press the mic, say what you want, and Sophon submits it to your chat session exactly as it would a typed message: the same queue, the same approval gates, the same session policy. Your words and the reply are stored in the thread like typed messages, and the turn shows up live in the chat while it runs.

Sophon 2.0 rebuilt the voice engine underneath. It is not a new product: the surfaces look much as they did, but listening, speaking, interrupting and answering prompts now follow one set of rules everywhere.

What a voice turn is

A spoken request becomes an ordinary agent task on the chat session the voice conversation is attached to.

  • It is queued like a typed message. The default session policy runs one task at a time and holds up to three more, so a new spoken request waits behind whatever is already running.
  • It is approval-gated like a typed message. Nothing skips the gate because it was asked for out loud, and answering by voice is held to a stricter standard at High and Critical risk.
  • It is stored like a typed message, in the same thread and the same history.
  • It runs with a smaller, voice-focused tool set than a chat turn. The core tools are datetime.now, memory.search, memory.write, message.send and web.search, with a base set of up to 16 tools; the agent can load more on demand through tool search. Do not assume a voice turn can reach everything a chat turn can.

See Conversations & Hands-Free for the mechanics.

Where you can talk

SurfaceWhat it is
Voice pageA full-screen page in the Dashboard: orb, live transcript, availability banner, approval and question cards
Chat voice ribbonA strip inside a chat, talking on that chat. Prompts stay on the chat's own cards
Command BridgeThe fleet view's hands-free voice bar. Arm it and it listens; it turns itself off and shows why
Sophon MobileThe Voice tab, using your phone's own speech recognition by default

Two more places touch voice without being places you talk:

  • The CLI manages voice; it cannot hold a voice conversation. sophon voice status, sophon voice settings, sophon voice stt-providers, sophon voice providers, sophon voice transcribe and sophon voice runtime configure and inspect voice. There is no spoken conversation inside the CLI, and none is planned for this release.
  • A paired node can speak, but only if an operator turns it on. A node reads text aloud through its operating system's own speech synthesiser. It cannot listen, has no microphone capture and no wake word, and the Gateway routes that drive it are off by default.

Per-surface rules, including what each one does and does not do with approvals and questions, are in Voice Surfaces.

What 2.0 changed

  • One engine. The Voice page, the chat voice ribbon and the Command Bridge run on the same state machine, so listening, playback and interrupting follow the same rules on all three.
  • Turns are queued agent tasks. A voice request is submitted the same way a typed one is, appears live in the chat thread, and a failure is spoken or shown rather than dropped.
  • Sessions belong to the chat, not the connection. Reconnects re-attach to the same voice session, and every open tab attached to that conversation hears the reply.
  • Speech starts before the reply is finished. With a text-to-speech provider, the reply is spoken sentence by sentence while the model is still writing it.
  • Hands-free works with every speech-to-text provider. A server-side silence detector supplies the end of an utterance for the five providers that cannot detect it themselves.
  • Approvals and questions can be answered out loud. High- and Critical-risk approvals need you to say "yes, approve"; questions are read out with numbered choices; an unclear answer is asked again at most twice and never guessed at.
  • Providers are watched. Configured providers are health-checked in the background, the Voice page tells "not set up" apart from "failing" and names the failing provider, and every provider call is time-limited.
  • sophon voice status lists every configured provider with its health and reports server and browser hands-free separately.

What you need

Nothing is preconfigured, and voice works with no speech provider at all. Adding providers changes what you get, not whether it works.

SetupListeningReplies
No speech providerYour browser's own speech recognition, in Chrome or Edge. It ends the utterance at a pause. Firefox has none.Your browser's built-in voice reads the whole reply once the turn has finished
+ a speech-to-text providerServer-side transcription. Hands-free works with any provider. With Deepgram your words appear on screen in the Dashboard while you are still speaking; the other five transcribe after you stopUnchanged
+ a text-to-speech providerUnchangedThe reply is spoken sentence by sentence while it is still being written, long multi-step turns can get a short spoken progress line, Markdown is rewritten for the ear, and repeated sentences are replayed from an in-memory cache instead of being synthesized again

With no text-to-speech provider there is no sentence-by-sentence speech, no spoken progress line and no rewriting for the ear: the browser voice reads the reply text as it stands.

The Voice page says which of these you are on. An info-tone banner means a side isn't configured and the browser is being used; a warning-tone banner means a configured provider is failing, and it names that provider. See Health & Troubleshooting.

Availability

Voice is available on every Sophon tier. No tier or licence check gates listening, speaking, hands-free, approvals by voice or any surface.

Provider management is Admin-only. Adding, removing and testing speech providers, and changing host-wide listening defaults, require the Admin role. Everyone can change their own personal preferences — language, speed, conversation mode, their own listening overrides — and everyone can use voice. Provider configuration is host-wide, not per tenant.

What voice is not

Stated plainly, because these are the things people reasonably expect:

  • It is not on-device. With a provider, your audio goes through your Gateway to that vendor. With no provider, the browser or phone recognizer decides where the audio goes, which for Chrome means Google's speech service and for Safari means Apple's. Never assume audio stays on your machine. See Privacy & Data.
  • There is no wake word. Sophon listens after you press the mic or arm the Command Bridge, and stops on its own after silence or inactivity.
  • There is no automatic barge-in. Talking over Sophon does not stop it. You interrupt by pressing the mic. Detecting speech over playback is on the roadmap; a configuration key is reserved for it and does nothing today.
  • Interrupting silences the reply, not the task. Work already under way keeps running, and its reply still lands in the chat thread.
  • The CLI cannot hold a conversation, and a node cannot listen.

The full list, with what to do instead, is on Limits.

Where to go next