Sophon Voice
Say it out loud. The same gates apply.
Voice is another way into the runtime you already run, not a second assistant. A spoken request runs as a regular task on your chat, with a smaller, voice-focused tool set. It is queued and approval-gated the same way as a typed message and appears live in the chat thread. When an action needs your approval, the Voice page reads it out, and High- or Critical-risk actions need you to say “yes, approve”.
On every tier. Nothing is preconfigured: with no speech provider, voice uses your browser’s own speech recognition and voice (Chrome or Edge).
Animation of the Sophon Voice page: a spoken request to delete a schedule raises a High-risk approval; Sophon reads it out; a bare “yes” is asked again; “yes, approve” approves it; the reply is spoken and the microphone reopens.
Recreated from the Dashboard’s Voice page with conversation mode on, Deepgram for speech-to-text and a text-to-speech provider. Lines under the frame caption what is said aloud, plus a few notes; none of them is on the page. Approval prompts are read in the browser’s own voice, and a reply spoken by a provider is heard, not printed: its text lands in the chat thread.
A voice turn
A spoken request is an ordinary task.
Voice isn’t a second assistant. Sophon runs what you say as a regular agent task on your chat, with a smaller, voice-focused tool set, stores it in the thread, gates it with the same approvals as a typed message, and then speaks the answer.
Same queue, same thread
A voice request runs as a regular agent task on your chat session. It is queued and approval-gated the same way as a typed message and appears live in the chat thread.
Voice turns use a smaller, voice-focused tool set, so some jobs are better typed.
Failures are said, not swallowed
If a voice request fails, Sophon tells you by voice or on screen instead of going quiet.
A turn you interrupted is the exception: it stays silent by design, and its text still reaches the chat.
Interrupting silences the reply, not the work
On the Voice page, press the mic while Sophon is thinking or speaking to silence it and talk. The interrupted reply isn’t spoken later, but it still appears in the chat.
Interrupting stops the speech, not the task: work already under way keeps running, and by default your next request waits behind it. There’s no talk-over detection; interrupting is always a press.
Approvals by voice
Risk decides what counts as yes.
Which actions ask at all is set by your approval policy, the same one typed requests answer to. Voice changes how you answer: on the Voice page and the Command Bridge, Sophon closes the mic, reads out the kind of action and its risk, and sets a higher bar for the riskier ones.
Low and Medium
A plain “yes” approves. A plain “no” rejects.
High and Critical
Only “yes, approve” approves. A bare “yes” isn’t enough, and Sophon asks again.
A yes with a but
A “yes” that also says “no”, “wait” or “don’t” counts as unclear, so Sophon asks again.
Two re-asks, then a tap
Unclear answers are asked again at most twice; after that the mic closes and the prompt waits for a tap or a mic press. Sophon never guesses.
The approval’s own timer keeps running, and an approval nobody answers is rejected when it expires.
What you can say
| You say | Low or Medium | High or Critical |
|---|---|---|
| “Yes.” | Approved | Asked again |
| “Yes, approve.” | Approved | Approved |
| “Yes… wait.” | Asked again | Asked again |
| “No.” | Rejected | Rejected |
Asked again means Sophon repeats what it needs, at most twice, then waits for a tap or a mic press.
When the agent asks
Numbered choices. Nothing guessed.
When an agent asks a question with options, the Voice page and the Command Bridge read it out in your browser’s voice with numbered choices. Answer by saying the number (one to four) or the option’s name, or tap it.
One question at a time
Multi-part questions are asked one at a time, and your answers are sent together at the end. If you can pick more than one, Sophon says so.
Say it the way you’d say it
Option names with symbols are read out and matched as spoken words, so you can say “C plus plus” or “C sharp”.
“Don’t” means don’t
If you negate an option (“don’t send”), Sophon won’t pick it. It asks again, or passes your own words to the agent when the question accepts a free-text answer.
Short answers for numbers
A spoken number counts in a short answer. Options past the fourth are picked by name or tap.
If your speech recognizer writes the symbol instead of the words, Sophon asks again; tap the option. Question cards come from plans and other tasks on the chat: a voice turn is set up to ask its own follow-ups out loud, in conversation, rather than as a card. Sophon Mobile shows questions on screen and doesn’t read them out.
A Sophon Voice question card: question one of two, “Which calendar should the review go on?”, with the options Work, Personal and Board. It is answered out loud or by tapping an option.
One decision, recorded once
Answer anywhere. The rest close.
An approval is one record on your Gateway, not a dialog on one screen. The first answer decides it.
Where a prompt is asked out loud
Voice only speaks prompts that belong to its own conversation, so an approval from an unrelated session never interrupts you.
Closed everywhere
Answer an approval in one place, whether by voice, on a chat card, in another tab or in the mobile app, and any other open chat card or voice prompt for it closes on its own.
The Dashboard’s Approvals page catches up on its next refresh.
Recorded exactly once
Each approval is recorded exactly once, so when two devices answer the same card only one answer counts, and an approval that timed out or was cancelled can never later become an approval.
Says how it ended
Approval cards on the Dashboard and Mobile now say how an approval ended: Approved, Rejected, Cancelled or Timed out.
If a plan pauses for your approval and finishes later, its result is spoken to a voice session still open on that chat. If you’re mid-conversation at that moment, it appears in the chat thread instead.
What changed in 2.0
Same orb. New engine.
If you used voice in 1.x, the Voice page will look familiar: the look barely changed. The work went underneath: one voice engine behind the Voice page, the chat ribbon and the Command Bridge, and a voice session that belongs to the conversation, not to a browser connection.
Drop the connection, keep the conversation
If your connection drops while Sophon is answering, voice picks back up when you reconnect: you hear the rest of the reply from that point, and the full reply is in the chat.
Audio sent while you were offline isn’t replayed.
Every tab on the conversation
Every open tab attached to the same voice conversation hears the reply.
Ending voice in one tab ends it in all of them, and each visit to the Voice page opens a new conversation.
Nothing looks fine when it isn’t
Every voice error now shows in a strip across the top of the Voice page, and the microphone is released every time you leave voice, so there’s no lingering recording indicator.
With server speech-to-text, each transcript is matched to what you said when it was said, so a slow transcription of an earlier sentence is much less likely to be taken as your answer or as a new request.
Where you can talk to it
The CLI manages voice — providers, settings and sophon voice status — but has no voice conversation of its own.
Listening and speaking
Hands-free with any provider. Or none.
Conversation mode works end to end, your browser’s own speech is enough to start, and with a text-to-speech provider Sophon starts talking before the reply is finished.
Listening
Conversation mode
Turn on Conversation mode in Voice settings and your first mic press starts a hands-free conversation: speak, pause, hear the answer, and the mic reopens only after the reply has finished playing. Silence and inactivity limits close the mic when you walk away.
It isn’t a wake word or always-on listening: it starts with your press.
Any speech-to-text provider
Hands-free works with any supported speech-to-text provider, not just ones with built-in end-of-speech detection: Sophon detects the pause itself.
With no provider, your browser’s recognition ends each utterance at a pause, in Chrome or Edge.
Words as you say them
In the Dashboard, with Deepgram as the speech-to-text provider, your words appear on screen while you’re still speaking. Your browser’s recognition shows them live too.
Other providers transcribe once you stop, one round trip per utterance.
Speaking
A sentence at a time
With a text-to-speech provider configured, Sophon starts speaking a reply sentence by sentence while it is still writing the rest of it.
A one-sentence reply is spoken once it’s complete. Without a provider, your browser’s voice reads the whole reply at the end.
Progress, out loud
During long, multi-step tasks, Sophon can speak a short progress line after a quiet stretch, for example when it moves on to a new step. Long silences are less likely to feel like a dropped connection.
Needs a text-to-speech provider. A single long step can stay quiet for its whole run.
Formatting stays on screen
When a text-to-speech provider streams the reply, Sophon strips Markdown formatting before speaking and reads links as their text or site name.
Code and tables are best read in the chat.
While a text-to-speech provider speaks, the Voice page shows your words and the orb, not the reply. The full reply is written to the chat thread.
For whoever runs the Gateway
When voice breaks, it says where.
Nothing is preconfigured. With no speech provider, voice uses your browser’s speech recognition and voice (Chrome or Edge). Add speech providers for server-side transcription and provider voices.
Not set up, or failing
When voice can’t use a speech provider, the Voice page explains why in plain words: either speech isn’t set up and your browser’s speech is used, or the configured provider is failing, and it names that provider.
The notice appears only when no working provider of that kind is left; while another one works, voice simply uses it.
Checked and time-limited
Configured speech providers are health-checked automatically in the background, and a single transient failure doesn’t flag a provider as broken. Speech provider requests — transcription, synthesis, health checks and connecting — have time limits, so a provider that stops responding shows up as a named error instead of a stuck conversation.
A provider is flagged only after repeated failed checks, so its status can trail an outage, and it resets when the Gateway restarts.
One command to check
sophon voice status shows the health of every configured speech provider and whether hands-free is available. Operators can track time to first audio, per-provider text-to-speech latency and text-to-speech provider failures on Sophon’s authenticated Prometheus metrics endpoint.
Those metrics are recorded when a text-to-speech provider speaks the reply.
Speech-to-text
Text-to-speech
In the Dashboard, Deepgram shows your words as you speak; the others transcribe each utterance after you stop. Admins add speech providers, set their priority and test them, and provider credentials are kept in Sophon’s credential vault. Priority is set when a provider is added, and providers are configured per Gateway, not per tenant.
Where your audio goes
To the providers you chose
In the Dashboard, with a speech-to-text provider active, your mic audio travels over your signed-in connection to your Gateway, which passes it to that provider. Reply sentences go to your text-to-speech provider. Those vendors’ terms apply.
Or to your browser’s vendor
With no provider, your browser’s recognizer does the listening and needs an internet connection; the browser vendor decides where that audio goes (Chrome, for example, uses Google’s speech service). On Sophon Mobile, your phone’s own recognizer listens first.
What’s stored is the text
Your words and the reply are saved in the chat thread like typed messages. The Gateway doesn’t write your recorded audio to disk. Reply sentences already spoken by a text-to-speech provider are kept in an in-memory cache and replayed instead of being synthesized and billed again, until the Gateway restarts.
Where it actually is
What we have not shipped.
2.0 rebuilt the engine. Several things people reasonably expect from voice software are not part of it yet. Here is the list, before you rely on it.
There’s no talk-over detection. Speaking over Sophon doesn’t stop it; pressing the mic does, and that silences the reply, not the task behind it.
Voice turns use a smaller, voice-focused tool set than typed chat. Some jobs are better typed.
Approval and question prompts are read in your browser’s built-in voice (approvals on Sophon Mobile in your phone’s), not the provider voice you picked.
While a text-to-speech provider speaks, the Voice page doesn’t show the reply; the full text is in the chat thread. Speech is streamed a sentence at a time, never word by word, and only with a text-to-speech provider.
In the Dashboard, only Deepgram among speech-to-text providers shows your words while you speak; the others transcribe after you stop.
Sophon Mobile doesn’t read questions aloud and has no live transcription through the Gateway; it relies on your phone’s recognizer.
The CLI manages voice but can’t hold a voice conversation, and the desktop app has no push-to-talk shortcut.
Out of the box, Sophon Node takes no part in voice conversations: it can’t listen, has no wake word and doesn’t speak replies, and its voice routes stay off until an operator turns them on.
The Voice page and the Command Bridge start a new conversation with the default agent each time, with no agent picker. Ending voice in one tab ends it in every tab.
Without a provider, voice needs a browser with speech recognition, such as Chrome or Edge. Firefox has none.
Speech providers are set per Gateway, not per tenant, and can’t be reordered after they’re added. Smarter end-of-turn detection and plugin speech providers aren’t built yet.
Start with nothing
Three steps and a status check.
sophon voice status lists every configured speech provider with its health, and shows server and browser hands-free separately. The CLI can add, list, test and remove the same providers and edit the same personal and host listening settings; provider priority and a provider’s voice are set in Voice settings. Adding providers needs an admin.
Talk to it. Keep the gates.
Voice ships with Sophon on every tier and needs no speech provider to start: open the Voice page in Chrome or Edge, and add providers when you want server transcription and provider voices. The Personal tier is free for individual, non-commercial use.