Sophon 2.0 is here
Rebuilt in 2.0

Sophon Voice

Say it out loud. The same gates apply.

Voice is another way into the runtime you already run, not a second assistant. A spoken request runs as a regular task on your chat, with a smaller, voice-focused tool set. It is queued and approval-gated the same way as a typed message and appears live in the chat thread. When an action needs your approval, the Voice page reads it out, and High- or Critical-risk actions need you to say “yes, approve”.

On every tier. Nothing is preconfigured: with no speech provider, voice uses your browser’s own speech recognition and voice (Chrome or Edge).

Animation of the Sophon Voice page: a spoken request to delete a schedule raises a High-risk approval; Sophon reads it out; a bare “yes” is asked again; “yes, approve” approves it; the reply is spoken and the microphone reopens.

Recreated from the Dashboard’s Voice page with conversation mode on, Deepgram for speech-to-text and a text-to-speech provider. Lines under the frame caption what is said aloud, plus a few notes; none of them is on the page. Approval prompts are read in the browser’s own voice, and a reply spoken by a provider is heard, not printed: its text lands in the chat thread.

A voice turn

A spoken request is an ordinary task.

Voice isn’t a second assistant. Sophon runs what you say as a regular agent task on your chat, with a smaller, voice-focused tool set, stores it in the thread, gates it with the same approvals as a typed message, and then speaks the answer.

01Transcribeda pause or a press ends it; your speech-to-text provider or your browser turns it into text
02Queuedit joins that chat’s queue like a typed message and, by default, waits behind anything already running
03Gateda tool that needs approval stops and asks before it runs; on the Voice page, out loud
04Spoken and keptthe answer is spoken, and your words and the reply are saved in the chat thread

Same queue, same thread

A voice request runs as a regular agent task on your chat session. It is queued and approval-gated the same way as a typed message and appears live in the chat thread.

Voice turns use a smaller, voice-focused tool set, so some jobs are better typed.

Failures are said, not swallowed

If a voice request fails, Sophon tells you by voice or on screen instead of going quiet.

A turn you interrupted is the exception: it stays silent by design, and its text still reaches the chat.

Interrupting silences the reply, not the work

On the Voice page, press the mic while Sophon is thinking or speaking to silence it and talk. The interrupted reply isn’t spoken later, but it still appears in the chat.

Interrupting stops the speech, not the task: work already under way keeps running, and by default your next request waits behind it. There’s no talk-over detection; interrupting is always a press.

Approvals by voice

Risk decides what counts as yes.

Which actions ask at all is set by your approval policy, the same one typed requests answer to. Voice changes how you answer: on the Voice page and the Command Bridge, Sophon closes the mic, reads out the kind of action and its risk, and sets a higher bar for the riskier ones.

Low and Medium

A plain “yes” approves. A plain “no” rejects.

High and Critical

Only “yes, approve” approves. A bare “yes” isn’t enough, and Sophon asks again.

A yes with a but

A “yes” that also says “no”, “wait” or “don’t” counts as unclear, so Sophon asks again.

Two re-asks, then a tap

Unclear answers are asked again at most twice; after that the mic closes and the prompt waits for a tap or a mic press. Sophon never guesses.

The approval’s own timer keeps running, and an approval nobody answers is rejected when it expires.

What you can say

You sayLow or MediumHigh or Critical
“Yes.”ApprovedAsked again
“Yes, approve.”ApprovedApproved
“Yes… wait.”Asked againAsked again
“No.”RejectedRejected

Asked again means Sophon repeats what it needs, at most twice, then waits for a tap or a mic press.

A spoken approval names the kind of action and its risk, not its exact details. When the details matter, read the card before you say “yes, approve”. Prompts are read in your browser’s or phone’s built-in voice, not your text-to-speech provider’s; turn that voice off and you get the card and the mic instead.

When the agent asks

Numbered choices. Nothing guessed.

When an agent asks a question with options, the Voice page and the Command Bridge read it out in your browser’s voice with numbered choices. Answer by saying the number (one to four) or the option’s name, or tap it.

One question at a time

Multi-part questions are asked one at a time, and your answers are sent together at the end. If you can pick more than one, Sophon says so.

Say it the way you’d say it

Option names with symbols are read out and matched as spoken words, so you can say “C plus plus” or “C sharp”.

“Don’t” means don’t

If you negate an option (“don’t send”), Sophon won’t pick it. It asks again, or passes your own words to the agent when the question accepts a free-text answer.

Short answers for numbers

A spoken number counts in a short answer. Options past the fourth are picked by name or tap.

If your speech recognizer writes the symbol instead of the words, Sophon asks again; tap the option. Question cards come from plans and other tasks on the chat: a voice turn is set up to ask its own follow-ups out loud, in conversation, rather than as a card. Sophon Mobile shows questions on screen and doesn’t read them out.

A Sophon Voice question card: question one of two, “Which calendar should the review go on?”, with the options Work, Personal and Board. It is answered out loud or by tapping an option.

One decision, recorded once

Answer anywhere. The rest close.

An approval is one record on your Gateway, not a dialog on one screen. The first answer decides it.

Where a prompt is asked out loud

Voice pageApprovals and questions are read out in your browser’s voice, with the card and its buttons on screen. Answer by voice or tap.
Command BridgePrompts are read out in your browser’s voice. Pressing the mic while one is waiting opens the mic so you can answer it out loud. A separate Stop button turns voice off without answering, and the prompt stays on its chat card.
Chat voice ribbonNever reads prompts aloud. Answer on the chat’s own cards.
Sophon MobileReads approvals and applies the same “yes, approve” rule, for its own voice conversation only. Questions are answered on screen.

Voice only speaks prompts that belong to its own conversation, so an approval from an unrelated session never interrupts you.

Closed everywhere

Answer an approval in one place, whether by voice, on a chat card, in another tab or in the mobile app, and any other open chat card or voice prompt for it closes on its own.

The Dashboard’s Approvals page catches up on its next refresh.

Recorded exactly once

Each approval is recorded exactly once, so when two devices answer the same card only one answer counts, and an approval that timed out or was cancelled can never later become an approval.

Says how it ended

Approval cards on the Dashboard and Mobile now say how an approval ended: Approved, Rejected, Cancelled or Timed out.

If a plan pauses for your approval and finishes later, its result is spoken to a voice session still open on that chat. If you’re mid-conversation at that moment, it appears in the chat thread instead.

What changed in 2.0

Same orb. New engine.

If you used voice in 1.x, the Voice page will look familiar: the look barely changed. The work went underneath: one voice engine behind the Voice page, the chat ribbon and the Command Bridge, and a voice session that belongs to the conversation, not to a browser connection.

01One enginelistening, playback and interrupting follow the same rules on the Voice page, the chat ribbon and the Command Bridge
02A session that holdsa dropped connection rejoins the reply in progress; audio sent while you were offline isn’t replayed
03Every tab hears itevery tab attached to the voice conversation hears the reply
04A clean exitleave voice and the microphone is released, every time

Drop the connection, keep the conversation

If your connection drops while Sophon is answering, voice picks back up when you reconnect: you hear the rest of the reply from that point, and the full reply is in the chat.

Audio sent while you were offline isn’t replayed.

Every tab on the conversation

Every open tab attached to the same voice conversation hears the reply.

Ending voice in one tab ends it in all of them, and each visit to the Voice page opens a new conversation.

Nothing looks fine when it isn’t

Every voice error now shows in a strip across the top of the Voice page, and the microphone is released every time you leave voice, so there’s no lingering recording indicator.

With server speech-to-text, each transcript is matched to what you said when it was said, so a slow transcription of an earlier sentence is much less likely to be taken as your answer or as a new request.

Where you can talk to it

Voice pageFull screen: the orb, a preview of what you said, and approval and question cards you answer out loud or with a tap.Each visit starts a new conversation with the default agent.
Chat voice ribbonTalks on the chat you have open and uses your speech-to-text provider when one is active. Switching chats ends voice on the old chat and restarts it, listening, on the new one, so replies aren’t spoken into the wrong conversation.It never reads prompts aloud.
Command BridgeThe Dashboard’s full-screen fleet view has a hands-free voice bar that listens once you arm it. When it stops — after silence, inactivity or an error — it switches itself off and shows why.Hands-free only, with no push-to-talk; each arm starts a new conversation with the default agent.
Sophon MobileThe Voice tab uses your phone’s own speech recognition by default, speaks replies in your provider’s voice or the device voice, and Conversation mode listens again after each reply.Questions are answered on screen, not read aloud.

The CLI manages voice — providers, settings and sophon voice status — but has no voice conversation of its own.

Listening and speaking

Hands-free with any provider. Or none.

Conversation mode works end to end, your browser’s own speech is enough to start, and with a text-to-speech provider Sophon starts talking before the reply is finished.

Listening

Conversation mode

Turn on Conversation mode in Voice settings and your first mic press starts a hands-free conversation: speak, pause, hear the answer, and the mic reopens only after the reply has finished playing. Silence and inactivity limits close the mic when you walk away.

It isn’t a wake word or always-on listening: it starts with your press.

Any speech-to-text provider

Hands-free works with any supported speech-to-text provider, not just ones with built-in end-of-speech detection: Sophon detects the pause itself.

With no provider, your browser’s recognition ends each utterance at a pause, in Chrome or Edge.

Words as you say them

In the Dashboard, with Deepgram as the speech-to-text provider, your words appear on screen while you’re still speaking. Your browser’s recognition shows them live too.

Other providers transcribe once you stop, one round trip per utterance.

Speaking

Needs a text-to-speech provider

A sentence at a time

With a text-to-speech provider configured, Sophon starts speaking a reply sentence by sentence while it is still writing the rest of it.

A one-sentence reply is spoken once it’s complete. Without a provider, your browser’s voice reads the whole reply at the end.

Progress, out loud

During long, multi-step tasks, Sophon can speak a short progress line after a quiet stretch, for example when it moves on to a new step. Long silences are less likely to feel like a dropped connection.

Needs a text-to-speech provider. A single long step can stay quiet for its whole run.

Formatting stays on screen

When a text-to-speech provider streams the reply, Sophon strips Markdown formatting before speaking and reads links as their text or site name.

Code and tables are best read in the chat.

While a text-to-speech provider speaks, the Voice page shows your words and the orb, not the reply. The full reply is written to the chat thread.

For whoever runs the Gateway

When voice breaks, it says where.

Nothing is preconfigured. With no speech provider, voice uses your browser’s speech recognition and voice (Chrome or Edge). Add speech providers for server-side transcription and provider voices.

Not set up, or failing

When voice can’t use a speech provider, the Voice page explains why in plain words: either speech isn’t set up and your browser’s speech is used, or the configured provider is failing, and it names that provider.

The notice appears only when no working provider of that kind is left; while another one works, voice simply uses it.

Checked and time-limited

Configured speech providers are health-checked automatically in the background, and a single transient failure doesn’t flag a provider as broken. Speech provider requests — transcription, synthesis, health checks and connecting — have time limits, so a provider that stops responding shows up as a named error instead of a stuck conversation.

A provider is flagged only after repeated failed checks, so its status can trail an outage, and it resets when the Gateway restarts.

One command to check

sophon voice status shows the health of every configured speech provider and whether hands-free is available. Operators can track time to first audio, per-provider text-to-speech latency and text-to-speech provider failures on Sophon’s authenticated Prometheus metrics endpoint.

Those metrics are recorded when a text-to-speech provider speaks the reply.

Speech-to-text

DeepgramOpenAIGoogle Cloud SpeechAzure SpeechElevenLabsSophon Managed Speech

Text-to-speech

ElevenLabsOpenAIGoogle CloudAzureDeepgramSophon Managed Speech

In the Dashboard, Deepgram shows your words as you speak; the others transcribe each utterance after you stop. Admins add speech providers, set their priority and test them, and provider credentials are kept in Sophon’s credential vault. Priority is set when a provider is added, and providers are configured per Gateway, not per tenant.

Where your audio goes

Not on-device

To the providers you chose

In the Dashboard, with a speech-to-text provider active, your mic audio travels over your signed-in connection to your Gateway, which passes it to that provider. Reply sentences go to your text-to-speech provider. Those vendors’ terms apply.

Or to your browser’s vendor

With no provider, your browser’s recognizer does the listening and needs an internet connection; the browser vendor decides where that audio goes (Chrome, for example, uses Google’s speech service). On Sophon Mobile, your phone’s own recognizer listens first.

What’s stored is the text

Your words and the reply are saved in the chat thread like typed messages. The Gateway doesn’t write your recorded audio to disk. Reply sentences already spoken by a text-to-speech provider are kept in an in-memory cache and replayed instead of being synthesized and billed again, until the Gateway restarts.

Where it actually is

What we have not shipped.

2.0 rebuilt the engine. Several things people reasonably expect from voice software are not part of it yet. Here is the list, before you rely on it.

There’s no talk-over detection. Speaking over Sophon doesn’t stop it; pressing the mic does, and that silences the reply, not the task behind it.

Voice turns use a smaller, voice-focused tool set than typed chat. Some jobs are better typed.

Approval and question prompts are read in your browser’s built-in voice (approvals on Sophon Mobile in your phone’s), not the provider voice you picked.

While a text-to-speech provider speaks, the Voice page doesn’t show the reply; the full text is in the chat thread. Speech is streamed a sentence at a time, never word by word, and only with a text-to-speech provider.

In the Dashboard, only Deepgram among speech-to-text providers shows your words while you speak; the others transcribe after you stop.

Sophon Mobile doesn’t read questions aloud and has no live transcription through the Gateway; it relies on your phone’s recognizer.

The CLI manages voice but can’t hold a voice conversation, and the desktop app has no push-to-talk shortcut.

Out of the box, Sophon Node takes no part in voice conversations: it can’t listen, has no wake word and doesn’t speak replies, and its voice routes stay off until an operator turns them on.

The Voice page and the Command Bridge start a new conversation with the default agent each time, with no agent picker. Ending voice in one tab ends it in every tab.

Without a provider, voice needs a browser with speech recognition, such as Chrome or Edge. Firefox has none.

Speech providers are set per Gateway, not per tenant, and can’t be reordered after they’re added. Smarter end-of-turn detection and plugin speech providers aren’t built yet.

Start with nothing

Three steps and a status check.

01Open VoiceIn the Dashboard, open Voice. In Chrome or Edge it works straight away with your browser’s speech.
02Go hands-freeSettings → Voice (Voice Center) → Conversation mode. Your first mic press then starts a hands-free conversation.
03Add providers when you want themAn admin adds speech-to-text and text-to-speech providers in Voice settings (where their priority is set) or from the CLI, and tests them. Speech that starts before the reply is finished needs a text-to-speech provider.
$sophon voice status$sophon voice settings set --conversation-mode true$sophon voice stt-providers add --name deepgram --type deepgram --api-key …$sophon voice providers add --name elevenlabs --type elevenlabs --api-key …

sophon voice status lists every configured speech provider with its health, and shows server and browser hands-free separately. The CLI can add, list, test and remove the same providers and edit the same personal and host listening settings; provider priority and a provider’s voice are set in Voice settings. Adding providers needs an admin.

Talk to it. Keep the gates.

Voice ships with Sophon on every tier and needs no speech provider to start: open the Voice page in Chrome or Edge, and add providers when you want server transcription and provider voices. The Personal tier is free for individual, non-commercial use.