Sophon 2.0 is here
Sophon Docs
Sophon Voice

Set Up Voice

The shortest path from nothing to a working setup — the zero-provider start, turning on Conversation mode, adding and testing speech providers, verifying with sophon voice status, and the two preferences everyone should set.

There is nothing to install and nothing to switch on. Voice ships inside Sophon on every tier, and with no speech provider configured at all it still listens and still speaks. Everything below is optional, in the order most people want it.

Step 1 — open Voice and talk

Open Voice in the Dashboard nav. In Chrome or Edge it works immediately, using the browser's own speech recognition and the browser's own voice.

You will see an info-tone banner at the top of the page, one line per side that is not configured:

SideBanner titleWhat it says
Speech-to-textSpeech-to-text isn't configured"Server speech recognition is not set up, so voice is using your browser's built-in recognition."
Text-to-speechText-to-speech isn't configured"Replies use your browser's built-in voice."

The banner carries an Open Voice Center link. It is informational, not an error: this is the intended zero-configuration path, and the page is fully usable in it. A warning-tone banner is a different thing — it means a configured provider is failing, and it names that provider. See Health & Troubleshooting.

Firefox has no speech recognition, so the no-provider path cannot listen there. The mic button still appears and fails when pressed, and hands-free is not offered — you get "This browser has no speech recognition. Use Chrome or Edge, or add a speech-to-text provider in Voice Center." Use Chrome or Edge, or add a speech-to-text provider.

Step 2 — turn on Conversation mode

Hands-free is a per-user preference and it is off by default. Turn it on in Settings → Voice (Voice Center) under Conversation mode, in Mobile's Voice Center, or from the CLI.

FieldRangeDefaultWhat it does
Auto-resume listeningon / offoffThe master switch. After each reply the mic re-opens by itself
Silence timeout3–30 s30 sStop listening after this much silence on your end
Inactivity limit30 s – 10 min2 minLeave Conversation mode after this much continuous inactivity

Both timers are only editable while Auto-resume listening is on.

sophon voice settings set --conversation-mode true --silence-timeout 15 --inactivity-limit 120

The flag is --inactivity-limit; --max-inactivity does not exist. settings set reads your current preferences, applies only the flags you passed and writes the whole object back, so your personal listening overrides survive.

Turning the preference on does not start listening. The hands-free toggle appears on the Voice page once the preference is on, and your first mic press arms hands-free and starts listening in one action.

Step 3 — add a speech-to-text provider

Optional. It gives you server-side transcription instead of the browser's, and it is what makes server-side hands-free available — the row sophon voice status reports as Hands-free · server STT. Hands-free works with any of the six provider types: Sophon supplies the end of an utterance itself for the five that cannot detect it.

In the Dashboard: Settings → Voice → the Speech-to-text providers panel → Add provider. The dialog asks for a provider type, a display name, an API key or access token, and a priority (1–100, lower wins). Two types want one more field: Azure Speech asks for a region (default eastus), and Sophon Managed Speech asks for the base URL of the hosted speech service.

From the CLI:

sophon voice stt-providers add --name deepgram --type deepgram --api-key ...
sophon voice stt-providers list
sophon voice stt-providers test <providerId>

Omit --api-key and the CLI prompts for it without echoing it.

Then test it. A test makes one real call to the vendor and reports whether it answered.

Priority is fixed at the moment a provider is added. There is no reorder control and no way to disable a provider in place — to change a priority, remove the provider and add it again. The CLI adds at the default priority of 1 and has no flag to set one, so if priority matters, add the provider in Voice Center. Adding, testing and removing providers is Admin-only, and the provider list is per Gateway, not per tenant.

Step 4 — add a text-to-speech provider and pick a voice

This is the one that changes how replies sound. With an active text-to-speech provider, the reply is spoken sentence by sentence while it is still being written, long multi-step turns can get a short spoken progress line, and Markdown is rewritten for the ear. Without one, the browser voice reads the finished reply once, at the end.

In the Dashboard: Settings → Voice → the TTS providers panel → Add provider. The same fields as above.

From the CLI:

sophon voice providers add --name elevenlabs --type elevenlabs --api-key ...
sophon voice providers list
sophon voice providers voices <providerId> [--language en-US]
sophon voice providers test <providerId>

Each row in the TTS providers panel has Test, Browse voices and Remove. Browse voices fetches that provider's voice list, lets you search it, and gives every voice a Preview button that synthesizes a sample through the provider itself. Picking a voice saves it on the provider. The CLI can list voices but cannot set one — that lives in Voice Center.

Finally, in General, set Default TTS provider. It is Automatic by priority out of the box; pinning a provider makes it the one tried first, and the others stay available as failover behind it.

Step 5 — verify

sophon voice status

It prints one table, Voice Status, with the columns Kind, Name and Status:

  • one row per configured speech-to-text provider, with its health (active, inactive or error)
  • one row per configured text-to-speech provider, the same way
  • (none configured) with the status browser fallback for a side with nothing registered
  • Hands-free · server STT — either available, or not available — no active speech-to-text provider
  • Hands-free · browser (Dashboard /voice) — reported separately, because the browser path works with no provider at all

The two hands-free rows are deliberately separate: a Gateway with no speech-to-text provider still has working hands-free on the Voice page in a browser that has speech recognition.

The two preferences worth setting

In Settings → Voice → General:

  • Language. A BCP-47 code, en-US by default. Recognition and synthesis both use your own language preference. The Admin-only Host defaults → Default language field is labelled "Not used yet" in the product and no code path reads it, so setting it changes nothing.
  • Speech speed. 0.5× to 2.0×, 1.0× by default. It applies to provider speech output.
sophon voice settings set --language en-GB --speed 1.15

Values outside a range are rejected by the CLI before anything is sent, with the range in the error.

On Sophon Mobile

Mobile has its own Voice Center at More → Voice Center. It covers the same personal preferences — language, speech speed, speech routing, Conversation Mode with its silence and inactivity limits, and which text-to-speech provider is your default — and shows the active provider with its status.

Two differences worth knowing:

  • Providers are not added on the phone. Add, test and remove them in the Dashboard or the CLI.
  • The Listening group edits the host-wide defaults, so it needs Admin or Owner. Everything above it is yours alone.

Mobile listens with the phone's own recognizer by default. The Sophon STT fallback toggle decides whether a failed or unavailable phone recognizer falls back to uploading the recording to your Gateway for transcription after you stop speaking. See Privacy & Data.

Where to go next