Voice Settings
A field-by-field reference for Voice Center — General, Listening, Conversation mode and TTS providers — with the real ranges, the settings that only apply to Sophon Mobile, and the three that do nothing where you would expect.
Voice settings live in the Dashboard at Settings → Voice, which the product calls Voice Center. Sophon Mobile has its own copy under More → Voice Center. The same values are readable and writable from the CLI.
Almost everything here is your own preference, stored per user. Two things are not: the host listening defaults and the provider lists, both of which are Admin-only and apply to the whole Gateway.
General
| Field | Values | Default | What it does |
|---|---|---|---|
| Language | BCP-47 code | en-US | Used for speech recognition and for synthesis |
| Speech speed | 0.5×–2.0× | 1.0× | Applies to provider speech output |
| Speech input | OS / browser first · OS / browser only · Sophon provider only | OS / browser first | Which recognizer to use |
| Sophon STT fallback (Mobile app) | on / off | on | Mobile only — see below |
| Speech output | Configured provider first · Configured provider only · OS / browser only | Configured provider first | Which synthesizer to use |
| Default TTS provider | a provider, or Automatic by priority | Automatic | Tried first; the rest stay as failover |
| Native speech fallback | on / off | on | Let the browser or device voice speak when no provider can |
Language is yours, not the host's. Recognition and synthesis always use the language on your own preference. Turning off Native speech fallback, or setting Speech output to Configured provider only, also silences spoken approval and question prompts, because those are read by the browser or device voice rather than by a provider.
Listening
The Listening panel holds your personal speech-to-text overrides. Anything left on automatic uses the host default, and a Reset to host defaults button clears all four at once.
| Field | Range | Host default | What it does |
|---|---|---|---|
| Speech-to-text provider | a provider, or Automatic | Automatic (highest priority active) | Which provider transcribes your audio |
| Interim results | on / off | on | Stream partial transcripts while you speak |
| Endpointing | 100–3000 ms | 300 ms | Silence before your utterance counts as complete |
| Max utterance length | 5–300 s | 60 s | A longer recording is force-finalized |
Below it, admins see a Host defaults panel with the same four fields plus a Default language. Those values apply to every user who has not set a personal override.
Interim results only produce live text with Deepgram, the one streaming speech-to-text provider. The other five transcribe after you stop talking, so the setting has nothing to stream. The browser recognizer shows its own live words regardless of this setting.
Conversation mode
Conversation mode is hands-free: after each reply, the mic re-opens by itself. It is off by default and is a per-user preference.
| Field | Range | Default | What it does |
|---|---|---|---|
| Auto-resume listening | on / off | off | The master switch for Conversation mode |
| Silence timeout | 3–30 s | 30 s | Stop listening after this much silence on your end |
| Inactivity limit | 30 s–10 min | 2 min | Leave Conversation mode after this much continuous inactivity |
Both timers are only editable while Conversation mode is on. The server clamps more loosely than the controls do — silence to 3–300 s, inactivity to 10–1800 s — so a value set through the API can sit outside the range the slider offers.
The two limits stop different things. The silence timeout mainly bites with server speech-to-text, where the mic stays open through a pause; with browser recognition the recognizer ends itself within seconds of any pause and hands-free just re-arms, so the inactivity limit is what eventually turns it off. Both stop with a message on screen rather than silently. See Conversations.
TTS providers
The provider list, with a status pill and Test, Browse voices and Remove on each row, plus Add provider for admins. Browsing voices fetches the provider's voice list and lets you preview one before setting it. Priority is shown on each row but is fixed once a provider is added.
Speech-to-text providers have their own list in the same page, under Speech-to-text providers.
See Speech Providers for what each type needs and how failover works.
Settings that only apply to Sophon Mobile
Two fields on this page do nothing in the Dashboard:
- Sophon STT fallback (Mobile app). Mobile listens with the phone's own recognizer first. When that fails, this decides whether Mobile uploads the recording to your Gateway for transcription instead. The Dashboard never reads it. The toggle is only enabled when Speech input is OS / browser first.
- Speech input, partly. Mobile honours all three values distinctly. The Dashboard only special-cases OS / browser only; OS / browser first and Sophon provider only behave identically there — both mean "use an active server provider if one exists, otherwise the browser".
Settings that do nothing where you would expect
Three fields are labelled in the product to say so, and the behaviour behind them is unchanged in 2.0. They are on the roadmap, not fixed.
- Host defaults → Default language is labelled "Not used yet". No code path reads it. Recognition and synthesis always use each person's own language from Voice Center.
- Sophon STT fallback is read only by Sophon Mobile, as above.
- Speech input values OS / browser first and Sophon provider only are indistinguishable in the Dashboard, as above.
From the CLI
sophon voice settings show
sophon voice settings set --conversation-mode true --silence-timeout 15
sophon voice settings listening --endpointing-ms 500
sophon voice settings listening --host-defaults --max-utterance 90settings set is read-modify-write: it reads your current preferences, applies only the flags you passed, and writes the whole object back. Fields it does not manage — including your personal listening overrides — are preserved. Earlier releases dropped them.
| Setting | CLI flag | Accepted values |
|---|---|---|
| Language | --language | BCP-47, for example en-US |
| Speech speed | --speed | 0.5–2.0 |
| Speech input | --input-mode | native-first · native-only · server-only |
| Speech output | --output-mode | provider-first · provider-only · native-only |
| Default TTS provider | --default-tts-provider | a provider id, or automatic |
| Native speech fallback | --fallback-native | true · false |
| Sophon STT fallback | --fallback-server-stt | true · false |
| Auto-resume listening | --conversation-mode | true · false |
| Silence timeout | --silence-timeout | 3–30 (seconds) |
| Inactivity limit | --inactivity-limit | 30–600 (seconds) |
The inactivity flag is --inactivity-limit. Some older product notes call it --max-inactivity; that flag does not exist.
sophon voice settings listening edits your personal overrides. Add --host-defaults to edit the Gateway-wide defaults instead, which requires Admin or Owner. Its flags are --language, --default-provider, --interim-results, --endpointing-ms (100–3000) and --max-utterance (5–300).
Values outside a range are rejected by the CLI before anything is sent, with the range in the error.
Where to go next
- Speech Providers — the six speech-to-text and six text-to-speech types, and how priority works
- Conversations — push-to-talk, hands-free, interrupting and reconnects
- Spoken Replies — what the output settings actually change
- Set Up Voice — a working setup from scratch
- CLI Commands — the full
sophon voicecommand surface
Speech Providers
The six speech-to-text and six text-to-speech provider types, how priority and failover work, where credentials are kept, and what the browser does when nothing is configured.
Health & Troubleshooting
What each availability banner means, how the provider health sweep works and why status can trail an outage, the per-call time limits, sophon voice status and the capabilities endpoint, the three voice metrics, the Sophon:Voice configuration keys, and a symptom-by-symptom table.