Sophon 2.0 is here
Sophon Docs
Sophon Voice

Voice Settings

A field-by-field reference for Voice Center — General, Listening, Conversation mode and TTS providers — with the real ranges, the settings that only apply to Sophon Mobile, and the three that do nothing where you would expect.

Voice settings live in the Dashboard at Settings → Voice, which the product calls Voice Center. Sophon Mobile has its own copy under More → Voice Center. The same values are readable and writable from the CLI.

Almost everything here is your own preference, stored per user. Two things are not: the host listening defaults and the provider lists, both of which are Admin-only and apply to the whole Gateway.

General

FieldValuesDefaultWhat it does
LanguageBCP-47 codeen-USUsed for speech recognition and for synthesis
Speech speed0.5×–2.0×1.0×Applies to provider speech output
Speech inputOS / browser first · OS / browser only · Sophon provider onlyOS / browser firstWhich recognizer to use
Sophon STT fallback (Mobile app)on / offonMobile only — see below
Speech outputConfigured provider first · Configured provider only · OS / browser onlyConfigured provider firstWhich synthesizer to use
Default TTS providera provider, or Automatic by priorityAutomaticTried first; the rest stay as failover
Native speech fallbackon / offonLet the browser or device voice speak when no provider can

Language is yours, not the host's. Recognition and synthesis always use the language on your own preference. Turning off Native speech fallback, or setting Speech output to Configured provider only, also silences spoken approval and question prompts, because those are read by the browser or device voice rather than by a provider.

Listening

The Listening panel holds your personal speech-to-text overrides. Anything left on automatic uses the host default, and a Reset to host defaults button clears all four at once.

FieldRangeHost defaultWhat it does
Speech-to-text providera provider, or AutomaticAutomatic (highest priority active)Which provider transcribes your audio
Interim resultson / offonStream partial transcripts while you speak
Endpointing100–3000 ms300 msSilence before your utterance counts as complete
Max utterance length5–300 s60 sA longer recording is force-finalized

Below it, admins see a Host defaults panel with the same four fields plus a Default language. Those values apply to every user who has not set a personal override.

Interim results only produce live text with Deepgram, the one streaming speech-to-text provider. The other five transcribe after you stop talking, so the setting has nothing to stream. The browser recognizer shows its own live words regardless of this setting.

Conversation mode

Conversation mode is hands-free: after each reply, the mic re-opens by itself. It is off by default and is a per-user preference.

FieldRangeDefaultWhat it does
Auto-resume listeningon / offoffThe master switch for Conversation mode
Silence timeout3–30 s30 sStop listening after this much silence on your end
Inactivity limit30 s–10 min2 minLeave Conversation mode after this much continuous inactivity

Both timers are only editable while Conversation mode is on. The server clamps more loosely than the controls do — silence to 3–300 s, inactivity to 10–1800 s — so a value set through the API can sit outside the range the slider offers.

The two limits stop different things. The silence timeout mainly bites with server speech-to-text, where the mic stays open through a pause; with browser recognition the recognizer ends itself within seconds of any pause and hands-free just re-arms, so the inactivity limit is what eventually turns it off. Both stop with a message on screen rather than silently. See Conversations.

TTS providers

The provider list, with a status pill and Test, Browse voices and Remove on each row, plus Add provider for admins. Browsing voices fetches the provider's voice list and lets you preview one before setting it. Priority is shown on each row but is fixed once a provider is added.

Speech-to-text providers have their own list in the same page, under Speech-to-text providers.

See Speech Providers for what each type needs and how failover works.

Settings that only apply to Sophon Mobile

Two fields on this page do nothing in the Dashboard:

  • Sophon STT fallback (Mobile app). Mobile listens with the phone's own recognizer first. When that fails, this decides whether Mobile uploads the recording to your Gateway for transcription instead. The Dashboard never reads it. The toggle is only enabled when Speech input is OS / browser first.
  • Speech input, partly. Mobile honours all three values distinctly. The Dashboard only special-cases OS / browser only; OS / browser first and Sophon provider only behave identically there — both mean "use an active server provider if one exists, otherwise the browser".

Settings that do nothing where you would expect

Three fields are labelled in the product to say so, and the behaviour behind them is unchanged in 2.0. They are on the roadmap, not fixed.

  • Host defaults → Default language is labelled "Not used yet". No code path reads it. Recognition and synthesis always use each person's own language from Voice Center.
  • Sophon STT fallback is read only by Sophon Mobile, as above.
  • Speech input values OS / browser first and Sophon provider only are indistinguishable in the Dashboard, as above.

From the CLI

sophon voice settings show
sophon voice settings set --conversation-mode true --silence-timeout 15
sophon voice settings listening --endpointing-ms 500
sophon voice settings listening --host-defaults --max-utterance 90

settings set is read-modify-write: it reads your current preferences, applies only the flags you passed, and writes the whole object back. Fields it does not manage — including your personal listening overrides — are preserved. Earlier releases dropped them.

SettingCLI flagAccepted values
Language--languageBCP-47, for example en-US
Speech speed--speed0.5–2.0
Speech input--input-modenative-first · native-only · server-only
Speech output--output-modeprovider-first · provider-only · native-only
Default TTS provider--default-tts-providera provider id, or automatic
Native speech fallback--fallback-nativetrue · false
Sophon STT fallback--fallback-server-stttrue · false
Auto-resume listening--conversation-modetrue · false
Silence timeout--silence-timeout3–30 (seconds)
Inactivity limit--inactivity-limit30–600 (seconds)

The inactivity flag is --inactivity-limit. Some older product notes call it --max-inactivity; that flag does not exist.

sophon voice settings listening edits your personal overrides. Add --host-defaults to edit the Gateway-wide defaults instead, which requires Admin or Owner. Its flags are --language, --default-provider, --interim-results, --endpointing-ms (100–3000) and --max-utterance (5–300).

Values outside a range are rejected by the CLI before anything is sent, with the range in the error.

Where to go next