Speech Providers
The six speech-to-text and six text-to-speech provider types, how priority and failover work, where credentials are kept, and what the browser does when nothing is configured.
Sophon ships with no speech provider configured. Voice still works in that state, using your browser's own recognition and voice. Providers are what you add when you want server-side transcription, consistent voices, and replies spoken while they are still being written.
Adding, testing and removing providers is Admin-only, and provider configuration is per Gateway, not per tenant: every user on the Gateway sees the same provider list.
Speech-to-text providers
Six types. Only one of them streams.
| Provider | Type string | How it transcribes | Ends an utterance itself |
|---|---|---|---|
| Deepgram | deepgram | True streaming over a WebSocket, with interim results while you speak | Yes |
| OpenAI | openai | Buffered — sends the whole recording once capture ends | No |
| Azure Speech | azure-speech | Buffered | No |
| Google Cloud Speech | google-speech | Buffered | No |
| ElevenLabs | elevenlabs | Buffered | No |
| Sophon Managed Speech | sophon-managed | Buffered | No |
Deepgram is the only provider that shows your words as you speak. The other five pay one round trip per utterance: nothing is sent until you stop talking, and the transcript arrives in one piece. In the Dashboard that means the preview card fills in after you finish rather than during.
The other five can still run hands-free. In Conversation mode, Sophon's own server-side silence detector supplies the end of the utterance for any provider that cannot detect it: it waits for at least 300 ms of real speech, then for the configured endpointing gap of silence. Deepgram is never wrapped that way — it keeps using its own detection.
The OpenAI provider transcribes with gpt-4o-mini-transcribe.
Text-to-speech providers
Six types, all called one sentence at a time.
| Provider | Type string | Voice list |
|---|---|---|
| ElevenLabs | elevenlabs | Fetched from your account |
| OpenAI | openai-tts | Six built-in voices (tts-1) |
| Google Cloud TTS | google-tts | Fetched per language |
| Azure Cognitive Services | azure-speech | Fetched per region |
| Deepgram | deepgram | Three built-in voices |
| Sophon Managed Speech | sophon-managed | Supplied by the managed service |
No provider streams audio within a sentence. Sophon splits the reply into sentences and synthesizes them one after another, so speech starts early because the sentences arrive early, not because any single sentence streams. See Spoken Replies for how that pipeline behaves.
Priority and failover
Each provider carries a priority, a number from 1 to 100 where lower wins. Sophon tries active providers in priority order and moves to the next one when a call fails — an error, a timeout, or a provider that returns empty audio all count as a failure. If every provider fails, Sophon sends a fallback notice and, unless you have turned it off, reads the reply with the browser voice instead.
You can also pin a default text-to-speech provider in your own settings. That provider is tried first, and the rest stay available as failover behind it.
Priority is set when a provider is added, and cannot be changed afterwards. There is no reorder control and no way to disable a provider in place. The only operations on an existing provider are test, change its voice (text-to-speech only), and remove. To change a provider's priority, remove it and add it again with the priority you want.
A provider can show as inactive in the capabilities response and the provider list, but nothing in the product sets that state — it only appears if stt.json or tts.json was hand-edited. What the health sweep sets is error, which is a different thing and is covered in Health & Troubleshooting.
Adding a provider
In the Dashboard, go to Settings → Voice (Voice Center). Speech-to-text providers are added from the Speech-to-text providers panel, text-to-speech providers from the TTS providers panel. Each dialog asks for a provider type, a display name, an API key or access token, a priority, and — for two types — an endpoint:
- Azure Speech asks for a region (
eastus,westeurope, and so on). It defaults toeastus. - Sophon Managed Speech asks for the base URL of the hosted speech service.
Every other type ignores the endpoint field.
From the CLI:
# Speech-to-text
sophon voice stt-providers add --name deepgram --type deepgram --api-key ...
sophon voice stt-providers list
sophon voice stt-providers test <providerId>
sophon voice stt-providers remove <providerId>
# Text-to-speech
sophon voice providers add --name elevenlabs --type elevenlabs --api-key ...
sophon voice providers list
sophon voice providers voices <providerId> [--language en-US]
sophon voice providers test <providerId>
sophon voice providers remove <providerId>The CLI adds at the default priority of 1 and has no flag to set one, and it cannot set a provider's voice. Both of those live in Voice Center. Everything else — adding, listing, testing and removing — works the same from either place. Omit --api-key and the CLI prompts for it without echoing it.
After adding, test the provider. A test makes one real call to the vendor and reports whether it answered. Text-to-speech providers also have a synthesize test, which is what the voice preview in Voice Center uses.
Where credentials are kept
Provider credentials go into Sophon's credential vault, encrypted at rest with the Gateway's vault key. They are never written into the provider config files.
~/.sophon/config/stt.json and ~/.sophon/config/tts.json hold non-secret configuration only: the provider id, display name, type, endpoint or region, voice id, priority, status, and the vault credential id the key is stored under. A Gateway upgraded from an older release migrates any key still embedded in those files into the vault the first time it loads them.
This matters for backups: copying stt.json and tts.json alone does not carry your keys. See Backup & Upgrade.
Sophon Managed Speech
sophon-managed points both speech-to-text and text-to-speech at one Sophon-hosted speech service. Configure the same base endpoint and the same scoped access token in both provider lists.
The service exposes three operations: a health check, a multipart transcription endpoint that returns { "text": ... }, and a synthesis endpoint that accepts text, voice, language, speed and output format and returns audio bytes or base64. Sophon treats any successful status on the health check as healthy.
Licensing, quotas, metering and token exchange belong to that external service, not to your Gateway. Sophon does not enforce or report them. Use short-lived tokens scoped to speech health, transcription and synthesis.
When no provider is configured
This is a supported configuration, not a broken one.
- Listening uses the browser's own speech recognition, which ends the utterance at a pause. Sophon's in-product hint names Chrome and Edge; the CLI's
voice statusoutput also lists Safari. Firefox has no speech recognition at all — the mic button is still shown and fails when pressed, and hands-free is not offered. - Replies are read by the browser's built-in voice after the whole turn finishes. There is no sentence-by-sentence speech, no spoken progress lines, no Markdown rewriting for the ear, and no voice metrics on this path.
- Hands-free still works, through the browser recognizer rather than the server.
- The Voice page shows an info-tone banner saying which side is not configured, with a link to Voice Center.
Browser recognition is not local processing. The browser vendor decides where that audio goes, and it needs an internet connection. See Privacy & Data.
Where to go next
- Voice Settings — every field in Voice Center, with its real range
- Set Up Voice — the shortest path from nothing to a working setup
- Health & Troubleshooting — health checks, timeouts, and what the banners mean
- Privacy & Data — where audio and text actually travel
- Credential Vault — how provider keys are stored
- CLI Commands — the full
sophon voicecommand surface
Set Up Voice
The shortest path from nothing to a working setup — the zero-provider start, turning on Conversation mode, adding and testing speech providers, verifying with sophon voice status, and the two preferences everyone should set.
Voice Settings
A field-by-field reference for Voice Center — General, Listening, Conversation mode and TTS providers — with the real ranges, the settings that only apply to Sophon Mobile, and the three that do nothing where you would expect.