Health & Troubleshooting
What each availability banner means, how the provider health sweep works and why status can trail an outage, the per-call time limits, sophon voice status and the capabilities endpoint, the three voice metrics, the Sophon:Voice configuration keys, and a symptom-by-symptom table.
Voice is built to say where it broke. A provider that is not set up reads differently from one that is failing, every provider call is time-limited so a hung vendor surfaces as a named error, and one CLI command prints the whole picture.
This page is the operator's view: what the banners mean, what is checked and how often, and what to do when something is wrong.
The availability banner
The Voice page reads GET /api/voice/capabilities and shows at most one message per side.
| What you see | Tone | Meaning |
|---|---|---|
| No banner | — | At least one active provider exists on that side |
| "Speech-to-text isn't configured" | Info | Nothing registered on that side, or everything registered is turned off. "Server speech recognition is not set up, so voice is using your browser's built-in recognition." |
| "Text-to-speech isn't configured" | Info | The same, for replies. "Replies use your browser's built-in voice." |
| "A speech provider is failing" | Warning | At least one registered provider on that side failed its health checks and none on that side is active. The description names the provider: "Deepgram is failing, so voice is using your browser's built-in speech recognition instead." |
The banner only reports whether a side has anything usable at all. A second, failing provider sitting next to a healthy active one shows no banner — voice simply uses the healthy one. To see the health of every provider, including the ones the banner hides, run sophon voice status or read the capabilities endpoint.
Both banners link to Open Voice Center. The info tone is the expected state on a fresh install, not a fault.
The health sweep
A background sweep calls every registered provider's own health check and updates its status in the Gateway's memory.
| Behaviour | Value |
|---|---|
| How often every registered provider is re-checked | 15 minutes (Sophon:Voice:HealthMonitor:IntervalMinutes) |
| Consecutive failures before a provider is marked failing | 2 |
| Successes needed to clear it | 1 — recovery is immediate |
| Providers that are explicitly turned off | Skipped |
| Where the status lives | In memory only |
| What a Gateway restart does | Resets every provider to whatever stt.json / tts.json last held, until the next sweep |
Two consequences follow, and both are deliberate:
- Status trails an outage. Two consecutive failures at a 15-minute interval means a provider that dies right after a sweep can take roughly 30 minutes to show as failing. Degradation requires confirmation so that one network blip does not make voice look broken; recovery does not, so a provider that comes back is trusted on the next sweep.
- Status does not survive a restart. After a Gateway restart a failing provider reads as active until it is checked again.
LastHealthCheckon each provider tells you when it was last looked at.
The sweep never invents an unhealthy entry for something that was never registered, so "not configured" can never render as "failing".
Per-call time limits
Every speech-provider request is bounded, and a timeout is reported as a named error that says which provider and which operation ran out of time — not as a stuck conversation.
| Call | Limit |
|---|---|
| Health check, and listing voices or languages | 10 s |
| Speech synthesis (one sentence) | 30 s |
| Transcription of a captured recording | 45 s |
| Connecting — a streaming handshake, or the first response of a streamed call | 15 s |
An open streaming session is deliberately not bounded. The 15-second limit covers establishing the channel; what flows afterwards has no aggregate cap, because a legitimately long reply would otherwise be cut off mid-sentence.
Separately, once audio capture ends the Gateway gives the provider a bounded window to produce the transcript — max(5 seconds, seconds of audio ÷ 4) for a streaming provider, never less than 55 seconds for a buffered one. See Conversations.
sophon voice status
sophon voice statusOne table — Kind, Name, Status — built from the capabilities endpoint:
| Kind | Name | Status |
|---|---|---|
| Speech-to-text | each configured provider | active, inactive or error |
| Speech-to-text | (none configured) | browser fallback |
| Text-to-speech | each configured provider | active, inactive or error |
| Text-to-speech | (none configured) | browser fallback |
| Hands-free | server STT | available, or not available — no active speech-to-text provider |
| Hands-free | browser (Dashboard /voice) | works without a provider in Chrome, Edge or Safari |
The two hands-free rows are reported separately on purpose. Collapsing them would read as "hands-free is broken" on a Gateway with no speech-to-text provider, even though the Voice page's browser path works there.
The capabilities endpoint
GET /api/voice/capabilities is what the banner, the Dashboard and the CLI all read. It is the right thing to poll from a script or a support runbook.
| Field | What it means |
|---|---|
serverTranscriptionAvailable | At least one speech-to-text provider is active |
providerSpeechAvailable | At least one text-to-speech provider is active |
handsFreeSupported | Server-side hands-free is available — identical to serverTranscriptionAvailable, because any active provider can be endpointed server-side. It says nothing about the browser path |
nativeInputSupported / nativeOutputSupported | Always true; the browser paths are always offered |
singleVendorOptions | Vendors that are active on both sides |
sttProviders / ttsProviders | Every registered provider, including inactive and failing ones |
Each provider entry carries id, name, providerType, vendor, supportsStreaming, status (active, inactive or error) and lastHealthCheck. Speech-to-text entries add supportsEndpointing and languages.
status: "inactive" is not something the product sets. The health sweep only ever sets error. An inactive provider means stt.json or tts.json was hand-edited.
Voice metrics
Three instruments on the Sophon.Voice meter, exposed on Sophon's Prometheus endpoint at /metrics, which requires authentication:
| Instrument | Type | Tags | What it measures |
|---|---|---|---|
sophon.voice.time_to_first_audio | Histogram, ms | — | Turn start to the first audio chunk reaching the client |
sophon.voice.tts_synthesis_latency | Histogram, ms | provider | Per-sentence synthesis latency |
sophon.voice.tts_provider_failures | Counter | provider, reason | Text-to-speech provider failures |
All three are recorded on the text-to-speech streaming path only. A turn with no active text-to-speech provider records nothing, and there are no speech-to-text metrics at all — no transcription latency, no speech-to-text failure counter. Watch the speech-to-text side with sophon voice status and the logs instead. Speech-to-text log lines carry the session id and an utterance id as structured fields, so one transcript can be followed end to end.
Configuration keys
Sophon:Voice
Read live, so an edit applies to the next turn without a Gateway restart.
| Key | Default | Effect |
|---|---|---|
NarrationGapSeconds | 12 | Silence, with no new reply text, before a spoken progress line becomes eligible |
NarrationIntervalSeconds | 20 | Minimum spacing between two spoken progress lines |
SpeakProgressUpdates | true | Master switch for spoken progress lines during long tool work |
MaxSpokenSentenceChars | 280 | Sentences longer than this are clause-split before synthesis |
BargeInMinWords | 2 | Reserved and unused. Declared for a future barge-in feature; no code path reads it, and setting it changes nothing |
None of these keys affects the browser path — they all govern the provider streaming path, so a Gateway with no text-to-speech provider is unaffected by every one of them. See Spoken Replies.
Sophon:Voice:HealthMonitor
| Key | Default | Effect |
|---|---|---|
Enabled | true | Master switch for the health sweep |
IntervalMinutes | 15 | How often every registered provider is re-checked |
Sophon:Voice:NodeWake
| Key | Default | Effect |
|---|---|---|
Enabled | false | Off: every node-token voice route answers 404 |
MaxTranscriptChars | 2000 | A longer wake transcript is rejected with 400 |
EventsPerMinute | 6 | Per-node wake budget, and the burst size; over it, 429 with Retry-After. Read once at startup, so changing it needs a restart |
Enabled is checked per request and takes effect without a restart. See Privacy & Data before turning any of this on.
Troubleshooting
| Symptom | What is actually happening |
|---|---|
| The banner says voice "isn't configured" | Info tone: no active provider on that side. Expected on a fresh install — voice still works through the browser's own recognition and voice. |
| The banner says a provider "is failing" | Warning tone: at least one registered provider on that side failed its health checks and none on that side is active. It is named on the banner. A failing provider next to a healthy active one shows no banner at all. |
| The mic button does nothing in Firefox | With no speech-to-text provider, capture uses the browser's own recognizer, which Firefox does not have. The button is still shown and fails when pressed, and hands-free is not offered. Use Chrome or Edge, or add a speech-to-text provider. |
| Hands-free stops after about two minutes | The inactivity limit (default 120 s) turned Conversation mode off. With browser recognition that is usually what ends a session: the recognizer ends itself within seconds of any pause and hands-free just re-arms, so the 30-second silence timeout rarely bites. With a speech-to-text provider the mic stays open through a pause, so the silence timeout stops it first. Both say so on screen. |
| A prompt stopped re-asking and is just waiting | Two unclear answers each triggered a re-ask; a third closes the mic rather than guessing or looping. On the Voice page and the Command Bridge a fresh mic press re-opens capture for one more attempt. Mobile caps re-prompts the same way, and an answer the recognizer ended in silence counts as an attempt. The card stays and its countdown keeps running — tap Approve or Reject, or press the mic. |
| A reply carries a red "not spoken" notice | Provider speech was expected and cannot happen. The notice names the actual cause: a registered provider failing its health checks, providers registered but all turned off, nothing registered with speech output set to Configured provider only, or nothing registered with Native speech fallback off. Each case has a second wording for "and nothing will speak it either". In hands-free the notice clears on the next re-arm; in push-to-talk it stays until you press the mic. The six exact strings are on Spoken Replies. |
| The reply is in the chat but genuinely nothing is spoken | Real causes of silence: Native speech fallback is off; speech output is Configured provider only, which silences the browser entirely; speech output is OS / browser only with native fallback off; the browser has no speech synthesis; or the voice session itself ended with an error. Check the last one before assuming the reply was never spoken. |
| A node's device page has no Voice tab | The tab needs both an Admin viewer and a node that reports at least one voice.* command capability. A non-admin owner never sees it, because the runtime routes it would call are Admin-only, and an older or not-yet-reconnected node build does not report the capability at all. |
| A cancelled approval or question is announced as "timed out" | Fixed in 2.0: both cancel paths now record and announce cancelled, and the cards show Cancelled. If you still hear the timeout line, the Gateway predates the fix, or the request's own timeout genuinely ended the wait first. |
| A provider that recovered still shows as failing | The sweep runs every 15 minutes, and status is in memory. Wait for the next sweep, or restart the Gateway, which resets every provider to the status on disk. |
Where to go next
- Speech Providers — priority, failover and where credentials live
- Set Up Voice — adding and testing a provider
- Spoken Replies — the "not spoken" notice in full, and the audio cache
- Privacy & Data — where audio goes, and the node voice routes
- Limits — what voice does not do yet
- Common Issues — the rest of the product
Voice Settings
A field-by-field reference for Voice Center — General, Listening, Conversation mode and TTS providers — with the real ranges, the settings that only apply to Sophon Mobile, and the three that do nothing where you would expect.
Privacy & Data
Where your audio actually travels on each listening and speaking path, what Sophon stores and for how long, what is off by default, and exactly what the node voice routes expose once an operator turns them on.