Sophon 2.0 is here
Sophon Docs
Sophon Voice

Health & Troubleshooting

What each availability banner means, how the provider health sweep works and why status can trail an outage, the per-call time limits, sophon voice status and the capabilities endpoint, the three voice metrics, the Sophon:Voice configuration keys, and a symptom-by-symptom table.

Voice is built to say where it broke. A provider that is not set up reads differently from one that is failing, every provider call is time-limited so a hung vendor surfaces as a named error, and one CLI command prints the whole picture.

This page is the operator's view: what the banners mean, what is checked and how often, and what to do when something is wrong.

The availability banner

The Voice page reads GET /api/voice/capabilities and shows at most one message per side.

What you seeToneMeaning
No banner—At least one active provider exists on that side
"Speech-to-text isn't configured"InfoNothing registered on that side, or everything registered is turned off. "Server speech recognition is not set up, so voice is using your browser's built-in recognition."
"Text-to-speech isn't configured"InfoThe same, for replies. "Replies use your browser's built-in voice."
"A speech provider is failing"WarningAt least one registered provider on that side failed its health checks and none on that side is active. The description names the provider: "Deepgram is failing, so voice is using your browser's built-in speech recognition instead."

The banner only reports whether a side has anything usable at all. A second, failing provider sitting next to a healthy active one shows no banner — voice simply uses the healthy one. To see the health of every provider, including the ones the banner hides, run sophon voice status or read the capabilities endpoint.

Both banners link to Open Voice Center. The info tone is the expected state on a fresh install, not a fault.

The health sweep

A background sweep calls every registered provider's own health check and updates its status in the Gateway's memory.

BehaviourValue
How often every registered provider is re-checked15 minutes (Sophon:Voice:HealthMonitor:IntervalMinutes)
Consecutive failures before a provider is marked failing2
Successes needed to clear it1 — recovery is immediate
Providers that are explicitly turned offSkipped
Where the status livesIn memory only
What a Gateway restart doesResets every provider to whatever stt.json / tts.json last held, until the next sweep

Two consequences follow, and both are deliberate:

  • Status trails an outage. Two consecutive failures at a 15-minute interval means a provider that dies right after a sweep can take roughly 30 minutes to show as failing. Degradation requires confirmation so that one network blip does not make voice look broken; recovery does not, so a provider that comes back is trusted on the next sweep.
  • Status does not survive a restart. After a Gateway restart a failing provider reads as active until it is checked again. LastHealthCheck on each provider tells you when it was last looked at.

The sweep never invents an unhealthy entry for something that was never registered, so "not configured" can never render as "failing".

Per-call time limits

Every speech-provider request is bounded, and a timeout is reported as a named error that says which provider and which operation ran out of time — not as a stuck conversation.

CallLimit
Health check, and listing voices or languages10 s
Speech synthesis (one sentence)30 s
Transcription of a captured recording45 s
Connecting — a streaming handshake, or the first response of a streamed call15 s

An open streaming session is deliberately not bounded. The 15-second limit covers establishing the channel; what flows afterwards has no aggregate cap, because a legitimately long reply would otherwise be cut off mid-sentence.

Separately, once audio capture ends the Gateway gives the provider a bounded window to produce the transcript — max(5 seconds, seconds of audio ÷ 4) for a streaming provider, never less than 55 seconds for a buffered one. See Conversations.

sophon voice status

sophon voice status

One table — Kind, Name, Status — built from the capabilities endpoint:

KindNameStatus
Speech-to-texteach configured provideractive, inactive or error
Speech-to-text(none configured)browser fallback
Text-to-speecheach configured provideractive, inactive or error
Text-to-speech(none configured)browser fallback
Hands-freeserver STTavailable, or not available — no active speech-to-text provider
Hands-freebrowser (Dashboard /voice)works without a provider in Chrome, Edge or Safari

The two hands-free rows are reported separately on purpose. Collapsing them would read as "hands-free is broken" on a Gateway with no speech-to-text provider, even though the Voice page's browser path works there.

The capabilities endpoint

GET /api/voice/capabilities is what the banner, the Dashboard and the CLI all read. It is the right thing to poll from a script or a support runbook.

FieldWhat it means
serverTranscriptionAvailableAt least one speech-to-text provider is active
providerSpeechAvailableAt least one text-to-speech provider is active
handsFreeSupportedServer-side hands-free is available — identical to serverTranscriptionAvailable, because any active provider can be endpointed server-side. It says nothing about the browser path
nativeInputSupported / nativeOutputSupportedAlways true; the browser paths are always offered
singleVendorOptionsVendors that are active on both sides
sttProviders / ttsProvidersEvery registered provider, including inactive and failing ones

Each provider entry carries id, name, providerType, vendor, supportsStreaming, status (active, inactive or error) and lastHealthCheck. Speech-to-text entries add supportsEndpointing and languages.

status: "inactive" is not something the product sets. The health sweep only ever sets error. An inactive provider means stt.json or tts.json was hand-edited.

Voice metrics

Three instruments on the Sophon.Voice meter, exposed on Sophon's Prometheus endpoint at /metrics, which requires authentication:

InstrumentTypeTagsWhat it measures
sophon.voice.time_to_first_audioHistogram, ms—Turn start to the first audio chunk reaching the client
sophon.voice.tts_synthesis_latencyHistogram, msproviderPer-sentence synthesis latency
sophon.voice.tts_provider_failuresCounterprovider, reasonText-to-speech provider failures

All three are recorded on the text-to-speech streaming path only. A turn with no active text-to-speech provider records nothing, and there are no speech-to-text metrics at all — no transcription latency, no speech-to-text failure counter. Watch the speech-to-text side with sophon voice status and the logs instead. Speech-to-text log lines carry the session id and an utterance id as structured fields, so one transcript can be followed end to end.

Configuration keys

Sophon:Voice

Read live, so an edit applies to the next turn without a Gateway restart.

KeyDefaultEffect
NarrationGapSeconds12Silence, with no new reply text, before a spoken progress line becomes eligible
NarrationIntervalSeconds20Minimum spacing between two spoken progress lines
SpeakProgressUpdatestrueMaster switch for spoken progress lines during long tool work
MaxSpokenSentenceChars280Sentences longer than this are clause-split before synthesis
BargeInMinWords2Reserved and unused. Declared for a future barge-in feature; no code path reads it, and setting it changes nothing

None of these keys affects the browser path — they all govern the provider streaming path, so a Gateway with no text-to-speech provider is unaffected by every one of them. See Spoken Replies.

Sophon:Voice:HealthMonitor

KeyDefaultEffect
EnabledtrueMaster switch for the health sweep
IntervalMinutes15How often every registered provider is re-checked

Sophon:Voice:NodeWake

KeyDefaultEffect
EnabledfalseOff: every node-token voice route answers 404
MaxTranscriptChars2000A longer wake transcript is rejected with 400
EventsPerMinute6Per-node wake budget, and the burst size; over it, 429 with Retry-After. Read once at startup, so changing it needs a restart

Enabled is checked per request and takes effect without a restart. See Privacy & Data before turning any of this on.

Troubleshooting

SymptomWhat is actually happening
The banner says voice "isn't configured"Info tone: no active provider on that side. Expected on a fresh install — voice still works through the browser's own recognition and voice.
The banner says a provider "is failing"Warning tone: at least one registered provider on that side failed its health checks and none on that side is active. It is named on the banner. A failing provider next to a healthy active one shows no banner at all.
The mic button does nothing in FirefoxWith no speech-to-text provider, capture uses the browser's own recognizer, which Firefox does not have. The button is still shown and fails when pressed, and hands-free is not offered. Use Chrome or Edge, or add a speech-to-text provider.
Hands-free stops after about two minutesThe inactivity limit (default 120 s) turned Conversation mode off. With browser recognition that is usually what ends a session: the recognizer ends itself within seconds of any pause and hands-free just re-arms, so the 30-second silence timeout rarely bites. With a speech-to-text provider the mic stays open through a pause, so the silence timeout stops it first. Both say so on screen.
A prompt stopped re-asking and is just waitingTwo unclear answers each triggered a re-ask; a third closes the mic rather than guessing or looping. On the Voice page and the Command Bridge a fresh mic press re-opens capture for one more attempt. Mobile caps re-prompts the same way, and an answer the recognizer ended in silence counts as an attempt. The card stays and its countdown keeps running — tap Approve or Reject, or press the mic.
A reply carries a red "not spoken" noticeProvider speech was expected and cannot happen. The notice names the actual cause: a registered provider failing its health checks, providers registered but all turned off, nothing registered with speech output set to Configured provider only, or nothing registered with Native speech fallback off. Each case has a second wording for "and nothing will speak it either". In hands-free the notice clears on the next re-arm; in push-to-talk it stays until you press the mic. The six exact strings are on Spoken Replies.
The reply is in the chat but genuinely nothing is spokenReal causes of silence: Native speech fallback is off; speech output is Configured provider only, which silences the browser entirely; speech output is OS / browser only with native fallback off; the browser has no speech synthesis; or the voice session itself ended with an error. Check the last one before assuming the reply was never spoken.
A node's device page has no Voice tabThe tab needs both an Admin viewer and a node that reports at least one voice.* command capability. A non-admin owner never sees it, because the runtime routes it would call are Admin-only, and an older or not-yet-reconnected node build does not report the capability at all.
A cancelled approval or question is announced as "timed out"Fixed in 2.0: both cancel paths now record and announce cancelled, and the cards show Cancelled. If you still hear the timeout line, the Gateway predates the fix, or the request's own timeout genuinely ended the wait first.
A provider that recovered still shows as failingThe sweep runs every 15 minutes, and status is in memory. Wait for the next sweep, or restart the Gateway, which resets every provider to the status on disk.

Where to go next