Sophon 2.0 is here
Sophon Docs
Sophon Voice

Conversations & Hands-Free

Push-to-talk and conversation mode, how a spoken turn runs, what interrupting does and does not stop, reconnects and session limits, and how long Sophon waits for a transcript.

There are two ways to talk: press the mic for each turn, or turn on conversation mode and let Sophon re-open the mic after every reply. Both put the same kind of task on the same chat.

Push to talk

This is the default, and it is what you get with conversation mode off.

  • With a speech-to-text provider. Click the mic to start. Click it again to send — the control reads Stop listening and send. While you talk, the preview card shows a Rec · Listening… header with an elapsed timer, and offers Send now and Cancel.
  • With the browser recognizer. There is no send step. The recognizer ends the utterance itself at a pause, and that sends it.
  • On Sophon Mobile, tap to talk and tap to stop.
  • A single utterance is capped: the default maximum is 60 seconds, adjustable from 5 to 300 seconds. Audio that hits the cap is still transcribed and answered.

Conversation mode

Conversation mode is hands-free: you speak, Sophon answers, and the mic re-opens on its own. It is a per-user preference and it is off by default.

Turn it on in Voice settings under Conversation mode, in Mobile's Voice Center, or with sophon voice settings set --conversation-mode true. On the Voice page the hands-free toggle only appears once the preference is on, and your first mic press starts the hands-free conversation.

It is not a wake word and not always-on listening. Sophon listens only after you press the mic or arm the Command Bridge, and it stops on its own.

How an utterance ends

PathWhat ends the utterance
DeepgramThe provider's own end-of-speech detection
The other five speech-to-text providersA server-side silence detector: at least 300 ms of speech must be heard first, then a run of silence at least as long as the endpointing gap
Browser recognitionThe recognizer ends itself at a pause

The endpointing gap is 300 ms by default host-wide. Your personal override accepts 100–3000 ms in Voice settings and with sophon voice settings listening --endpointing-ms.

When the mic re-opens

The mic re-opens only when the server has finished sending the reply and local playback has finished. It never opens over Sophon's own voice, which is what used to make hands-free record the assistant.

An utterance that produced nothing just re-listens — and an empty re-listen does not reset the inactivity limit.

What closes the mic

LimitDefaultRange in the Dashboard and CLIWhat you see
Silence timeout30 s3–30 s"Listening stopped — no speech was heard."
Inactivity limit120 s30–600 s"Hands-free was paused after a period of inactivity."

The silence timeout mainly bites with a speech-to-text provider, where the mic stays open until something ends the utterance. The API clamps both values more loosely than the Dashboard and CLI allow, so a value set through the API may be wider than the range above.

If the microphone comes back busy while hands-free is running, Sophon retries after 1 s, 2 s and 5 s and then gives up with a message. A microphone whose permission was denied is not retried at all — fix the permission and start again.

How a turn runs

  1. Your words become a regular agent task on the chat session the voice conversation is attached to, and the message is recorded in the thread.
  2. It is queued. The default session policy runs one task at a time and holds up to three pending, so a new spoken request waits behind whatever is already running rather than replacing it. A session whose queue is full rejects the new request instead of dropping the running one.
  3. It runs with the voice tool set. Core tools are datetime.now, memory.search, memory.write, message.send and web.search; the turn starts with a base set of up to 16 tools and can load more on demand through tool search. There is no tool index, so a voice turn reaches a narrower surface than a chat turn by default.
  4. Approvals and policies are the same as for a typed message. The gate does not move because you spoke. Answering out loud has its own rules — see Approvals & Questions.
  5. The reply is delivered to every tab attached to the conversation, and the text lands in the chat thread. A turn that fails is spoken or shown; it is never dropped in silence.

The voice turn itself never raises a question card: the voice tool set has no tool for asking one, so it asks conversationally instead. Question cards you hear read out come from plans and other tasks on the same chat.

Interrupting

Press the mic while Sophon is thinking or speaking to silence it and start talking. On the Command Bridge, a mic press disarms instead, unless a prompt is waiting, in which case it opens the mic to answer.

What that does:

  • Speech stops at the next sentence boundary. Synthesis happens one sentence at a time, so the sentence already in flight finishes.
  • The interrupted reply is not spoken later, but it still appears in the chat thread.
  • The task keeps running. Interrupting stops the speech, not the work: tools already under way run to completion, and by default your next spoken request queues behind that turn.
  • There is no automatic barge-in. Talking over Sophon does not stop it — only pressing the mic does. Detecting speech over playback is roadmap; a configuration key is reserved for it and has no effect today.

Reconnects and sessions

A voice session belongs to the chat, not to the connection that opened it.

  • Reconnects are automatic, at 0, 2, 5, 10 and 30 seconds, and re-attach to the same voice session. You hear the rest of the reply from the point you rejoin.
  • Audio sent while you were offline is not replayed. The full reply text is in the chat thread.
  • Every tab attached to the conversation hears the reply. Ending voice in one tab ends the shared session for all of them.
  • A session survives a disconnect. It is reclaimed only after 15 minutes idle with no turn in flight. After that, a client that comes back is told "Your voice session has ended or could not be found. Please restart voice mode."
  • A Gateway holds at most 50 concurrent voice sessions. Past that you get "Maximum concurrent voice sessions reached."
  • The microphone is released every time you leave voice, so no recording indicator lingers.

Utterance ids

With server-side speech-to-text, each utterance is announced before its transcript arrives, and a client accepts a final transcript only for the most recently announced utterance. That is what stops a slow transcription of something you said earlier from being taken as your answer to the next question.

The protection applies to the server path only. A client using browser recognition has seen no announcement, so it forwards every final transcript. Even on the server path it is not absolute: a result that lands before the next utterance is announced can still be accepted.

How long Sophon waits for a transcript

Once audio capture ends, the Gateway gives the provider a bounded window to produce the transcript:

  • max(5 seconds, seconds of audio ÷ 4) for a streaming provider, which only has to flush what it has already been transcribing.
  • Never less than 55 seconds for the five buffered providers, because they do the entire transcription inside that window. That floor is the provider's own 45-second transcription limit plus a 10-second margin.

Past the window the utterance is abandoned and you get: "Transcription took too long and was abandoned. Please try again."

Provider calls are separately bounded — health checks at 10 seconds, synthesis at 30, transcription at 45 and connecting at 15 — so a provider that stops responding shows up as a named error instead of a stuck conversation. An open streaming session is deliberately not bounded, or a long reply would be cut off mid-sentence. See Health & Troubleshooting.

Where to go next