Sophon 2.0 is here
Sophon Docs
Sophon Voice

Privacy & Data

Where your audio actually travels on each listening and speaking path, what Sophon stores and for how long, what is off by default, and exactly what the node voice routes expose once an operator turns them on.

Voice is not local processing. On every path, audio or text leaves the machine you are speaking into — the only question is which vendor receives it. This page says which one, for each path, so you can decide what is acceptable before you turn anything on.

Voice is never on-device. With a speech provider, your audio goes through your Gateway to that vendor. With no provider, the browser's or the phone's own recognizer decides where the audio goes, and it needs an internet connection to work. Do not assume audio stays on your machine on any path.

Where your audio goes

PathWhat leaves your machineWho receives it
Dashboard, with a speech-to-text providerMicrophone audio as 16 kHz mono PCM, streamed over your signed-in connectionYour Gateway, which passes it to the speech-to-text provider you configured
Dashboard, with no providerWhatever the browser's own speech recognition sendsThe browser vendor's speech service — Google's for Chrome, Apple's for Safari. Sophon never sees the audio, only the text the browser hands back
Sophon Mobile, by defaultWhatever the phone's own recognizer sendsThe phone platform's speech service
Sophon Mobile, with the fallback onThe recording, uploaded once you stop speakingYour Gateway, which passes it to your speech-to-text provider
Speaking a reply, with a text-to-speech providerThe text of each reply sentence, one sentence per requestThe text-to-speech provider you configured
Speaking a reply, with no providerThe reply text, handed to the browser's or phone's own speech synthesisThe device — but the voice you picked may be a network voice, in which case the vendor receives the text

Those vendors' own terms and retention policies apply to everything they receive. Sophon does not negotiate or override them.

The server speech-to-text path, in detail

In the Dashboard, microphone audio is downsampled in the browser to 16 kHz mono PCM and streamed to your Gateway over the same authenticated connection the chat uses. The Gateway holds it in memory:

  • A streaming provider receives the audio as it arrives and returns interim and final transcripts.
  • A buffered provider receives one complete recording per utterance. The Gateway assembles a WAV file in memory once capture ends and posts it.

Sophon Mobile cannot stream audio to the Gateway. Its fallback uploads the finished recording to POST /api/voice/transcribe after you stop speaking — a normal file upload, capped at 10 MB, accepting .wav, .mp3, .m4a, .ogg, .webm and .flac. There is no live Gateway transcription on the phone.

The browser and phone recognizers

Browser speech recognition is not local processing, even though it needs no Sophon configuration. It requires an internet connection, and the browser vendor decides where that audio goes. The same is true of the phone's recognizer on Mobile. This is the path you are on with no speech-to-text provider configured, and it is the default on Mobile regardless of what is configured.

What Sophon stores

ThingWhere it livesHow long
Your words and the replyThe chat thread, as ordinary messagesLike any typed message — your session history
Reply sentences already spoken by a text-to-speech providerA 32 MB in-memory cache in the Gateway, keyed by the whole sentence plus voice, language, speed and provider typeUntil the Gateway restarts. Nothing is written to disk
Provider API keys and tokensThe credential vault, encrypted at restUntil you remove the provider
Non-secret provider config~/.sophon/config/stt.json and tts.json — id, name, type, endpoint or region, voice, priority, statusUntil you remove the provider
A paired node's voice runtime rowThe Gateway database: the node's last transcript, its last agent reply, its last error and its current stateOverwritten by the next wake event

Sophon does not save your recorded audio. On the Dashboard's streaming path nothing touches the disk at all — audio is held in memory, transcribed, and dropped. On the upload path (Mobile's fallback and sophon voice transcribe) the recording arrives as a normal HTTP file upload, which the web server may spill to a temporary file for the lifetime of the request before deleting it; Sophon itself never writes it to a stored location and keeps no copy.

Speech-to-text log lines carry a session id and an utterance id so a transcript can be correlated end to end. They do not carry the audio.

The spoken-audio cache key carries no tenant. Two users on the same Gateway who share a provider, voice, language and speed will share a cache entry for an identical sentence. Nothing of one user's conversation is readable from another's, but the synthesized audio for a sentence they both happened to hear is the same object. Per-tenant caching is on the roadmap.

What is off by default

  • No speech provider is configured. A fresh Gateway sends no audio and no reply text to any speech vendor. Everything runs on the browser or the phone until an admin adds a provider.
  • Conversation mode is off, per user. Sophon listens only after you press the mic or arm the Command Bridge, and it stops on its own. There is no wake word and no always-on listening anywhere in the product.
  • The node voice routes are off, Gateway-wide. See below.
  • Adding, testing and removing speech providers is Admin-only, and the provider list is per Gateway, not per tenant — every user on that Gateway uses the same providers. Everyone can change their own preferences, and everyone can use voice.

The microphone is released every time you leave voice, so no recording indicator lingers.

The node voice routes

A paired Sophon Node can read text aloud through its operating system's own speech synthesiser. It cannot listen: there is no microphone capture on a node and no wake word. The routes below exist for a future node-side capture feature, and no shipped client calls them.

They are off by default. With Sophon:Voice:NodeWake:Enabled unset or false, all three answer 404 — before the node's token is even examined, so a disabled Gateway does not reveal whether a token is valid.

Once an operator turns them on, the three routes are not protected equally:

RouteWhat it takes
GET /api/nodes/me/voice/runtimeA valid node token. Nothing more
POST /api/nodes/me/voice/stateA valid node token. Nothing more
POST /api/nodes/me/voice/wakeA valid node token, plus the node's voice.runtime scope (or the node.command wildcard), the node's own tenant, a per-node budget of 6 events per minute (over it, 429 with Retry-After), and a transcript of at most 2000 characters (over it, 400)

The wake route runs a full agent task as the node's owner, from a transcript the node supplies. That is why it carries the scope check, the rate limit and the length cap, and why none of them applies to the two read-and-report routes next to it. voice.runtime is not granted by default — see Node permissions. Treat enabling this as granting a device the ability to start work in your name.

Approvals are never spoken to a node. A task started from a node voice session still raises its approval in the usual places — the Dashboard, Sophon Mobile, the CLI, or the channel the conversation came from — and waits there.

Where to go next