Sophon 2.0 is here
All Posts
announcement
release

Sophon 2.0: A Rebuilt Voice Engine, Sophon Graph, a Full-Screen CLI, and Node Control Runs

Sophon Team·

Sophon 2.0 is a release about the substrate: how work runs, where its state is kept, and who has the authority to let it act. Voice requests now run as regular tasks on your chat, queued and approval-gated like a typed message, with a smaller, voice-focused tool set. Approved multi-step plans run as durable graphs that save their progress as they go. The terminal gets a full-screen chat with approvals and questions built in. And desktop control on Sophon Node now comes as short, gated runs, with the limits stated up front.

Voice, Rebuilt Underneath

Open the Voice page after upgrading and it will look familiar. The orb, the controls and Voice settings are much as they were in 1.15. The 2.0 work is in the engine that decides when Sophon listens, when it speaks, and where a reply goes.

The Voice page, the in-chat voice ribbon and the Command Bridge now run on the same voice engine, so listening, playback and interrupting follow the same rules on all three. A voice request runs as a regular agent task on your chat session. It is queued and approval-gated the same way as a typed message, and it appears live in the chat thread. Voice turns use a smaller, voice-focused tool set than typed chat. If a voice request fails, Sophon tells you by voice or on screen instead of going quiet.

Voice sessions now survive reconnects. If your connection drops while Sophon is answering, voice picks back up when you reconnect. You hear the rest of the reply from that point (audio sent while you were offline isn't replayed), and the full reply is in the chat. Every open tab attached to the same voice conversation hears the reply. The ribbon talks on the chat it was opened in and uses your speech-to-text provider when one is active; switch chats and it ends there and starts again, listening, on the new chat. The microphone is released every time you leave voice, so no recording indicator is left on.

With a text-to-speech provider configured, Sophon starts speaking a reply sentence by sentence while it is still writing the rest of it. During long, multi-step tasks it can speak a short progress line after a quiet stretch, for example when it moves on to a new step, so long silences are less likely to feel like a dropped connection. A single long step can stay quiet for its whole run, and a one-sentence reply is spoken once it's complete. Before speaking a streamed reply, it strips Markdown formatting and reads links as their text or site name. Code and tables are best read in the chat. All three of these need a text-to-speech provider. Without one, voice still works with your browser's speech recognition and voice (Chrome or Edge), and the reply is spoken once it's complete.

Conversation mode (hands-free) now works end to end: speak, pause, hear the answer, and the mic re-opens only after the reply has finished playing. It works with any supported speech-to-text provider, not just ones with built-in end-of-speech detection, because Sophon detects the pause itself. Turn on Conversation mode in Voice settings, and your first mic press starts a hands-free conversation. Silence and inactivity limits close the mic when you walk away. To cut in, press the mic while Sophon is thinking or speaking. That silences it so you can talk. The interrupted reply isn't spoken later, but it still appears in the chat. Interrupting stops the speech, not the task: work already under way keeps running, and by default a new request you speak waits in the queue behind it.

On the Voice page and the Command Bridge, you can now answer approvals and questions out loud. Approvals follow the risk level. Low- and Medium-risk actions take a plain "yes". High- and Critical-risk actions need you to say "yes, approve", and a bare "yes" isn't enough. A "yes" that also says "no", "wait" or "don't" counts as unclear, so Sophon asks again. A plain "no" rejects. Unclear answers are asked again at most twice; after that the mic closes and the prompt waits for a tap. Sophon never guesses. When an agent asks a question with options, Sophon reads it out with numbered choices. Answer by saying the number or the option's name, or tap it. Spoken numbers work for the first four options; beyond that, say the name or tap. Prompts are read in the browser's built-in voice, not your provider voice. If you've turned the browser voice off in Voice settings, or chosen provider-only speech output, they're shown but not read aloud. The in-chat ribbon leaves approvals and questions on the chat's own cards.

Prompts stay with their own conversation. Voice only speaks prompts that belong to its own conversation, so an approval from an unrelated session never interrupts you. Answer an approval in one place, whether by voice, on a chat card, in another tab or in the mobile app, and any other open chat card or voice prompt for it closes on its own (the Dashboard's Approvals page catches up on its next refresh). The Command Bridge's hands-free voice bar listens once you arm it. When it stops, after silence, inactivity or an error, it switches itself off and shows why instead of looking ready while nothing is listening. If an approval or question is waiting, pressing the Bridge mic lets you answer it out loud, and a separate Stop button turns voice off without answering, and the prompt stays on its chat card. Sophon Mobile applies the same "yes, approve" rule to High- and Critical-risk approvals by voice, and only answers approvals for its own voice conversation. In the mobile app, questions are answered on screen.

Voice is also easier to operate. Sophon checks the health of configured speech providers periodically in the background, and a single transient failure doesn't flag a provider as broken. When voice can't use a speech provider, the Voice page explains why in plain words. Either speech isn't set up and your browser's speech is used, or the configured provider is failing, and the page names that provider. The notice appears only when no working provider of that kind is left, and a provider is flagged only after repeated failed checks, so its status can trail an outage. Speech provider requests have time limits, so a provider that stops responding shows up as a named error instead of a stuck conversation. From the terminal:

sophon voice status

This shows the health of every configured speech provider and whether hands-free is available, with server and browser hands-free listed separately. sophon voice runtime start now works, and sophon voice settings set no longer resets your personal listening settings.

Where your audio goes: nothing is preconfigured. An admin adds speech providers, sets their priority and tests them, and provider credentials are kept in Sophon's credential vault. In the Dashboard, with a speech-to-text provider, your audio goes through your Gateway to that vendor for transcription; Sophon Mobile uses your phone's own recognizer first. Without one, your browser handles recognition and may send the audio to its vendor's speech service. Voice is available on every tier. What it doesn't do yet: there's no automatic talk-over detection (you interrupt with the mic), the CLI manages voice but doesn't hold a conversation, and paired nodes don't listen. See Voice or the Voice docs.

Sophon Graph: Plans That Run as Durable Graphs

Sophon Graph is on by default. Approved multi-step plans now run as durable graphs. Steps that don't depend on each other run in parallel, each step runs in its own isolated session, and progress is saved to the database as the run goes. Plans made by scheduled, heartbeat and webhook automations also run in Graph mode by default. Each plan step links to its own session, so you can drill in while it runs. Steps don't stream their tokens into the main chat.

Restarts are handled conservatively. If the Gateway restarts mid-run, finished plan steps are never re-run. With default settings, a run that was mid-step pauses, notifies you, and waits for you to resume or cancel it from the plan card in the Dashboard. After a page reload, running plans, revised plans, and the Resume and Cancel controls on interrupted plans all come back correctly.

When a step keeps failing, Sophon retries it with a suggested alternative approach, then revises the failed part of the plan. Revised plans show a version badge. The model writes the plan and its revisions; the substrate enforces the approval rule around them. A revision waits for your approval if it raises the plan's risk, brings in a new kind of tool, or adds steps that may have side effects. With default settings, every revision asks.

Graph-mode plans run under a spending limit, with a cost cap and a token limit. When a run goes over, Sophon stops starting new steps once the running ones finish and tells you in the chat which limit was hit, so a runaway plan can't keep spending. The limit is checked between steps, so a run can overshoot by whatever its running steps cost. A run stopped by its budget shows as failed in the Dashboard, with no Resume control.

Delegation to sub-agents is more structured. Tasks handed to sub-agents can now depend on each other, so they run in ordered waves that pass their results forward. This needs Graph turned on, and the cost cap doesn't cover these waves. Approvals and questions raised inside a plan step or by a sub-agent now appear in your main conversation in the Dashboard, Mobile and the CLI, and in the messaging channel the conversation came from. Plan-step requests are marked with the step they came from. Previously, a sub-agent's approval request could time out without ever being shown to you. One limit: those requests wait in memory, so a Gateway restart fails that step attempt, and an unanswered approval is rejected.

Plan-step and sub-agent sessions now sit in an expandable Agent activity group under their conversation in the Dashboard sidebar and update live. Deleting the conversation deletes its activity too. Failed or interrupted plan-step attempts are archived out of the group and kept for audit. You can also save a completed or failed Graph-mode plan as an editable workflow in the visual builder. It is a one-way export, and a notes dialog lists anything that couldn't carry over.

You choose between Graph and Classic. Switch a single conversation between them, set your own default in Settings → Execution, or use /mode in the CLI. An admin can turn Graph off for the whole workspace, which puts everyone on Classic, the previous step-by-step engine. The Graph runtime supports a single Gateway instance; running multiple replicas isn't supported. See Sophon Graph and Planning and Execution.

The CLI Goes Full Screen

Type sophon and you now get a full-screen chat by default. It has its own scrollback, a live status header and footer, streamed Markdown with live code highlighting, and a resume card on exit so you can always pick the session back up. Approvals and questions arrive as cards in the terminal. Approvals show risk, parameters, a scrollable preview and a live countdown, with one-key Approve, Reject, Skip, and Edit where the tool allows it. Questions show numbered options and also accept a free-text answer typed into the composer. If a card times out, the transcript tells you so, and an approval that goes unanswered is treated as rejected.

Common tasks open as dialogs on the same screen. /theme previews the whole screen live before you commit, /resume opens a session picker, /plan shows the full plan as an overlay, and a progress row keeps the running plan step visible. Shift-drag or F2 selects transcript text, and resumed sessions show your recent messages behind a "resumed" divider. Once the provider reports usage, each reply ends with a usage line (agent, model, tokens). The footer keeps the running total, and /tokens shows this session's breakdown. Typing / lists every command that works in the current chat view, /export writes a file again, and /quit is new.

The older prompts are still there. Small windows, redirected output and the Dashboard's built-in terminal automatically get the pinned-input or classic prompt instead. When a Windows screen reader is active, or SOPHON_CLI_ACCESSIBLE=1 is set, the chat uses the linear classic prompt. In the full-screen chat, Alt+Enter inserts a newline, the mouse wheel scrolls but doesn't click, and there's no Ctrl+R history search yet. To opt out, or to recover a stuck terminal:

SOPHON_CLI_FULLSCREEN=0 sophon      # skip full screen for this session
sophon doctor --reset-terminal      # recover a stuck terminal

Set "fullScreen": false in cli.json to turn it off for good. See Sophon CLI and Interactive Chat.

Sophon Node: Desktop Control as a Gated Run

The agent now drives a paired desktop through Control Runs, short batches of actions. Each batch stays on one machine and stops at the first failure. When every action succeeds, the batch ends with a fresh screenshot of the result, shown in the Dashboard's live panel. Some machines stop checking in, because they're asleep or the Node agent isn't running. A run aimed at one of them now fails at once with a clear "machine is offline" message instead of waiting out the timeout. The Agent Drive panel in Dashboard chat now has a Stop control that ends the agent's turn, so no further actions are sent to the machine. An action already running may finish first, and Stop is available in the Dashboard only.

Only Claude models receive the screenshots as images. With Claude through the Anthropic API or a directly connected Claude subscription, the latest screenshot from each desktop step goes to the model as an image. The screen geometry comes with it, so clicks map back to real pixels. Other providers, including the Claude Code CLI, Bedrock and Copilot, get a text description only. On Windows, the agent can read the real buttons, text boxes and menu items inside windows, with their names and on-screen positions. That lets it find the control it needs instead of guessing from pixels. Password fields are never returned. macOS and Linux keep their existing, more limited support. The agent can also ask a paired device for its foreground app, active window and focused control without taking a screenshot.

Device security is tighter. The agent can reach only devices owned by the requesting user, and launching, closing or focusing an app now asks for approval. The ungated single-command passthrough tool is removed. At the default approval threshold, the first desktop action in a conversation asks for approval, even a screenshot. In an interactive chat, approving a desktop action also covers later desktop actions up to that risk level in the same conversation, for a limited time. That consent never covers Critical actions, and a Gateway restart clears it. Shell commands always ask again, and workflows, scheduled runs and MCP calls ask every time.

File access needs a separate opt-in permission. With it, the agent can list, read and write text files, but only inside folders you allow (Desktop, Documents and Downloads by default). Writing a file is a high-risk action that needs your approval, binary files are refused, and each read or write has a size cap.

Activity history is opt-in per device. A device you enable records which app and window you're working in to a local database on that machine. Common password managers and private-browsing windows are skipped before anything is written. They are matched by app name and window title, and you can edit both lists. Reading the history needs a separate permission that isn't granted by default. Once a device has that permission and is online, you manage recording from the device's settings page in the Dashboard. There you can turn it on or off, pause it, set the interval and retention, edit the exclusion lists, or erase the history. A day-by-day timeline lets you delete any single entry. Activity screenshots are a second, separate opt-in. With it on, a screenshot is saved on the device each time a new activity entry starts. Screenshots are kept for less time than the text history, within a disk cap. The history is stored on the device, but any part the agent reads, and any screenshot it views, goes through the Gateway to your model provider for that turn and stays in the conversation. Only a Claude model can look at a saved screenshot; other providers see the text timeline.

What hasn't shipped: Node binaries aren't code-signed, and there's no installer package or working tray icon. There's no kill switch on the machine itself, and nodes don't listen: there's no microphone input or wake word. See Sophon Node, Control Runs and desktop history.

Also in 2.0

Approvals That Ask Once, Where You Are

Tools that need approval now ask once, before they run, instead of twice, and a rejected action is reported back to the agent as rejected. (A shell command on a paired desktop can still show two prompts at the default threshold.) Approval requests raised during a conversation on a messaging channel such as Telegram or WhatsApp are now sent to that channel. You answer by replying approve or reject with the request's id. Those replies are accepted only from the chat the request was delivered to, so members of other groups or chats on the same bot can no longer answer it. Anyone in the delivery chat still can. After a Gateway restart, a pending approval is checked against the channel's default chat instead.

Threshold and per-tool changes saved in the Dashboard take effect immediately, with no Gateway restart. Each approval is recorded exactly once, so when two devices answer the same card, only one answer counts. An approval that timed out or was cancelled can never later become an approval. Cards on the Dashboard and Mobile now say how an approval ended: Approved, Rejected, Cancelled or Timed out. The approval quiet-hours controls, which were never enforced, are gone. See Approval Gates.

Prompt-Injection Modes

An admin can now choose Off, Warn (the default) or Block in the Dashboard under Settings → Security. Block stops a turn only when third-party content matches a high-severity pattern. Third-party content means a tool result, a webhook payload, or a channel message from a group chat or from someone other than you. Your own direct messages and the memory snapshot are never blocked. A detection appears as a card in the conversation on Dashboard, CLI and Mobile, and stays after a reload. It is recorded in the audit log, and you get an alert when it stops a scheduled run. A plan step the guard stops fails with the guard's explanation and is not retried or re-planned around. Be clear about what this is: a phrase filter, not a security boundary. Obfuscated or split instructions can get through, and tools called through the embedded MCP server aren't scanned. See Prompt Injection Defense.

Chat That Keeps Up

The model's context is now built from the most recent part of a long session instead of its oldest messages. A repeated short reply like "yes" no longer causes the current message to be skipped. Your next message no longer waits for after-reply bookkeeping, memory indexing runs off the reply path, and stuck tasks are marked as lost automatically. Plans, approvals, questions, and the live desktop and browser panels now share one card design across all four themes, and risk badges show the real level.

Marketplace Updates

Settings → Marketplace has an Updates panel where you can check installed packages for updates, update one, or update all. Automatic updates are opt-in and off by default. When an admin turns them on, the regular update check also installs new package versions through the same verified install path, and rolls back if an update fails. See Installing Packages.

A Bigger Model Catalog

Settings → Models now covers well over a hundred models across many more provider types. It shows each exact model ID (copyable) and its aliases, with a provider filter. Usage of models newly added to the catalog used to be recorded at zero cost. It now carries an estimated cost based on list prices, so spending budgets count it. Models the catalog lists as vision-capable are now used for image understanding and document text recognition (OCR). See the model catalog.

Developer Tools

sophon dev validate --json and sophon dev build --json each print a single JSON document, including the exact list of files in the built package. Failing sophon dev commands now exit with a non-zero code, so scripts and CI no longer treat a broken project as a success. Global flags such as --json and --gateway-url can now go after any subcommand. Builds and installs no longer include .env files (including .env.example), a .venv folder, or editor folders such as .vscode and .idea. Packages built on Windows now install on macOS and Linux. sophon dev new skill creates Python and C# skills that run in the sandbox without changes and come with tests. Before you upload, sophon dev validate warns about skill manifest problems that the Marketplace's publishing rules flag. Sophon DevStudio, an IDE for building skills, agents and plugins, is being developed as a separate product. It's in early access with no public release yet. The sophon dev commands remain the supported way to scaffold, validate, build and install packages. See the command reference.

For Operators

The Gateway's authenticated Prometheus metrics endpoint now exports Graph and voice metrics. For Graph runs, it counts runs started, completed, paused and resumed, plus replans and budget stops. For voice, it records time to first audio, text-to-speech latency per provider, and text-to-speech provider failures, whenever a text-to-speech provider speaks a reply. The SQLite database now uses write-ahead logging with a busy timeout, so reads no longer block behind writes. Each conversation event gets a unique sequence number, even when events are written at the same time.

Upgrading

The first time the 2.0 Gateway starts, it automatically applies five database migrations. There is no automatic downgrade path, so stop the Gateway and back up your data directory before upgrading. One migration renumbers duplicate conversation-event numbers left by earlier versions; no events are removed. SQLite now keeps write-ahead-log side files next to the database, so take manual backups with the Gateway stopped, or include those files.

A few defaults and interfaces change:

Sophon Graph is on. An admin can turn it off for the workspace in Settings → Execution, or with Sophon:GraphRuntime:Enabled=false, which puts everyone on Classic. The Graph runtime supports a single Gateway instance only.

The CLI opens full screen. Opt out with SOPHON_CLI_FULLSCREEN=0, or "fullScreen": false in cli.json.

Prompt-injection detection defaults to Warn, which shows detections in the conversation. Block is opt-in.

Integrations. The session list API now returns only top-level conversations by default. Request the full flat list, or one conversation's agent activity, explicitly when you need it. Deciding an approval that is already decided now returns the recorded outcome instead of overwriting it, and unrecognized outcome values are rejected.

Node. The single-command passthrough tool is gone, and app launch, close and focus ask for approval. Every desktop command now returns camelCase fields, and the app-launch parameter is named name. The node-facing voice routes are off by default and return 404 until an operator turns them on with Sophon:Voice:NodeWake:Enabled.

Budgets. Newly catalogued models now carry an estimated cost, so spending budgets may trigger where they did not before.

Marketplace automatic updates stay off until you turn them on, and Sophon Mobile keeps working without an update. Full steps are in Backup and Upgrade.

Get Sophon 2.0

Installers are on the downloads page, new installs start at Installation, and the complete list of changes is in the changelog. Then start with the change you'll hear first: open the Voice page, turn on Conversation mode in Voice settings, and press the mic.