Changelog
A history of updates and improvements to Sophon.
v2.0.0
Sophon 2.0 rebuilds the voice engine behind the Voice page, the chat voice ribbon and the Command Bridge; Sophon Graph becomes the default way approved plans run; the CLI opens full screen; and Sophon Node gains Control Run.
Voice
Available on every tier. With no speech provider configured, voice uses your browser’s speech recognition and voice (Chrome or Edge), and the browser may send audio to its vendor; admins add speech providers in Voice settings.
- One voice engine: the Voice page, the in-chat voice ribbon and the Command Bridge keep their familiar look but now run on one rebuilt engine, so listening, playback and interrupting follow the same rules on all three.
- Speech that starts early (needs a text-to-speech provider): with a provider configured, Sophon starts speaking a reply sentence by sentence while it is still writing the rest, and a one-sentence reply is spoken once it’s complete; without a provider, the browser’s voice reads the finished reply.
- Hands-free that works: Conversation mode now runs end to end — speak, pause, hear the answer — with the mic re-opening only after the reply has finished playing, and Sophon detects the pause itself, so hands-free works with any supported speech-to-text provider.
- Voice turns are regular tasks: a voice request runs as an agent task on your chat session, queued and approval-gated the same way as a typed message and shown live in the chat thread, with a smaller, voice-focused tool set.
- Interrupt with the mic: press the mic while Sophon is thinking or speaking to silence it and talk; the interrupted reply isn’t spoken later but still appears in the chat, and work already under way keeps running.
- Approvals by voice follow risk: on the Voice page, the Command Bridge and Sophon Mobile, Low- and Medium-risk actions take a plain “yes”, High- and Critical-risk actions need “yes, approve”, a “yes” that also says “no”, “wait” or “don’t” is asked again, and a plain “no” rejects; prompts are read in the browser’s or phone’s built-in voice, not your provider voice.
- Questions by voice: when an agent asks a question with options, the Voice page and the Command Bridge read it out with numbered choices, and you answer by saying the number (one to four) or the option’s name, or by tapping it; option names with symbols can be said as words, such as “C sharp”.
- No guessing, no crosstalk: an unclear answer is asked again at most twice before the mic closes and the prompt waits for a tap, negating an option (“don’t send”) never selects it, and voice only speaks prompts that belong to its own conversation.
- Voice that stays with the conversation: voice started from a chat talks on that chat, and switching chats ends it on the old one and restarts it, listening, on the new one instead of speaking into the old one; every open tab attached to the same voice conversation hears the reply, and after a dropped connection you hear the rest of the reply from the moment you reconnect, with the full reply in the chat.
- Command Bridge voice bar: the hands-free bar listens once you arm it and switches itself off with the reason shown after silence, inactivity or an error, and while an approval or question is waiting the mic opens so you can answer out loud, and a separate Stop button turns voice off without answering, leaving the prompt on its chat card.
- Provider health you can see: configured speech providers are health-checked in the background, speech provider requests have time limits so a provider that stops responding shows up as a named error instead of a stuck conversation, and when voice can’t use a provider the Voice page says whether speech isn’t set up or which provider is failing; the notice appears only when no working provider of that kind is left, and a provider is flagged only after repeated failed checks.
- Progress lines and speakable replies (needs a text-to-speech provider): during long, multi-step tasks Sophon can speak a short progress line after a quiet stretch (a single long step can stay quiet for its whole run), and a streamed reply is spoken without Markdown formatting, with links read as their text or site name; code and tables are best read in the chat.
- Fewer silent failures: a voice request that fails is spoken or shown instead of going quiet, the microphone is released every time you leave voice, and with server speech-to-text a slow transcript of something said earlier is much less likely to be taken as your answer or as a new request.
- Voice from the command line: the new sophon voice status command shows the health of every configured speech provider and whether hands-free is available, sophon voice runtime start now works with --agent and --session, and sophon voice settings set no longer resets your personal listening settings.
Sophon Graph
Graph supports a single Gateway instance; Classic, the previous step-by-step engine, remains a supported choice.
- Sophon Graph, on by default: approved multi-step plans now run as durable graphs, with steps that don’t depend on each other running in parallel, each step running in its own isolated session, and progress saved to the database as the run goes.
- Graph or Classic, your choice: switch a single conversation between Graph and Classic, set your own default on the new Settings → Execution page, or use /mode in the CLI, while admins set parallelism and the cost cap there or turn Graph off for the whole workspace, which puts everyone on Classic.
- Safe restarts: if the Gateway restarts mid-run, finished plan steps are never re-run, and with default settings a run that was mid-step pauses, notifies you and waits for you to resume or cancel it from the plan card in the Dashboard, where plans and their Resume and Cancel controls now come back correctly after a page reload.
- Self-correcting plans: when a step keeps failing, Sophon retries it with a suggested alternative approach and then revises the failed part of the plan; revised plans show a version badge and wait for your approval if they raise the plan’s risk, bring in a new kind of tool or add steps that may have side effects, and with default settings every revision asks.
- Plan budgets: Graph-mode plans run under a cost cap and a token limit, and when a run goes over, Sophon stops starting new steps once the running ones finish and tells you in the chat which limit was hit; the limit is checked between steps, so a run can overshoot by what its running steps cost, and a run stopped by its budget shows as failed in the Dashboard, with no Resume control.
- Dependent delegation: with Graph on, tasks handed to sub-agents can depend on each other, so they run in ordered waves that pass their results forward; the cost cap doesn’t cover these waves.
- Save a plan as a workflow: turn a completed or failed Graph-mode plan into an editable workflow in the visual builder, with a notes dialog listing anything that couldn’t carry over. It is a one-way export.
Sophon CLI
- Full-screen terminal chat: sophon now opens a full-screen chat by default, with its own scrollback, a live status header and footer, streamed Markdown with live code highlighting, and a resume card on exit so you can always pick the session back up.
- Approval and question cards in the terminal: approvals show risk, parameters, a scrollable preview and a live countdown, with one-key Approve, Reject, Skip, and Edit where the tool allows it; questions show numbered options and also accept a typed free-text answer, and an approval that goes unanswered is treated as rejected.
- Dialogs and selection: /theme previews the whole screen live before you commit, /resume opens a session picker, /plan shows the full plan as an overlay while a progress row keeps the running step visible, Shift-drag or F2 selects transcript text, and typing / lists every command that works in the current view.
- Token usage: each reply in the full-screen chat ends with a usage line (agent, model, tokens), the footer keeps the running total and /tokens shows this session’s breakdown, and nothing is shown until the provider reports usage.
- Automatic fallback: small windows, redirected output and the Dashboard’s built-in terminal get the pinned-input or classic prompt, an active Windows screen reader (or SOPHON_CLI_ACCESSIBLE=1) gets the linear classic prompt, SOPHON_CLI_FULLSCREEN=0 turns full screen off, and sophon doctor --reset-terminal recovers a stuck terminal.
Sophon Node
Desktop actions run through a paired Sophon Node and the approval gate; activity or screenshots the agent reads are sent to your model provider for that turn and stay in the conversation.
- Control Run: the agent drives a paired desktop in short batches of actions that stay on one machine and stop at the first failure, and a batch whose actions all succeed ends with a fresh screenshot in the Dashboard’s live panel.
- Screenshots to the model, on Claude only: with Claude models reached through the Anthropic API or a directly connected Claude subscription, the latest screenshot from each desktop step goes to the model as an image, with the screen geometry it needs to map clicks back to real pixels; other providers, including the Claude Code CLI, Bedrock and Copilot, get a text description.
- Plan-level desktop consent: in an interactive chat, approving a desktop action also covers later desktop actions up to that risk level in the same conversation for a limited time; that consent never covers Critical actions and a Gateway restart clears it, shell commands always ask again, and workflows, scheduled runs and MCP calls ask every time.
- Stop and fail fast: the Agent Drive panel in Dashboard chat now has a Stop control that ends the agent’s turn so no further actions reach the machine (an action already running may finish first), and a run aimed at a machine that has stopped checking in fails at once with a “machine is offline” message.
- Windows UI grounding: on Windows the agent can read the real buttons, text boxes and menu items inside windows, with their names and on-screen positions, so it can find the control it needs instead of guessing from pixels; password fields are never returned, and macOS and Linux keep their existing, more limited support.
- File access on your own machine: with a separate opt-in permission, the agent can list, read and write text files only inside folders you allow (Desktop, Documents and Downloads by default), binary files are refused, each read or write is size-capped, and writing a file is a high-risk action that needs your approval (in an interactive chat, one approval can cover later writes in the same conversation for a limited time).
- Activity history (opt-in): a device you enable records which app and window you’re working in to a local database on that machine, skipping common password managers and private-browsing windows (matched by app name and window title, both editable) before anything is written, and once the device has the separate activity-history permission (not granted by default) and is online, the Dashboard lets you browse it by day, filter it, delete single entries and change recording settings without restarting the node.
- Activity screenshots (separate opt-in): the device can save a screenshot as an ordinary image file each time a new activity entry starts, by default kept for less time than the text history and within a disk cap, with Compact, Balanced and Sharp quality presets; when you ask what you were working on, a Claude model can look at the saved screenshot, while other providers see only the text timeline.
Approvals
- One prompt, before the action: tools that need approval now ask once, before they run, instead of twice, and a rejected action is reported back to the agent as rejected (a shell command on a paired desktop can still show two prompts at the default threshold).
- Approvals on messaging channels: approval requests raised in a conversation on a channel such as Telegram or WhatsApp are now sent to that channel, where you answer with “approve <id>” or “reject <id>”, and replies are accepted only from the chat the request was delivered to, though anyone in that chat can answer; after a Gateway restart, a pending approval is checked against the channel’s default chat.
- Plan-step and sub-agent requests in your main chat: approvals and questions raised inside a plan step or by a sub-agent now appear in your main conversation in the Dashboard, Mobile and the CLI and in the messaging channel it came from, marked with the step they came from; these requests wait in memory, so a Gateway restart fails that step attempt, an approval left unanswered is rejected, and an unanswered question lets the step continue on its best judgment.
- Live approval policy: threshold and per-tool changes saved in the Dashboard take effect immediately with no Gateway restart, security settings saves are atomic, and the approval quiet-hours controls, which were never enforced, are removed.
- Resolved everywhere, recorded once: answer an approval by voice, on a chat card, in another tab or in the mobile app and any other open chat card or voice prompt for it closes on its own, cards on the Dashboard and Mobile now say Approved, Rejected, Cancelled or Timed out, and a timed-out or cancelled approval can never later become an approval.
Prompt-injection guard
The guard matches known instruction phrasing; it is a filter, not a security boundary.
- Off, Warn or Block: an admin chooses the prompt-injection mode in the Dashboard under Settings → Security, Warn is the default, and Block stops a turn only when third-party content (a tool result, a webhook payload, or a channel message from a group chat or from someone other than you) matches a high-severity pattern, never for your own direct messages or memory.
- Injection notices: a detection appears as a card in the conversation on Dashboard, CLI and Mobile, stays after a reload, is recorded in the audit log, and alerts you when it stops a scheduled run.
- Blocked steps stay stopped: a plan step the guard stops fails with its explanation and is not retried or re-planned around, and a stopped sub-agent reports a failure instead of passing its refusal on as a finished result.
Marketplace
- Updates panel: check installed packages for updates, update one, or update all from Settings → Marketplace.
- Automatic updates (opt-in, off by default): when an admin turns them on, the regular update check also installs new package versions through the same verified install path, and rolls back if an update fails.
Models and cost
- Expanded model catalog: Settings → Models now covers 181 models from 18 provider types, showing each exact model ID (copyable) and its aliases, with a provider filter, and models the catalog lists as vision-capable are now used for image understanding and document OCR.
- Cost tracking for more models: usage of models newly added to the catalog, previously recorded at $0, now carries an estimated cost based on list prices, so spending budgets count it and may trigger where they did not before.
Chat, sessions and operations
- Long conversations stay current: the model’s context is now built from the most recent part of a long session instead of its oldest messages, and a repeated short reply like “yes” no longer causes the current message to be skipped.
- Faster back-to-back turns: a turn now completes as soon as the answer is delivered, memory indexing runs afterwards off the reply path so your next message isn’t held up, and stuck tasks are marked as lost automatically.
- Unified chat cards: plans, approvals, questions, live desktop and browser panels now share one design across all four themes, risk badges show the real level, and a locked session can no longer be toggled by a stray click.
- Agent activity groups: plan-step and sub-agent sessions now sit in an expandable “Agent activity” group under their conversation in the Dashboard sidebar and update live, failed or interrupted plan-step attempts are archived out of the group (kept for audit), and deleting the conversation deletes its activity too.
- Sturdier storage and diagnostics: the SQLite database now uses write-ahead logging with a busy timeout so reads no longer block behind writes, each conversation event gets a unique sequence number even when events are written at the same time, and the Dashboard’s runtime, log and doctor pages no longer fail on Windows while the Gateway writes its log.
- Metrics for operators: Graph run counts (started, completed, paused and resumed, plus replans and budget stops) and, when a text-to-speech provider speaks a reply, time to first audio and per-provider text-to-speech latency and failures are exported on the Gateway’s authenticated Prometheus metrics endpoint.
Developer tools
- Scriptable CLI: global flags such as --json, --format and --gateway-url can now follow any subcommand, sophon dev validate --json and sophon dev build --json each print a single JSON document including the built package’s exact file list, and failing sophon dev commands now exit with a non-zero code.
- Safer packages: sophon dev build and sophon dev install no longer include .env files (including .env.example), a .venv folder, or editor folders such as .vscode and .idea, and packages built on Windows now install on macOS and Linux.
- Working skill templates and Marketplace pre-checks: sophon dev new skill now creates Python and C# skills that run in the sandbox without changes and come with tests, and sophon dev validate warns before you upload about skill manifest problems the Marketplace’s publishing rules flag, such as a missing author, license, README or LICENSE file.
Upgrade notes
- Back up before upgrading: the first start of the v2.0.0 Gateway applies five database migrations automatically, and there is no automatic downgrade path.
- Sophon Graph is on by default: to keep the previous step-by-step engine, choose Classic per conversation, as your default in Settings → Execution or with /mode in the CLI, or have an admin turn Graph off for the workspace.
- The CLI opens full screen by default: set SOPHON_CLI_FULLSCREEN=0, or "fullScreen": false in cli.json, to keep the pinned-input chat.
- Prompt-injection detection now defaults to Warn, which shows a notice card in the conversation; installs that had turned detection off stay off.
- Spending budgets may trigger where they did not before, because usage of newly catalogued models is no longer recorded at $0.
- Skills that declare the deprecated “proxy” credential mode now receive the stored credential inside their sandbox, the same as “injected”.
- Node voice routes return 404 until an operator turns them on, and once on, the wake route requires the node’s voice permission, is rate-limited per node and rejects over-long transcripts.
- Paired devices: the agent can reach only devices owned by the requesting user, launching, closing or focusing an app now asks for approval, the ungated single-command passthrough tool is removed, and desktop command results use camelCase fields with the app-launch parameter named “name”.
- Custom realtime clients: the voice status event now only means a reply finished (it is no longer sent when speech-to-text starts), transcripts carry an utterance id, and approvals now broadcast a resolved event with outcomes approved, rejected, timed out or cancelled. See the realtime API reference.
- APIs: the session list returns only top-level conversations by default (request view=all for the previous flat list), and deciding an approval that is already decided returns the recorded outcome instead of overwriting it, while unrecognized outcome values are rejected.
- Failing sophon dev commands now exit with a non-zero code, so scripts and CI no longer treat a broken project as a success.
v1.15.0
- Sophon Plugins: extend the platform with out-of-process plugins — add new messaging channels and document formats as local processes, hot-started and stopped from Settings → Plugins with health checks and automatic restart (off by default, admin allowlist)
- Memory Graph v2: a force-directed graph explorer with agent colors, kind filters, date windows, and a co-mention overlay — plus Mermaid diagram export from the Dashboard, chat, agent tools, and API
- Memory that maintains itself: scheduled consolidation promotes daily-log facts into long-term memory, superseded facts drop out of default views, and near-duplicate entities can be reviewed and merged from the Dashboard
- Ask-the-user v2: agents ask structured multi-question forms with tappable options on Dashboard, CLI, and Mobile — and numbered-reply questions on messaging channels; timeouts are explicit, never a silently assumed answer
- Plan approval modes: interactive sessions still gate multi-step plans, while heartbeat, cron, webhook, and sub-agent runs proceed automatically — configurable per session or host-wide
- Documents v2: grounded Q&A with citations — ask a single document or your whole library and get answers that cite their sources, or an honest "no answer" when the documents don't contain one
- Rich in-app document viewer for PDFs, spreadsheets, DOCX, images, audio, Markdown, Mermaid, and code — plus save-by-URL ingestion and automatic ingestion of files sent over WhatsApp, Telegram, and Slack
- Version-preserving document updates with per-version download, an async processing pipeline with a durable queue, and OCR for scanned PDFs
- Claude subscription providers: run Sophon on your Claude Pro/Max plan — direct chat (anthropic-subscription), text-only via the Claude Code CLI (claude-code), or full tool-using agent runs with no per-token cost (claude-code-agent)
- Canvas v3: per-frame version history with a stepper, a frame toolbar with width presets, and library documents that open directly on the canvas
- Backup from the Dashboard: create, monitor, and delete full data-directory backups with live progress
- MCP per-user identity: MCP tokens are now bound to a Sophon user, with per-user memory, document, and agent resources — regenerate tokens bound to a real user after upgrading
- Parallel delegation: agent.delegate can fan out 2–5 sub-agent tasks in one call, and its model and budget overrides now apply
v1.14.0
- Sophon Marketplace: browse, search, and install community skills and plugins straight from the Dashboard (/skills) and CLI — every install human-approved with an up-front permission preview
- Verified installs: marketplace downloads are SHA-256-checked, safely extracted, and hot-registered with no restart; plugins unload and update cleanly
- Device linking and reviews: link a Sophon instance to your marketplace account and leave ratings and written reviews from the Dashboard
- Leaner bundle: Jira, Confluence, and Gmail skills now ship via the marketplace instead of being bundled (google → google-gmail)
- Refreshed model catalog: added the Claude 5 family, GPT-5.6 / 5.4, and Gemini 3.x
- Smarter memory: keyword search across stored memories (portable across databases), a vector scope contract, and memory-write for scoped sub-agents
- Agent rebinding: switch the agent on an existing conversation from the Dashboard picker or the CLI /agent command — the session adopts it on resume and switch
- Codex reasoning replay: encrypted provider reasoning is captured and replayed across agentic-loop iterations, with prompt-cache keys for cheaper multi-turn runs
- More accurate usage and cost accounting: cache-token counting across Codex, Claude Code CLI, Bedrock Converse, Gemini implicit cache, and Copilot
v1.11.0
- Heartbeat autonomy: agents run standing checklists on a schedule with noop suppression, failure backoff, and a Medium-risk auto-approve ceiling
- Real-time speech-to-text: streaming server-side transcription with live interim results on Dashboard and Mobile (Deepgram, OpenAI, Azure, Google, ElevenLabs, Sophon Managed Speech)
- Model catalog: capability and pricing matrix for 20+ bundled models, hidden-model controls, and per-session model selection via Dashboard picker or CLI /provider
- Group chat controls: per-channel policies (mention-only, allowlist, all-messages, disabled) with mention, reply, quote, and thread triggers across six platforms
- Thread-aware sessions: Slack threads, Discord forum posts, Telegram topics, and more each get an isolated session with automatic expiry and cleanup
- Inbound message debouncing: rapid-fire messages coalesce into a single agent turn, with per-channel tuning
- Reply and quote context: quoted messages are captured and shown to the agent on Telegram, Discord, Slack, Mattermost, and Matrix
- Media understanding: inbound images, audio, and video are automatically transcribed or captioned so any model can reason over them
- PowerPoint generation: build complete themed .pptx decks in one call — 10 layouts, 6 themes, speaker notes, and deck editing operations
- Mobile Material 3 redesign: Daylight, Indigo Core, and AMOLED themes, Material You dynamic color (Android 12+), edge-to-edge layout, predictive back
- Mobile chat reliability: per-session stream buffers, mid-task reconnection with stream reconciliation, and approval recovery
- CLI v2 by default: pinned input zone, inline slash and @-mention pickers, live subagent status band, and four terminal themes
- New CLI chat commands: /provider (per-session model override), /thinking, /auto-approve, /canvas — thinking and auto-approve settings persist across sessions
- Session-scoped realtime events: per-session badges and approvals across Dashboard, Mobile, CLI, and Desktop — no more cross-device noise
- Hardened multi-tenant realtime isolation: all live events are tenant-scoped on shared gateways
- Production hardening: automatic credential migration into the vault, channel config rollback on failure, SSRF protections, and configurable rate limits
v1.0.0
- Initial release of Sophon
- Multi-agent runtime with LLM-based routing
- 12+ channel adapters (WhatsApp, Telegram, Slack, Discord, etc.)
- Visual workflow builder with 20+ node types
- Intelligent memory system with agent-scoped isolation
- Document processing pipeline (PDF, DOCX, XLSX, images, audio)
- Secure code sandbox (Docker + gVisor, process fallback)
- Credential vault with OAuth 2.1 + PKCE
- Human approval gates with risk classification
- React 19 Dashboard with 18+ pages
- Electron desktop app (Windows, macOS, Linux)
- Three tiers: Personal (free), Pro, Enterprise
- Enterprise features: SSO, RBAC, audit logs, multi-tenancy
v0.9.0-beta
- Agent-scoped memory with database storage
- Channel file attachment support
- Process sandbox fallback for non-Docker environments
- Workflow expression engine with recursive descent parser
- Dynamic service connections with manifest-driven configuration
v0.8.0-alpha
- Core agent runtime and session management
- Basic channel adapters (WebChat, Telegram)
- Memory engine with short-term and long-term storage
- Skill engine with bundled skills
- SQLite database for Personal tier