Sophon Node
One run. One approval. One machine.
Sophon Node is a background service on your desktop holding one outbound connection to your Gateway, with no inbound ports and no reasoning on the machine. The agent sends an ordered batch of actions; the run pins to a single node, stops at the first failure, and ends with a fresh screenshot so the model can check its own work.
The unit of work
A run, not a click.
One call carries an ordered list of actions and executes them as a single unit. One approval covers the batch, at the highest risk in it, and the run halts at the first failure, with everything after it marked Not executed.
One approval, at the batch’s highest risk
Approve a desktop run in chat and the consent lasts 30 minutes, scoped to that one conversation, held in memory and cleared by a Gateway restart. Critical is never covered, so a shell command asks again every time. Workflows, scheduled runs and MCP calls get no such cover at all: they ask on every action.
Over 30 command types, one ordered list
A batch draws on the screen, input, window, app, clipboard, file and system commands the node implements. There are no per-run budgets and no stable element references: a run is bounded by its action list and by you stopping it.
The passthrough is gone
The old node.command tool was deleted, not deprecated. It let an action reach the desktop around the approval gate and the shell denylist, so it no longer exists. Every desktop action now arrives as a Control Run.
Sight
The model sees the screen it clicks.
Screenshots used to stop at the chat. On Claude models reached through the Anthropic API or a connected Claude subscription they now travel into the model’s own context, carrying the geometry that turns a click on a downscaled image into a real pixel on your display.
Captures that carry their geometry
A capture ships the origin and size of the screen region it came from, so a coordinate read off a downscaled image maps back to the real display instead of landing two hundred pixels away. That frame is emitted on Windows; on macOS and Linux the model gets the picture without it.
Grounded in the accessibility tree
Element queries run on real UI Automation on Windows, capped at 300 elements, with password fields dropped at the source before anything leaves the machine. When UI Automation cannot answer, it falls back to the window APIs.
The picture reaches the model
On Claude models reached through the Anthropic API or a connected Claude subscription, a capture travels on its own image channel into the model’s context instead of stopping at the chat transcript. That is the difference between an agent that is told what happened and one that can look. On every other provider the run still works, but the model gets the text result and not the image.
Fully supported. Native input and capture, per-monitor DPI awareness, element-level grounding, and physical screen coordinates on every capture.
Supported with caveats. Every desktop action works, and needs macOS Accessibility permission. Element queries reach inside windows but are shallower and slower than on Windows.
Supported with caveats, X11 recommended. Missing helper tools surface as actionable errors. Element queries see top-level windows only on X11, and are unavailable on Wayland.
Windows is the only desktop with element-level grounding; on macOS and Linux the agent sees windows, not controls. Releases publish self-contained builds for win-x64, osx-arm64 and linux-x64.
What it knows about your desktop
Nothing ambient. Durable only if you say so.
A context action reports the foreground app, the active window and the focused element. It is a light read rather than a picture, cheap enough to ask for often, and it runs under the screen scope a device already has. The agent has to ask for it: your desktop is not described to the model until it does.
It asks; it is never told
Nothing about your desktop is attached to a turn on its own. When the agent needs to know what is in front of you, it runs a context action like any other action: inside a run, through the same gates, on the record.
Treated as untrusted input
A window title is something an attacker can write, so the result comes back wrapped as untrusted content. Text inside it is data to be read, never an instruction to be followed.
Looking back is a separate grant
Asking what is on screen now and asking what someone did this morning are different questions. There is no rolling history unless the recorder below is switched on for that device, and reading it needs a scope of its own.
A real history, only on a device where you turn it on.
The context action above only ever sees the desktop at the moment it is asked, so “what was I doing before lunch?” had no answer at all. The recorder is the missing half: a background sampler notes the foreground app and window every 20 seconds whether or not anything is talking to the agent, and keeps 30 days of it. Consecutive samples of the same window collapse into one row with a duration, so a day at a desk is a handful of rows rather than thousands of ticks.
- It does not start on its own. Recording is off until you enable it for a specific device, and it can be paused without losing its configuration.
- The history stays on the machine. It is written to a local database on the device itself. The Gateway holds no copy.
- Password managers and private windows are excluded before the write. 1Password, Bitwarden, KeePass and similar apps, and titles containing incognito, inprivate, private browsing, are dropped at sample time, so an excluded window’s title never reaches disk rather than being filtered out when something reads it back. The lists are yours to extend.
- Screenshots are a second opt-in. Turning recording on does not turn them on. Saved separately and off by default, one image is written when a new entry starts, because a window title is metadata and a screenshot is whatever happened to be on the screen. They are kept for 7 days rather than the history’s 30, under a 2 GB disk ceiling that deletes the oldest first, as ordinary image files you can open and delete yourself. Viewing one puts it in front of the model on Claude models reached through the Anthropic API or a connected Claude subscription; elsewhere the agent gets the entry, not the image.
- Reading it is a separate grant. The agent needs a scope of its own, not granted by default. Being allowed to drive a desktop is a different decision from being allowed to read a record of the day.
- It is all in the Dashboard, and none of it needs a restart. Turn recording on or off, pause it, change the interval and retention, edit both exclusion lists, switch screenshots on, erase the history: the panel shows what the recorder is actually doing, and a change takes effect on the machine without restarting the node.
- Browse it, and delete one row of it. The device page carries the timeline itself, grouped by day, filterable, with any screenshots inline and a breakdown of where the time went. Deleting a single entry deletes its image with it.
- Queries are bounded. The agent asks for a period and gets that slice, never the whole history. When recording is off it is told so plainly, so an empty answer is never reported back to you as “you did nothing”.
Before it acts
Four gates before anything touches your desktop.
Every command runs the same gauntlet, in the same order. Fail any one and it never runs.
Risk decides who gets asked
Low-risk reads run inside the scope you granted. Launching, closing or focusing an app, and writing a file, each need an approval. Shell execution is Critical on its own: an editable preview of the exact command, a hardcoded denylist of destructive patterns, and an approval that a run-level consent can never pre-cover.
Ownership is checked where the work is queued
A node belongs to one owner, and that check lives inside the path that queues the command rather than only at the edge. Ask for a machine that is not yours and the answer is the same one you get for a machine that does not exist. The tenant is captured when the turn starts, not when the command lands.
Every outcome on the record
Accepted, refused by scope, rate-limited, denied, completed, failed, cancelled. Every outcome streams into the Gateway’s audit log, and every action shows up inline in the Dashboard chat and on the Devices page.
Scoped per device
Grant exactly what you mean.
Approving a device grants 8 scopes. The ones that reach past the screen are not among them. A Node without a scope can be asked for it all day, and the Gateway refuses before the command ever leaves the server.
Granted when you approve a device
Off until you turn them on
A further 4 scope names are recognised by the permission gate but have no commands behind them on this release. They are placeholders, not features waiting to be switched on.
Files, inside folders you name
File access is confined to the roots you configure: Desktop, Documents, Downloads by default. An empty roots list denies everything, not the reverse. Reads and writes cap at 512 KB, a listing caps at 500 entries, binary reads are refused, and symlinks are resolved before containment is tested, so a link pointing out of a root fails the check instead of escaping it. There is no file transfer between machines: reads and writes happen on the node, in place.
How file access is boundedWhile it runs
Watch it work. Stop it mid-run.
A live panel sits in the Dashboard chat while a run is in flight: the action list, the step it is on, the risk on that step, and the screenshot it just took.
Stop means both things
Stop cancels the in-flight node command and the agent turn behind it. Not a pause, and not a request the model gets to reconsider.
Heartbeats, and absence
A node heartbeats every 30 seconds and is treated as offline after 2 minutes. A run fails fast by default: ask a machine that is not connected and you get that answer straight back, rather than a timeout half a minute later.
No local kill switch
Stopping happens from the Dashboard or the CLI. There is no hotkey on the machine itself and no tray icon. That is a gap we would rather name than let you assume away.
Where it actually is
What we have not shipped.
Several things people reasonably assume arrived with all of the above did not. Here is the list, before you decide what to hand it.
Node binaries are not code-signed. Windows and macOS will treat them as unsigned software, and you should check what you downloaded before you run it.
There is no installer package, no tray icon, and no doctor, install or uninstall subcommand. The node CLI is 5 verbs: pair, start, status, unpair, version.
A node verifies a release manifest’s signature before it trusts it, and refuses one it cannot verify, but updates are installed by hand. Nothing swaps a binary for you.
Element-level grounding is real on Windows only. On macOS and Linux an agent works from screenshots and coordinates.
Element references are not stable between turns. The model re-queries the screen; it does not hold a handle to a button.
Per-turn desktop awareness has not shipped. Nothing tells the agent what is on your screen at the start of a turn; it has to ask, with a context action inside a run. The rolling in-session timeline that would be built from those automatic readings is empty on this release, which is why a durable answer means switching the recorder on.
There is no file transfer between machines, and no per-run cost or time budget. A run ends when its action list ends, when a step fails, or when you stop it.
A visible AI-operation disclosure surface and an exportable per-run replay bundle are planned. Neither has shipped.
Paired in minutes
Three steps and a terminal.
The Dashboard hands you the full pairing command. The node CLI is 5 verbs in total: pair, start, status, unpair and version. It runs as a service under your own user account, because a system-level service has no desktop session to drive.
Give it a desktop.
Pair a machine, keep the default scopes, and watch the first Control Run land, with one approval, one node, and a screenshot at the end.