Platform Support & Limits
What each operating system supports today, what this release does not include, and what is planned.
Every command in this section is implemented on all three operating systems. There are no stubs. What differs is how well the underlying platform lets a node see a screen, and that difference is large enough to plan around.
This page collects every limitation in one place, so the rest of the Node documentation does not have to hedge.
Platform parity
| Capability | Windows | macOS | Linux |
|---|---|---|---|
| Screen capture | Yes | Yes | Yes (X11 and Wayland) |
| Mouse and keyboard | Yes | Yes, needs a helper tool | Yes, needs a helper tool |
| Window management | Yes | Yes | Yes |
| Application management | Yes | Yes | Yes |
| Clipboard | Yes | Yes | Yes |
| Files | Yes | Yes | Yes |
| Notifications | Yes | Yes | Yes |
| Canvas | Yes | Yes | Yes |
| Shell execution | Yes | Yes | Yes |
| Element queries | Full, inside windows | Inside windows, shallower | Top-level windows only on X11; unavailable on Wayland |
| Capture coordinates | Yes | No | No |
The last two rows are the ones that decide how you write automation.
Element queries are how an agent finds a control rather than guessing at pixels, and the three platforms are not close to each other here.
- Windows runs the query against UI Automation, the platform's real accessibility tree, so it reaches individual controls inside the foreground window — or every window, when the call asks for background ones too. It is capped at 300 elements, and a query that hits the cap comes back marked as truncated, with advice to narrow it by name, role, or depth rather than silently returning a partial tree as if it were whole. Password fields are dropped at the source, so their contents never leave the machine. If UI Automation cannot be used at all, the query falls back to enumerating window handles, which still answers "which windows are open" instead of failing the run.
- macOS walks the front window's element tree and needs macOS Accessibility permission; without it the OS refuses the query and the node says so, naming the setting to open. With it, the tree is shallower and slower than the Windows one, and background windows are not covered.
- Linux is X11 only, where the query enumerates top-level windows with their geometry, using the helper tools listed below. Under Wayland it returns nothing at all, with a warning saying why, so an agent must fall back to screenshots and coordinates.
Capture coordinates are the geometry that lets a click on a downscaled screenshot land on the right pixel. Windows captures carry the origin and size of the region they came from; macOS and Linux captures do not, so a model working from them has the picture without the frame.
Release status
Windows is generally available. Screen, input, window, application, clipboard, file, and element-level control are all supported and covered by the release test suite.
macOS and Linux are preview. The node runs, pairs, and executes commands, but automation is window-level: an agent can move, focus, and capture windows, and click and type at coordinates, but it cannot address individual UI elements the way it can on Windows. Plan macOS and Linux automation around screenshots and coordinates, not element names.
Host requirements
- Windows. Nothing extra. Input, capture, and element queries use the platform's own APIs.
- macOS. Grant Accessibility permission to the node, or input and element queries will be refused by the OS. Some input requires a helper utility to be installed.
- Linux. X11 is recommended. Helper utilities are needed for input, capture, and clipboard, and they differ between X11 and Wayland. A missing tool produces an actionable error naming what to install, and the node reports which tools it found in its heartbeat.
Not in this release
Each of these is a real gap. Where there is a way to achieve the same end today, it is named.
| Capability | Status | What to do instead |
|---|---|---|
| Per-turn desktop awareness | Not shipped | Nothing tells the agent what is on your screen at the start of a turn. The agent asks for desktop context when it needs it, and the activity recorder is what answers questions about earlier in the day. |
| Screenshots the model can see on every provider | Not shipped | An image reaches the model itself only on Claude models, through the Anthropic API or a connected Claude subscription. On other provider paths the model gets the text result and cannot see the picture. |
| Node lifecycle webhooks | Not shipped | There is no node.online or node.offline event to subscribe to. Watch the Devices page, or poll sophon nodes list --json. |
| Signed binaries | Not shipped | Node binaries are not code-signed; Windows and macOS will flag them. Verify what you downloaded before running it. |
| Installer package | Not shipped | Install as a global tool, or take a self-contained build from /downloads/node. |
| Verified update delivery | Not shipped end to end | A node verifies a release manifest's signature and refuses one it cannot verify, but manifests are provisioned by an operator and updates are installed by hand. Nothing swaps a binary for you. |
| Tray application | Not shipped | Manage the node with sophon-node status and your service manager. |
| Local kill switch | Not shipped | Nothing on the machine itself stops a run. Press Stop in the Dashboard chat (Mobile and the CLI can watch a node command but cannot cancel one), remove the device's scopes, revoke the device, or stop the service. |
| Presence and takeover signalling | Not shipped | There is no on-screen indicator that an agent is operating the machine, and no instant local-input takeover. |
| Stable element references | Not shipped | An agent re-queries the screen each time it looks; it does not hold a handle to a control across actions. |
| File transfer | Not shipped | File access reads and writes in place, on the machine. |
| Per-run budgets | Not shipped | A run ends when its actions end, when one fails, or when you stop it. |
Taken together these mean a node is best given to a machine you own and supervise. The gates are real and the audit trail is complete, but there is no indicator on the machine itself telling a bystander that an agent is driving it.
Planned hardening
Stated as intent, not as behaviour you can rely on today:
- A visible AI-operation disclosure surface and an exportable per-run replay bundle, for reconstructing exactly what a run did. Both are planned; neither has shipped.
- OS keyring storage for the node token, replacing today's plain-text config file with the platform credential store on each OS.
- Native accessibility parity on macOS and Linux, which is what would move those platforms out of preview.
- Wider vision coverage on tool results, so a captured screen reaches the model on every provider that can accept one.
- Ambient desktop context each turn, so an agent starts a turn already knowing which application is in front of you instead of having to ask for it.
Where to go next
- Install & Pair a Node — prerequisites per platform
- Running as a Service — why the node runs as your own user
- Control Runs — what a run does with what it can see
- Commands & Actions — the full command surface