Sophon 2.0 is here
Sophon Docs
API Reference

SignalR

Real-time hubs — chat streaming, task status, approvals, voice, canvas, terminal, and node commands.

SignalR is Sophon's real-time transport. The Dashboard, Mobile, Desktop, and Sophon Node all connect to SignalR hubs for streaming, live status, and bidirectional commands. This page is the protocol reference.

Hubs

Sophon maps four hubs under /hubs/*:

HubPurpose
/hubs/chatChat streaming, task lifecycle, tool progress, session updates, approvals and questions, voice, and discussion runs
/hubs/canvasCanvas frames and canvas actions
/hubs/terminalInteractive terminal sessions
/hubs/nodeSophon Node command dispatch + heartbeats

There is no separate tasks, approvals, discussion or notifications hub. Task lifecycle, approval and question events, voice, and discussion runs all ride /hubs/chat — a custom client only needs that one connection unless it drives the canvas, a terminal or a node.

A single client can connect to multiple hubs over one WebSocket (SignalR multiplexes).

Authentication

Same JWT bearer as REST. Pass via query string (SignalR's WebSocket limitation):

import { HubConnectionBuilder } from "@microsoft/signalr";

const connection = new HubConnectionBuilder()
  .withUrl("https://gw.example.com/hubs/chat", { accessTokenFactory: () => jwt })
  .withAutomaticReconnect()
  .build();

Chat hub

Hub methods take positional arguments — SignalR matches by arity, so pass them in order, not as one object:

// SendMessage(sessionId, text, agentId, thinking?)
await connection.invoke(
  "SendMessage",
  "sess_abc",
  "What's on my calendar today?",
  "default"
);

// SendMessageWithAttachments(sessionId, text, agentId, documentIds, thinking?)
await connection.invoke(
  "SendMessageWithAttachments",
  "sess_abc",
  "Summarize this",
  "default",
  ["doc_1"]
);

Server → client events. As of 1.11.0, task stream, status, lifecycle, and approval events are delivered to the session group rather than the whole user — see Session scoping.

EventPayloadMeaning
TaskQueued{taskId, sessionId, policyMode, pendingTaskCount, replacedActiveTask, droppedPendingTaskCount}Message accepted, background task queued (sent to the caller)
TaskRejected{taskId, sessionId, errorCode, error, policyMode, pendingTaskCount}The session's execution policy refused the message (sent to the caller)
TaskStarted{taskId, sessionId}Worker picked up the task
AgentStatus{taskId, sessionId, phase, description, isComplete, planMetadata}Progress checkpoint for the running turn
StreamChunk{taskId, sessionId, content, isComplete, modelId, providerId, providerName, inputTokens, outputTokens}Token stream (if streaming enabled)
ToolProgress{taskId, sessionId, toolName, eventType, summary, data}Live progress from a running tool (e.g. browser screenshots)
ReceiveMessage{id, sender, content, timestamp, sessionId, channelType, proactive}Final agent message
TaskCompleted{taskId, sessionId, modelId, providerId, providerName}Task done
TaskFailed{taskId, sessionId, error}Task errored
TaskCancelled{taskId, sessionId, success}Task cancelled
SessionUpdated{id, …}Session metadata changed — carries the session id plus whichever field changed (e.g. status, agentId)
SessionNotification{sessionId, kind, refId, success, excerpt, timestamp}User-wide activity event (see SessionNotification)
ApprovalRequired{id, toolName, description, risk, sessionId, agentId, channelType, parameters, previewContent, allowEdit, timeoutSeconds, escalationSessionId, graphRunId, graphNodeId}An approval is waiting for a decision
ApprovalResolved{id, outcome, sessionId, escalationSessionId}That approval is no longer pending (see Approvals and questions)
InfoRequestRequired{id, question, options, questions, context, sessionId, timeoutSeconds, createdAt, escalationSessionId, graphRunId, graphNodeId}The agent is asking a question
InfoRequestAnswered{id, answer, timedOut?, cancelled?, escalationSessionId, graphRunId, graphNodeId}That question is settled
VoiceUtteranceStarted{sessionId, utteranceId}A new speech-to-text utterance is open — sent only to the connection streaming the audio
VoiceInterimTranscript{sessionId, utteranceId, text}Live partial transcript — sent only to the connection streaming the audio
VoiceFinalTranscript{sessionId, utteranceId, text, confidence}Committed utterance — sent only to the connection streaming the audio

Voice replies (audio, status, fallback notices) go to the session's voice group instead, so every attached tab hears them — see Voice.

The hub carries more than the table above: session-list events (SessionCreated, SessionDeleted, SessionTitleUpdated, ChildSessionCreated), BudgetAlert, DocumentProgress, WorkflowProgress, the ClaudeCode stream events, and the discussion events below.

Client → server methods:

MethodPurpose
SendMessageSubmit a user message
SendMessageWithAttachmentsSubmit a user message with document ids attached
CancelTaskCancel in-flight task
CancelPlanCancel an active plan
JoinSessionSubscribe to a session's events — the primary subscription mechanism (see Session scoping)
LeaveSessionStop receiving events for a session
RespondToApproval(approvalId, approved) — decide a pending approval
RespondToApprovalWithEdit(approvalId, editedPlanJson) — approve with modifications
RespondToInfoRequestV2(infoRequestId, answers) — answer a question; answers maps each wire question id (q1…qN) to the list of chosen values
VoiceStart(sessionId?, agentId) — open a voice session on a chat session; returns the session id it settled on (a new one if you passed none, or one you don't own)
VoiceMessage(sessionId, text) — submit an already-transcribed utterance as a voice turn
StreamAudioStream microphone audio for realtime transcription (see StreamAudio)
VoiceInterrupt(sessionId) — stop the current speech-to-text stream and silence the turn's reply. It does not cancel the task: the work keeps running and its reply still lands in the chat thread
VoiceEnd(sessionId) — end the voice session and leave its voice group

Session scoping

As of 1.11.0, JoinSession(sessionId) is the primary subscription mechanism, and per-task events target the session group instead of broadcasting user-wide:

  • Clients are auto-joined to sessions they create — sending the first message on a new session subscribes you implicitly.
  • Ownership is validated on every join — a client can only join sessions owned by its user (and tenant). The Canvas (whiteboard) hub validates session ownership on join the same way.
  • Legacy clients that never call JoinSession fall back to user-wide events, so older integrations keep working unchanged.

This is what keeps multiple devices from interfering: a Dashboard tab watching one session never receives another session's stream chunks, and approvals shown in the CLI aren't disturbed by activity on other devices. Devices not joined to a session still learn about its activity via SessionNotification.

SessionNotification

SessionNotification is a lightweight, user-wide activity event — every connected device receives it, whether or not it has joined the session:

FieldMeaning
sessionIdThe session with activity
kindWhat happened — one of the seven kinds below
refIdID of the thing the event refers to (task, approval, or info request)
successOutcome flag, where the kind has one (e.g. false on taskFailed)
excerptShort content excerpt, when one is available
timestampWhen the event occurred
KindFires when
taskStartedA task begins running
taskCompletedA task finishes successfully
taskFailedA task errors
taskCancelledA task is cancelled
approvalPendingAn approval is waiting for a decision
infoRequestPendingThe agent is waiting on a user answer
heartbeatActionedA heartbeat run took action — see Heartbeat

The excerpt gives clients a short preview to render; full message content and detail still require JoinSession or REST. Clients use the event to update unread badges and pending-approval counters. The Desktop app's native notifications are driven by this event.

Voice

A voice session is opened with VoiceStart(sessionId?, agentId). It attaches your connection to the chat session's group and to its voice group, then returns the chat session id it settled on. Several connections can attach to the same voice session, which is why replies reach every open tab and survive a reconnect (reconnecting mints a new connection id, so the client re-invokes VoiceStart).

Server → client voice events. Transcript events are addressed to the one connection streaming the audio; everything else goes to the session's voice group:

EventPayloadMeaning
VoiceUtteranceStarted{sessionId, utteranceId}A new utterance is open (connection-addressed)
VoiceInterimTranscript{sessionId, utteranceId, text}Partial transcript; each interim replaces the previous one (connection-addressed)
VoiceFinalTranscript{sessionId, utteranceId, text, confidence}Committed utterance (connection-addressed)
VoiceStatus{sessionId, status}Session state: listening, thinking, speaking, idle, expired
VoiceAudioChunk{sessionId, audioBase64, chunkIndex, isLast}Synthesized reply audio, chunk by chunk
VoiceNativeTtsRequested{sessionId, text, language}No provider audio for this reply — speak it with the client's own speech synthesis
VoiceProviderFallback{sessionId, reason}Why provider speech isn't being used for this reply
VoiceTranscript{sessionId, text, isAgent}Transcript line for the voice UI — the user's own words, or the agent's reply
VoiceError{sessionId, error}Voice or transcription failure

Two VoiceStatus rules changed in 2.0 and matter to custom clients:

  • VoiceStatus is no longer sent when a speech-to-text stream opens. The producer was removed. A client that treated VoiceStatus{listening} from the speech-to-text path as "the mic is open" must stop; open the mic on its own action instead.
  • A VoiceStatus of listening after a turn means the reply is finished — that is the signal to re-arm the mic, together with playback having finished on the client side.

StreamAudio

StreamAudio(sessionId, audioChunks, mode) is a client-to-server streaming method for realtime speech-to-text. SignalR matches by positional arity, so all three parameters are always sent:

import { Subject } from "@microsoft/signalr";

const audioChunks = new Subject<string>();
await connection.send("StreamAudio", sessionId, audioChunks, "auto");

// feed base64-encoded PCM chunks as the mic produces them
audioChunks.next(base64Chunk);
// manual mode: end the utterance explicitly
audioChunks.complete();
  • Audio format — 16 kHz mono linear16 PCM, base64-encoded per chunk. This works over the standard SignalR JSON protocol; no binary transport is required.
  • mode: "auto" (the default for any unrecognised value) — hands-free. The end of the utterance is detected for you: natively where the provider supports it, otherwise by a server-side silence detector, so hands-free works with any active provider.
  • mode: "manual" — push-to-talk; the client ends the stream by completing it.
  • mode: "answer" (new in 2.0) — the server transcribes and returns the final transcript, but runs no agent turn and does not resolve the pending prompt. The client matches the transcript to an answer itself and then calls RespondToApproval or RespondToInfoRequestV2. Use it for a capture opened to answer an approval or a question, so answering never spawns a chat turn.

Mode matching is case-insensitive. The arity is unchanged, so a client that only ever sent "auto" or "manual" keeps working untouched.

Utterance ids

Every transcript now carries an utteranceId, announced first by VoiceUtteranceStarted. A client should accept a final transcript only for the most recently announced utterance and drop the rest — that is what keeps a slow transcription of something you said earlier from being taken as your answer to the next question.

This is advisory, not enforced: a client that ignores VoiceUtteranceStarted and forwards every final keeps working exactly as before. It also only applies to server-side speech-to-text — a client using its own browser recognition never sees an announcement, so it has nothing to match against.

In auto and manual modes the committed utterance enters the agent pipeline like a typed message; in answer mode it does not. See Voice for provider setup and per-user listening preferences.

Approvals and questions

Approval and question events ride this hub. They are sent to the request's session group, to the escalation root's session group when the request was raised in a child session, and to the user group when the request has no session.

ApprovalResolved tells every surface an approval is no longer pending — decided on a chat card, over REST, from a channel reply, by voice, in another tab, by timeout, or by a cancelled run — so an open card or a waiting voice prompt stops asking:

FieldMeaning
idThe approval
outcomeapproved | rejected | timed_out | cancelled
sessionIdThe session the approval belongs to
escalationSessionIdThe escalation root's session, when the approval was raised in a child session

cancelled is not a decision: it is written when the caller's own wait ends (the chat's Cancel, a cancelled plan or step) or when a graph run is cancelled while one of its approval gates is still pending. Before 2.0 that case was recorded and announced as timed_out. An approval that ends without a decision is still never treated as an approval.

InfoRequestAnswered gained the matching field: a cancelled question now carries cancelled: true with timedOut: false, so a client can tell a cancel from a real expiry. The caller-facing answer shape is unchanged.

Node hub (/hubs/node)

Bidirectional command channel between the Gateway and Sophon Nodes. Nodes connect as clients.

Gateway → Node (methods Nodes implement):

MethodPayload
ExecuteCommand{commandId, command, params} — Node runs it, returns result
RefreshScopes{scopes} — update permission scopes without reconnect
RevokeAccess{} — force disconnect

Node → Gateway:

MethodPayload
Heartbeat{nodeId, status, uptime} — every 30 s
CommandResult{commandId, ok, result, error?} — response to ExecuteCommand
ScopeChanged{scopes} — ack of RefreshScopes

See Sophon Node for command semantics and Node Protocol for field schemas.

Discussion events

Discussion runs have no hub of their own — their events arrive on /hubs/chat, addressed to the user group rather than a session group, so clients filter by runId.

Server → client:

EventPayload
DiscussionStarted{runId, definitionId, definitionName, rounds, seedPrompt, startedAt, panelists, judge, …}
DiscussionPhaseChanged{runId, phase}
DiscussionRoundStarted{runId, round, totalRounds}
DiscussionTurnStarted{runId, turnId, round, participantId, agentId, side, displayName, startedAt}
DiscussionTurnStreamChunk{runId, turnId, participantId, agentId, side, …} — live tokens for one turn
DiscussionTurnToolProgressTool progress inside one turn
DiscussionTurnCompleted{runId, turn}
DiscussionJudgeStarted{runId, turnId, participantId, agentId, side, displayName, startedAt}
DiscussionVerdict{runId, verdict}
DiscussionInterjectionAppliedAn interjection was folded into the run
DiscussionCostUpdated{runId, accruedUsd, capUsd}
DiscussionCompleted{runId, status, completedAt, skippedPanel, timedOut, verdict}
DiscussionFailed{runId, error, completedAt, timedOut}
DiscussionCancelled{runId, reason, completedAt}

Client → server: CancelDiscussion(runId) on the chat hub. Starting and listing runs is REST — see REST API.

Groups

SignalR groups are the mechanism for user / tenant / session scoping. As of 1.11.0, all group names are tenant-qualified — session groups are scoped to the tenant that owns the session, and user groups are prefixed with the tenant ID. A shared Gateway hosting multiple tenants can never deliver one tenant's events to another tenant's connections. See Tenants & Multi-Tenancy.

GroupScope
User groupUser-wide events for one user in one tenant (SessionNotification, session list updates, discussion run events)
Session groupEvery event in one chat session — joined via JoinSession, ownership-validated
Voice groupVoice delivery for one chat session — joined by VoiceStart, left by VoiceEnd; this is why every attached tab hears the reply
Canvas groupCanvas frames for one session, on the canvas hub
Node groupCommand / heartbeat events for one node, on the node hub

Clients are auto-joined to their user group on connect and to session groups for sessions they create. Group names are server-internal — clients only ever pass IDs (sessionId, runId) to subscribe methods.

Reconnection

SignalR's withAutomaticReconnect() handles network blips. Reconnection retries with exponential backoff (0 ms, 2 s, 10 s, 30 s, then every 30 s).

On reconnect:

  • Subscriptions are re-established (groups rejoined)
  • Missed events are not replayed — use REST to catch up (e.g., GET /tasks/{id} for task state)
  • Streaming tasks that were in flight continue; the client sees a possibly-incomplete stream followed by TaskCompleted once the task finishes

Rejoining mid-task: after JoinSession, fetch the persisted message (GET /sessions/{id}/messages) and replace any partial stream buffer with it — appending renders duplicated text. Guard for staleness: the probe may discover the task already finished while you were reconnecting; in that case render the final message and don't re-arm any waiting or streaming state. The Dashboard and Mobile apps do this automatically.

Backpressure

Server buffers outgoing events per-connection (default: 100 events). Beyond that, the slowest consumer is dropped. This prevents a slow client from OOMing the Gateway.

The mobile app and Dashboard rarely hit this. Custom SignalR clients that can't keep up should either fan out to a queue or switch to polling REST.

Transport selection

Preference order:

  1. WebSockets — default; lowest latency, bidirectional
  2. Server-Sent Events — for browsers that block WebSockets
  3. Long polling — fallback for restrictive proxies

The Gateway also exposes a node-token command queue over plain HTTP (GET /api/nodes/me/commands), but the shipped Sophon Node connects over the hub — it does not fall back to polling.

Client examples

TypeScript (browser)

import { HubConnectionBuilder } from "@microsoft/signalr";

const conn = new HubConnectionBuilder()
  .withUrl("https://gw.example.com/hubs/chat", { accessTokenFactory: () => jwt })
  .withAutomaticReconnect()
  .build();

conn.on("StreamChunk", (c) => process.stdout.write(c.delta));
conn.on("TaskCompleted", (c) => console.log("\nDone in", c.durationMs, "ms"));

await conn.start();
await conn.invoke("SendMessage", sessionId, "hi", "default");

.NET / C#

var conn = new HubConnectionBuilder()
  .WithUrl("https://gw.example.com/hubs/chat", o => o.AccessTokenProvider = () => Task.FromResult(jwt))
  .WithAutomaticReconnect()
  .Build();

conn.On<StreamChunkEvent>("StreamChunk", c => Console.Write(c.Delta));
await conn.StartAsync();
await conn.InvokeAsync("SendMessage", sessionId, "hi", "default");

Python (signalrcore)

from signalrcore.hub_connection_builder import HubConnectionBuilder

hub = (HubConnectionBuilder()
  .with_url(f"https://gw.example.com/hubs/chat?access_token={jwt}")
  .with_automatic_reconnect()
  .build())

hub.on("StreamChunk", lambda msg: print(msg[0]["delta"], end=""))
hub.start()
hub.send("SendMessage", [sid, "hi", "default"])

Where to go next

  • REST API — everything SignalR doesn't stream
  • Node Protocol — the Node-specific hub contract
  • Webhooks — for consumers that prefer push over WebSockets
  • Voice — speech provider configuration and listening preferences behind StreamAudio