Planning
How Sophon decomposes complex requests into a DAG of steps and runs them on Sophon Graph — in parallel, with checkpoints, retries, replanning, and cancellation.
When you ask Sophon something complex — "research three competitors, build a comparison spreadsheet, and email it to my team" — the agent doesn't just start typing. It plans: it breaks the work into steps, figures out which steps depend on which, and executes independent steps in parallel. If a step fails, it either retries with an alternative approach or cascades cleanly so other steps still run.
Planning is middleware #7 in the orchestration pipeline. Since v2.0.0, approved plans run on Sophon Graph by default.
When planning kicks in
Four modes, set per agent or per request:
| Mode | Behavior |
|---|---|
Auto (default) | Heuristic classifier scores the message; score ≥ 3 triggers planning |
LlmDecides | Fast-tier LLM is asked "is this complex?" (3-token answer, temp 0). Falls back to heuristic on error |
Always | Every request gets planned |
Never | Planning disabled; always flat agentic loop |
This Auto decides whether to make a plan at all. It is unrelated to the execution mode that decides how an approved plan runs — that one has exactly two values, Graph and Classic.
The heuristic
The heuristic classifier adds signal scores:
| Signal | Points |
|---|---|
| Message length > 100 chars | +1 |
| Contains conjunctions ("and then", "after that", "first", "then") | +2 (once) |
| Contains multi-step phrases ("steps", "plan", "sequence") | +2 (once) |
| Contains ≥ 2 action verbs (search, create, send, write, …) | +1 |
Score ≥ 3 means "complex, plan it."
Planning is skipped if the pipeline is already three levels deep. This prevents infinite plan-nesting when a plan step itself triggers planning.
What a plan looks like
A plan is a goal plus a DAG of steps. Each step carries an id, a description, the step ids it depends on, suggested tools, an estimated risk level, and optionally a model capability hint:
{
"id": "plan_abc123",
"goal": "Compare three competitors and email the team",
"steps": [
{ "id": "s1", "description": "Research Acme Corp",
"dependsOn": [], "toolHints": ["web.search", "browser.navigate"] },
{ "id": "s2", "description": "Research Globex",
"dependsOn": [], "toolHints": ["web.search", "browser.navigate"] },
{ "id": "s3", "description": "Research Initech",
"dependsOn": [], "toolHints": ["web.search", "browser.navigate"] },
{ "id": "s4", "description": "Build comparison spreadsheet",
"dependsOn": ["s1", "s2", "s3"], "toolHints": ["document.create"] },
{ "id": "s5", "description": "Email the team",
"dependsOn": ["s4"], "toolHints": ["gmail.send"],
"estimatedRisk": "Medium" }
]
}s1, s2, s3 have no dependencies → they run in parallel. s4 waits for all three research steps. s5 waits for the spreadsheet. And in an interactive session, the whole plan is shown for user approval before anything runs — see Approval below.
Execution
An approved plan is compiled into an immutable, versioned graph and handed to Sophon Graph, the durable execution engine. In each wave, every step whose dependencies are satisfied is dispatched together.
Checkpointing. Every state transition is written to the database as it happens, so a plan's progress outlives the process running it. If the Gateway restarts mid-run, completed steps are never re-run. A step that was in flight is a different matter: chat-plan steps are assumed to have side effects, so with the default recovery mode the run comes back Interrupted and waits for you to press Resume or Cancel on the plan card — Sophon will not replay a half-finished email.send on its own.
Isolated steps. Each step runs in its own child session with a small, clean context instead of one ever-growing transcript. Dependency outputs are wrapped and marked as untrusted before they reach the consuming step, and a step fed external content gets a lower auto-approve ceiling. Step sessions appear under their conversation in the Dashboard sidebar as Agent activity, and each step links to its own session while it runs — steps do not stream their tokens into the main chat.
Retries. On step failure, Sophon asks a fast-tier LLM for an alternative approach and retries. If there's no alternative, the step is marked failed.
Replanning. When a step keeps failing, Sophon revises the failed part of the plan into a new version, and the plan card shows a version badge. A revision that raises the plan's risk, reaches for a new kind of tool, or adds a step with side effects parks the run behind a durable approval gate — that gate survives a restart and waits hours rather than seconds. Frozen, completed steps are never re-gated.
Smart cascade. A failed step doesn't poison the whole plan. Steps that depended on it are marked Skipped, but unrelated steps continue. If the spreadsheet step above failed, the email step would be skipped but nothing upstream would be rolled back.
Budgets. A plan run carries a default cost cap and a token backstop, checked between waves. When a run goes over, no new steps are dispatched and the chat says which limit was hit. The caps are soft by design — steps already in flight are allowed to finish — and a budget-stopped run is reported as failed rather than left resumable.
Cancellation. Cancelling an active plan stops dispatching new steps; in-flight steps observe the cancellation token and stop at their next cancellation point.
Status events
The engine emits events the UI can subscribe to:
plan_start,plan_complete,plan_failed,plan_interrupted,plan_updatedplan_step_start,plan_step_complete,plan_step_failed,plan_step_skipped,plan_step_retry
Clients (Dashboard, Mobile, CLI) subscribe via SignalR and render live progress — you see each step tick from pending → running → done, and a revised plan replaces the step list in place.
Graph or Classic
Execution mode is a choice, resolved per conversation:
| Surface | Scope |
|---|---|
| The execution-mode pill in the Dashboard composer | This conversation |
/mode graph|/mode classic|/mode default in the CLI | This session |
| Settings → Execution | Your personal default |
| Settings → Execution (admin panel) | The workspace default — and the ceiling |
Classic is the previous step-by-step engine: no checkpointing, no durable gates, and a crash mid-run ends that run. It remains the opt-out, both per conversation and host-wide via Sophon:GraphRuntime:Enabled=false. An admin turning Graph off for the workspace puts everyone on Classic regardless of their own preference.
Execution mode governs chat plans. A plan made inside a plan step or a sub-agent always runs Classic — a nested graph run would deadlock on the host's concurrency budget — and plans made by scheduled, heartbeat and webhook automations run on Graph like any other.
Approval
Whether a plan waits for approval is governed by the plan-approval mode (config Sophon:AgentExecution:PlanApproval, live-reloaded). The default for interactive sessions is Always: the user sees the full plan with all steps and can approve, reject, or edit (timeout: 10 minutes). Editing opens the plan in the Workflow Builder pre-populated. Unattended runs — heartbeat, cron, subagent, webhook — auto-proceed past the plan gate; a per-session override and a host default let you tune the behavior.
Plan approval is up-front: you approve once and aren't interrupted step by step. Three things can still ask for you mid-run:
- A plan revision that raises risk, adds a new tool family, or adds a side-effecting step — parked durably, as described above.
- A tool approval raised inside a running step, for a call above the step's auto-approve ceiling.
- A question a step asks you outright.
The last two are mirrored into your main conversation — Dashboard, Mobile, CLI, and the messaging channel the conversation came from — tagged with the step they came from, and the plan card flips that step to a waiting state. Unlike a revision gate, they wait in memory on a short timer: an unanswered approval is treated as rejected, an unanswered question tells the step to proceed on its best judgment, and a Gateway restart mid-wait fails that step attempt so the retry ladder picks it up again.
See Approval Gates for the risk model and thresholds.
Model hints
A plan step can declare capability requirements:
{
"id": "s4",
"description": "Analyze the results and write a two-paragraph executive summary",
"modelHint": { "requiresReasoning": true }
}The Capability Routing middleware picks a provider/model that advertises reasoning support for that step. If no provider qualifies, the step errors rather than silently falling back to an unqualified model.
Dynamic tool chains
Related to planning is the tool.create_chain meta-tool, which lets the LLM compose existing tools into a new tool at runtime:
{
"name": "research_and_save",
"description": "Search, scrape, and save to memory",
"steps": [
{ "toolName": "web.search", "arguments": { "query": "{{input.topic}}" } },
{ "toolName": "browser.navigate", "arguments": { "url": "{{previous_result.top_url}}" } },
{ "toolName": "memory.write", "arguments": { "content": "{{previous_result.text}}" } }
]
}Chain risk is the max of all constituent tool risks, and every referenced tool is validated before the chain is registered.
Configuration
{
"Sophon": {
"AgentExecution": {
"PlanMaxSteps": 20
},
"GraphRuntime": {
"Enabled": true
}
}
}Sophon:GraphRuntime:Enabled is the host-wide master switch for graph execution; setting it to false restores the previous step-by-step engine for everyone. The planning mode is per-agent; set it in the agent's config or via sophon agents edit <id> --planning-mode Auto|LlmDecides|Always|Never.
Where to go next
- Sophon Graph — the durable execution engine approved plans run on
- Orchestration Pipeline — how planning fits into the broader request flow
- Approval Gates — the risk model and how plans, revisions, and tool calls are gated