Sophon 2.0 is here

Give it a goal. Get finished work.

Sophon reasons through complex requests, remembers people and projects across every conversation, and runs real workflows — the same assistant on Slack, mobile, and the dashboard. Self-hosted and model-agnostic, enterprise-ready when you are.

Self-hostedModel-agnosticEnterprise-ready17 channels
sophon.local/chat
Sophon
Chat
Agents
Workflows
Memory
Skills
Approvals2
Morning Briefing
Live
Research our top 3 competitors and email me a comparison by 9am.
Plan · 3 steps
Find competitors via web search
Summarize pricing + positioning
Draft email to you
web.search
"acme pricing", "globex pricing", "initech pricing"
running

The interchangeable part

The model is a config change.

Sophon Gateway ships with no pre-configured provider. Point it at any of 30+ models, cloud or local, on your own keys — and override the choice for a single conversation without touching your routing rules.

Cloud

Anthropic, OpenAI, Gemini, Azure, Bedrock, Mistral, Groq, DeepSeek, and more.

Local

Ollama and LM Studio, by base URL. No key, no account, no egress.

OpenAI-compatible

Anything that speaks the API, including your own endpoint.

Layer one — state

Everything it knows about you is a row you can read.

A context window is not memory — it's a buffer that empties. Sophon keeps what matters outside the window, in a store you can inspect, correct, and delete.

An entity graph, not a transcript
People, projects, places, and dates are linked through the [[wikilinks]] you write and background extraction. “Who's involved in Atlas?” is a graph traversal, not a text scan.
Facts that can be superseded
When something changes, the new entry supersedes the old one: it drops out of default views and search and its vector is deleted, but the row is never destroyed. The Dashboard shows how a fact evolved.
A nightly sweep, on a schedule you can read
MemoryConsolidationJob runs daily at 03:00 UTC, promoting durable facts out of daily logs into long-term entries through the same write path memory.write uses.

What the model actually receives

Up to 50 user-scoped entries, up to 50 agent-scoped entries, the last 2 days of daily logs, and the last 100 messages. Past those caps, entries are ranked against the current message and the most relevant float up.

None of this is learning. Sophon moves no weights — no training, no fine-tuning, no reinforcement learning. Consolidation is a cron job writing rows, and you can read every one.

Semantic search runs on Pro and Enterprise. On Personal, memory search is keyword-only.

How memory works
Reasoning
thinks
Claude
Memory
remembers
Orchestration
acts

Layer two — execution

A run that outlives the process running it.

NEW

Sophon Graph compiles your request into a durable graph of steps. Forty invoices, a crash at number twenty-five, and the twenty-four already done are not processed — or billed — a second time.

Checkpointed, so a restart isn't a rerun
Every transition is committed, so work that already finished is never repeated. A restart brings the run back at the step it interrupted and waits for you to resume it — and a plan revision you haven't signed off on is still parked when it returns.
Parallel by default
Independent steps run at the same time in isolated child sessions with clean context. Three research steps finish in roughly the time of the slowest one, not the sum.
Bounded on purpose
A $2.00 default cost cap per plan run, checked between steps — in-flight steps finish, so a run lands near the cap rather than on it. Plus loop detection: repeated calls, polling with no progress, and A→B→A ping-pong all get caught, with a circuit breaker at 30 repeats.

Graph mode is on by default, switchable per conversation, and drives a run from a single Gateway instance. Classic mode is still a single loop — a crash mid-turn marks that run failed.

Layer three — authority

Authority is enforced outside the model.

A system prompt is a request. A registry is a rule. Sophon puts the decision to act where the model can't reach it — and keeps your credentials in a vault the model's prompt never touches.

“The authority to act is enforced outside the model — approval gates, tool risk levels, and Node scopes mean a jailbroken persona still cannot run a Critical action unattended.”— from our published threat model
See the control layer

The credential path — it never enters the model's prompt

Agent
asks for a result
Vault
encrypted at rest
Sandbox
holds the credential
API
the skill calls it

Risk declared per tool, not per prompt

Five levels — None, Low, Medium, High, Critical — declared in each tool's manifest. A call at or above the threshold you set raises an approval request, and the shipped default is Medium. Critical prompts whatever the threshold or a per-tool override says. Per-agent allowlists are enforced at the registry, so a tool that isn't on the list is refused even if the model asks for it.

Silence is not consent

Every request resolves once and only once — Approved, Edited, Rejected, Cancelled, or Timed out — and a timeout or a cancel can never later become an approval. Answer from the dashboard, an actionable mobile push, out loud in a voice session, or by replying in the channel the request came from.

A veto at the keyboard

Sophon Node raises a native prompt on its own screen before it will run a shell command — 15 seconds, fail-closed. The person sitting at that machine can block it even if the Gateway path is compromised.

Where the ceiling actually sits

  • Approval is a ceiling you set, not a switch the model flips. None and Low run unattended by default.
  • Unattended runs — heartbeat, cron, subagent, webhook — proceed past the plan gate by design. High and Critical still raise their own approvals, and no heartbeat setting lifts that.
  • An auto-approving heartbeat self-approves up to Medium, and message.send is Medium — so it can send mail on your behalf unless your checklist says otherwise.
  • In a multi-tenant deployment, any user in a tenant can answer any pending request.
  • A channel reply is accepted only from the chat the request was delivered to — but anyone in that chat can answer it. Binding a reply to one sender is roadmap.
  • The vault keeps credentials out of the model's prompt, not out of the skill. A sandboxed skill holds the real credential while it runs; a broker that makes the call on its behalf is roadmap.

Voice — the same three layers, out loud

Voice is another way in. Not a way around.

Rebuilt in 2.0

In Sophon 2.0 a spoken request becomes the same kind of task a typed one does: queued on your chat, gated by the same approvals, and live in the same thread.

Kept like typing
Your words and the reply are saved in the chat thread, like typed messages.
Queued like typing
A regular task on that chat, behind anything already running by default, with a smaller, voice-focused tool set.
A higher bar out loud
Answered by voice, High- and Critical-risk actions need “yes, approve”; a bare “yes” isn’t enough. Unclear answers are asked again at most twice, then the prompt waits for a tap.

Interrupting silences the reply, not the task behind it. On every tier; with no speech provider, voice uses your browser’s own speech (Chrome or Edge).

What the mind can do

Four faculties, one assistant.

Everything Sophon does flows from the same core — here's what that looks like in practice.

It remembers — and keeps its memory clean.

A self-maintaining graph of the people, projects, and dates in your world. New facts supersede old ones, everything stays searchable, and every entry is yours to inspect.

  • Entity graph of people, projects, places, and dates
  • Temporal supersession keeps facts current
  • Keyword search on every tier; hybrid semantic on Pro and Enterprise
  • User- and agent-scoped, fully inspectable
Explore Memory
project deadline...
semantickeyword
User lives in Vienna, Austriafact
Prefers dark mode everywherepreference
Use formal tone in emailspreference
Project Atlas deadline: Mar 15fact
Morning standup at 9:30 AMobservation
shared
agent

Automations you can see.

Drag-and-drop a workflow, or just describe it and let Sophon draft it. Triggers, logic, approvals, sub-workflows — 20+ node types on an infinite canvas.

  • 8 trigger types, 20+ node types
  • AI-authored from natural language
  • Human approval gates for sensitive steps
  • Versioning and sub-workflows
Explore Workflows
Cron Trigger
Every day 7 AM
Fetch Emails
Gmail — unread
Summarize
Claude Sonnet
Send Briefing
WhatsApp

One request. Many steps. Zero micromanagement.

Sophon plans the work, runs independent steps in parallel, checks with you on anything sensitive, and recovers when a step fails — while you watch in real time.

  • 12-stage middleware pipeline per message
  • Parallel execution of independent steps
  • Smart retries and cascades on failure
  • Approval gates on anything sensitive
Explore Orchestration
Request
“Prep my 9am”
Check calendar
parallel
Scan inbox
parallel
Draft briefing
merge
Approve send
waiting for you

When one opinion isn't enough.

Put three agents with different perspectives in a room, let them argue, and get a judge's structured verdict — a decision, not a decision tree.

  • Three agents, distinct perspectives
  • A judge renders a structured verdict
  • Streams live, end to end
Explore Discussions
Optimist

Ship it — the upside is real.

Skeptic

Not until we cover the edge cases.

Pragmatist

Phase it: pilot, then roll out.

Judge · verdict

Pilot with two teams, then decide on the full rollout.

Up and running in minutes

Terminal

For one, or for all

Personal enough for one. Accountable enough for a company.

The same platform, one config change apart.

For you

Your own assistant, on your own hardware — running the moment you are.

  • Free for individual, non-commercial use
  • Self-host with one Docker command
  • Bring any model — Claude, GPT, or local
  • Your data stays in your deployment — only the providers you choose see it
Get Started

For your organization

The assistant your security team will actually approve — on your infrastructure, under your controls.

  • SSO / OIDC / SAML and RBAC
  • Audit logging and multi-tenancy
  • Kubernetes + Helm, Vault-backed secrets
  • Dedicated support and custom SLA
0

Risk levels

0

Pipeline stages

0

Channel adapters

0

ATLAS techniques mapped

Counted from the documentation, not rounded up. Every one of these is configurable, and every one is written down.

Trusted by design

Run on your terms. Built to be trusted.

Self-hosted from day one, model-agnostic by design, and accountable end to end. Join the community, browse the marketplace, and follow every release.

4 steps · 3 parallel

Morning briefing

Pulls your calendar and inbox in parallel, drafts a briefing, and asks before it sends.

Runs on a heartbeat

Inbox triage

Reads new mail, flags what's urgent, and files the rest — on a schedule you set once.

Research → draft

Release deck

Give it a topic and three competitors; it researches, positions, and returns a draft.

Run the layer yourself.

Docker Compose, a model key, a channel. The Personal tier is free for individual, non-commercial use — commercial deployment needs Pro or Enterprise.