Give it a goal. Get finished work.
Sophon reasons through complex requests, remembers people and projects across every conversation, and runs real workflows — the same assistant on Slack, mobile, and the dashboard. Self-hosted and model-agnostic, enterprise-ready when you are.
The interchangeable part
The model is a config change.
Sophon Gateway ships with no pre-configured provider. Point it at any of 30+ models, cloud or local, on your own keys — and override the choice for a single conversation without touching your routing rules.
Cloud
Anthropic, OpenAI, Gemini, Azure, Bedrock, Mistral, Groq, DeepSeek, and more.
Local
Ollama and LM Studio, by base URL. No key, no account, no egress.
OpenAI-compatible
Anything that speaks the API, including your own endpoint.
What Sophon is
Three things a model doesn't come with.
Swap the model as often as you like. These three stay — and they are the part that decides whether an agent is useful on Friday for something you asked on Monday.
Layer one — state
Everything it knows about you is a row you can read.
A context window is not memory — it's a buffer that empties. Sophon keeps what matters outside the window, in a store you can inspect, correct, and delete.
What the model actually receives
Up to 50 user-scoped entries, up to 50 agent-scoped entries, the last 2 days of daily logs, and the last 100 messages. Past those caps, entries are ranked against the current message and the most relevant float up.
None of this is learning. Sophon moves no weights — no training, no fine-tuning, no reinforcement learning. Consolidation is a cron job writing rows, and you can read every one.
Semantic search runs on Pro and Enterprise. On Personal, memory search is keyword-only.
Layer two — execution
A run that outlives the process running it.
Sophon Graph compiles your request into a durable graph of steps. Forty invoices, a crash at number twenty-five, and the twenty-four already done are not processed — or billed — a second time.
Graph mode is on by default, switchable per conversation, and drives a run from a single Gateway instance. Classic mode is still a single loop — a crash mid-turn marks that run failed.
Voice — the same three layers, out loud
Voice is another way in. Not a way around.
In Sophon 2.0 a spoken request becomes the same kind of task a typed one does: queued on your chat, gated by the same approvals, and live in the same thread.
Interrupting silences the reply, not the task behind it. On every tier; with no speech provider, voice uses your browser’s own speech (Chrome or Edge).
What the mind can do
Four faculties, one assistant.
Everything Sophon does flows from the same core — here's what that looks like in practice.
It remembers — and keeps its memory clean.
A self-maintaining graph of the people, projects, and dates in your world. New facts supersede old ones, everything stays searchable, and every entry is yours to inspect.
- Entity graph of people, projects, places, and dates
- Temporal supersession keeps facts current
- Keyword search on every tier; hybrid semantic on Pro and Enterprise
- User- and agent-scoped, fully inspectable
Automations you can see.
Drag-and-drop a workflow, or just describe it and let Sophon draft it. Triggers, logic, approvals, sub-workflows — 20+ node types on an infinite canvas.
- 8 trigger types, 20+ node types
- AI-authored from natural language
- Human approval gates for sensitive steps
- Versioning and sub-workflows
One request. Many steps. Zero micromanagement.
Sophon plans the work, runs independent steps in parallel, checks with you on anything sensitive, and recovers when a step fails — while you watch in real time.
- 12-stage middleware pipeline per message
- Parallel execution of independent steps
- Smart retries and cascades on failure
- Approval gates on anything sensitive
When one opinion isn't enough.
Put three agents with different perspectives in a room, let them argue, and get a judge's structured verdict — a decision, not a decision tree.
- Three agents, distinct perspectives
- A judge renders a structured verdict
- Streams live, end to end
Ship it — the upside is real.
Not until we cover the edge cases.
Phase it: pilot, then roll out.
Pilot with two teams, then decide on the full rollout.
Up and running in minutes
For one, or for all
Personal enough for one. Accountable enough for a company.
The same platform, one config change apart.
For you
Your own assistant, on your own hardware — running the moment you are.
- Free for individual, non-commercial use
- Self-host with one Docker command
- Bring any model — Claude, GPT, or local
- Your data stays in your deployment — only the providers you choose see it
For your organization
The assistant your security team will actually approve — on your infrastructure, under your controls.
- SSO / OIDC / SAML and RBAC
- Audit logging and multi-tenancy
- Kubernetes + Helm, Vault-backed secrets
- Dedicated support and custom SLA
The other half
Here's what Sophon doesn't do.
Everything above is checkable. So is this. Each line links the page in our own documentation that says the same thing.
Risk levels
Pipeline stages
Channel adapters
ATLAS techniques mapped
Counted from the documentation, not rounded up. Every one of these is configurable, and every one is written down.
Trusted by design
Run on your terms. Built to be trusted.
Self-hosted from day one, model-agnostic by design, and accountable end to end. Join the community, browse the marketplace, and follow every release.
Morning briefing
Pulls your calendar and inbox in parallel, drafts a briefing, and asks before it sends.
Inbox triage
Reads new mail, flags what's urgent, and files the rest — on a schedule you set once.
Release deck
Give it a topic and three competitors; it researches, positions, and returns a draft.
Run the layer yourself.
Docker Compose, a model key, a channel. The Personal tier is free for individual, non-commercial use — commercial deployment needs Pro or Enterprise.