Sophon 2.0 is here

Opinion

AGI will need an operating system

By Enes Hoxha, Co-Founder, Buildersoft LLC · 19 September 2026

What this is, and what it isn't

This is an opinion piece and I want the label on it before the argument starts. I run a company that sells the thing this essay says is underrated, so read it the way you would read any argument from someone with an obvious interest in it being true. I don't train models. I couldn't if I wanted to — Buildersoft is a small company and there is no research lab here.

Nothing in this essay says Sophon is artificial general intelligence, approaches it, or is a step toward it. Sophon calls models it did not train and does not host; it has no continual learning, no fine-tuning, and no weights of its own. Where I cite our product below, every claim is a shipped mechanism with a number attached, and the things we haven't built are listed at the end of this page rather than left out of it.

The field agrees on the gaps

The barriers people name out loud are continual learning, long-horizon memory, and sustained agency. What I notice is that when researchers write down what would have to be built to close them, the list reads like the programme of an operating-systems conference.

The AgenticOS workshop at SOSP 2026 asks for long-lived state abstractions for agent context and episodic memory, dynamic sandboxing for agent-generated code, fault tolerance across distributed agents, provenance and debugging of agent executions, and isolation for agent-invoked tools and data flows. Not one of those is a model-architecture problem. Every one is a storage, scheduling, or permissions problem.

A model is a function. A process is something else.

A model call is stateless. Text goes in, text comes out, and whatever the model appeared to know evaporates when the request closes. Everything we call an agent is scaffolding around that fact: something that decides what to put in, something that holds on to what came out, something that decides whether to call again.

That space between one model call and the next is the least engineered part of the entire stack. Usually it is a for-loop somebody wrote in an afternoon. It has no durability story, no memory model beyond a list that gets truncated, and no permission system beyond hoping the system prompt holds. We have spent five decades learning how to build processes — scheduling, isolation, persistence, capabilities — and almost none of it has been applied here.

My argument is narrow, and I think it is boring in the way true things often are: whatever the models turn out to be capable of, they will still be stateless functions, and someone has to build the process around them. That is an engineering problem with a literature, and it is available to work on today.

Memory is a storage problem

The interesting operation in a memory system is not writing. It is forgetting on purpose. A store that only accumulates gets worse over time, because stale facts compete with current ones and the retrieval layer cannot tell them apart.

Sophon writes memory entries into an entity graph — people, projects, places — built from the wikilinks you write and background extraction. When a fact changes, the new entry supersedes the old one: the superseded row drops out of default views and search and its vector is deleted, but the row itself is never destroyed, so you can inspect how a fact evolved. Duplicate entities are flagged at 0.85 trigram similarity into a review sheet where you merge or dismiss them. A cron job at 03:00 UTC promotes durable facts out of daily logs into long-term entries. What the model actually receives is capped and published: 50 user-scoped entries, 50 agent-scoped entries, the last 2 days of logs, the last 100 messages.

None of that is learning. It is a cron job writing rows and a retrieval budget. And on the Personal tier there is no vector store at all — search is keyword-only. The field's memory problem is weight-level continual learning. I am using “memory” in the systems sense. These are different problems and I am only claiming the easy one.

Continuity is a durability problem

An agent that cannot survive a restart cannot be trusted with anything that takes longer than a restart. This sounds obvious and is almost universally unsolved, because the for-loop keeps its state in memory and the memory goes away.

Sophon Graph compiles a request into a versioned graph and checkpoints every transition, so a run that dies at step twenty-five of forty picks up at twenty-five rather than at one. Independent steps run concurrently in isolated child sessions. Runs are bounded on purpose: a default cost cap per run, a per-node wall clock, and loop detection that catches repeated calls, polling with no progress, and A-to-B-to-A ping-pong, with a circuit breaker at thirty repeats. Unattended work runs a standing checklist and replies with a literal HEARTBEAT_OK when there is nothing to say.

Sustained agency, in practice, looks less like ambition and more like a process supervisor. The interesting question was never how to make an agent keep going. It is how to make it stop.

That durability is Graph mode, and it is not automatic: a restart parks a run that was mid-step and waits for you to press Resume. The finished steps are never re-run, and the cost cap is checked between steps, so a run overshoots it by whatever the steps already running cost. Classic mode is still a single loop, and a task caught mid-turn by a restart is marked failed. The task queue is an in-memory bounded channel on every tier today — Redis- and RabbitMQ-backed queues, and therefore horizontal scaling, are roadmap. And Sophon Graph shipped before its documentation did, which is a gap I own.

Control is the part nobody wants to write

A system that can write its own tools and put them on a schedule is either governed or it is a liability. There is no third state. This is the least glamorous work in the field and it is the work that decides whether any of the rest of it is deployable.

The load-bearing idea is that authority lives outside the model. Every tool declares one of five risk levels in its manifest; a call at or above the threshold you set raises an approval request, the shipped default is Medium, and Critical prompts whatever the threshold or a per-tool override says; a request that times out is recorded as a rejection rather than a yes. Per-agent allowlists are enforced at the registry, so a tool that is not on the list is refused no matter how the model asks. Credentials sit in an encrypted vault and are delivered into the skill's sandbox at call time, never into the model's prompt. An agent can author its own skill, but it is tested in a sandbox and then waits for a human before installation. On a Sophon Node, a shell command raises a native prompt on that machine's own screen — fifteen seconds, fail-closed — so the person at the keyboard can veto it even if the Gateway path is compromised.

The ceiling is configurable, which means it is also lowerable. Unattended runs proceed past the plan gate by design; an auto-approving heartbeat self-approves up to Medium, and sending mail is Medium. In a multi-tenant deployment any user in the tenant can answer any pending request. Injection detection defaults to Warn, which annotates and logs; Block is opt-in and stops a turn only on third-party content, and it matches phrasing, so an obfuscated or split payload still gets through. The sandbox holds the real credential while a skill runs — a broker that makes the call on the skill's behalf is roadmap. Plugins run full-trust and are not a security boundary. Audit logging is Enterprise-only.

Why "operating system" and not "framework"

I don't mean the property of being impossible to turn off. By that test almost nothing qualifies, including us. I mean the property of being the default, and of every override being deliberate, scoped, and logged.

A framework is something a developer imports and configures per call site. The difference I care about is where the decision lives. When permission is a function of the prompt, every new capability re-opens the question. When it is a function of the registry, the question is asked once, in one place, by a person. High and Critical actions page you regardless of any heartbeat setting; everything below that is a ceiling you set on purpose and can see.

Sophon's topology is hub-and-spoke and I would not want anyone to read more into the word than that. There is no federation and no peer-to-peer. A Node is a dumb executor: it receives a command, runs it, returns a result, and never decides anything.

Where this goes

We are building Sophon OS, shipping September 2026 — two shells on a Linux foundation with a Rust shell-host, under a commercial licence. The reason is scope, not capability. Today our consent layer stops at our own process boundary. Below that, the machine still does what it is told, and an approval gate that only governs one application is a gate with a wall missing on one side.

I don't think that makes the operating system smart. It makes consent a system property rather than an application feature, which is a much smaller and much more achievable claim.

What would prove me wrong

Three things would make this argument obsolete, and I would rather name them than be quietly wrong for a couple of years.

If a frontier lab ships weight-level memory that is durable, user-scoped, auditable, and revocable, the storage argument collapses into the model and my fourth section is a footnote. If context windows become large and cheap enough to hold a working lifetime, the retrieval budget stops mattering. And if a consent primitive lands at the inference layer that a caller cannot bypass, the control argument moves down the stack and out of my hands.

I don't know whether the argument is right. It is dated and versioned so you can check later.

What Sophon doesn't do

Every claim above has a boundary. Here they are in one place, as of this revision.

  • It does not learn. No training, no fine-tuning, no reinforcement learning, no weight updates of any kind. Consolidation is a Quartz cron job that writes rows to a database at 03:00 UTC.
  • Its anomaly detection is rule-based. Not ML. It catches an obvious spike against the prior period and misses subtle drift.
  • Its prompt-injection detection is a phrase filter. Warn, the default, annotates and logs. Block is opt-in and stops a turn only on third-party content. Obfuscated or split instructions get through; the approval gates are still the prevention.
  • Its plugins are not sandboxed. They run full-trust, in-process privileges, out of process only for stability. Skills are the sandboxed extension point.
  • Its memory is keyword-only on the free tier. Vector-backed semantic recall needs a vector store, and the Personal tier ships none.
  • Its distribution is hub-and-spoke. No federation, no peer-to-peer. A Node is a dumb executor; every decision happens on the Gateway.
  • Its execution engine shipped ahead of its documentation. Sophon Graph has a product page and no reference page. That is a gap, and it is ours.

The product pages carry the same list, and the threat model carries the residual risks we haven't closed.

Revision 1.1 · 19 September 2026 · Disagreements: hello@buildersoft.io

Enes Hoxha, Co-Founder, Buildersoft LLC