Agentic OS (AIOS)
A personal operating layer built on top of a coding agent, structured as four stacked levels — codified skills, persistent memory/state, a visual interface, and distribution to other people. The organizing claim is that the value sits in the two invisible levels, not in the dashboard that makes an AIOS look like one.
Definition
An Agentic OS (AIOS) is the bundling of an agent’s codified workflows, a structured knowledge vault, an optional visual shell, and an optional distribution mechanism into a single customized product for one operator or one team. The framing deliberately separates what an AIOS looks like from what makes it work: the dashboards, metrics tiles, and voice interfaces that define the genre visually are, by this account, the last and least valuable layer.
The source frames two failure modes symmetrically. One audience sees an AIOS dashboard and wants the dashboard; the other sees it and dismisses the whole thing as “smoke and mirrors.” Both miss the same thing — the codification and state layers underneath, which are what actually change how the agent performs.
Implications: For this vault, the AIOS framing is a warning against measuring progress by visible surface. This vault already has levels 1 and 2 in substance — synaptic-* skills as codified workflows, raw/ + wiki/ as state — and no level 3 at all. By this model that is the correct order, not a gap.
The Four Levels
- Skills and loop engineering — every repeated task is turned into a skill, then into an automation, then optionally into a self-improving loop.
- Memory and state — a vault or database the agent reads from and writes back to, structured so the agent can navigate it cheaply.
- Interface and customization — a web app or Obsidian plug-in wrapping levels 1 and 2 in buttons, metrics, and optionally voice.
- Distribution — handing the whole construct to teammates or clients who never touch a terminal.
The load-bearing claim is the value distribution: levels 1 and 2 hold roughly 90% of the value, and both are fully achievable inside a plain terminal session with no interface work at all. Levels 3 and 4 are described as “the cherry on top.” The stack is also presented as model-agnostic — the same construct runs on claude-code, Codex, or a local model.
Implications: This inverts the usual build order. If levels 1 and 2 carry the value and require no UI, then the honest sequencing for anyone building an AIOS is to postpone the interface until the codification and state layers are proven — which also means an AIOS that never grows a dashboard is not an incomplete AIOS.
ARMS Framework
The newer source presents a related bottom-up model called ARMS: Applications, Routines, Memory, and Skills. It starts with Skills, adds Memory once the workspace becomes difficult to navigate, then adds scheduled Routines, and finally connects or builds Applications. The source treats the order as a learning sequence rather than four independent features.
The two models describe overlapping territory from different angles. The earlier four-level model ranks value by visibility and distribution; ARMS describes the operating components an individual assembles. In ARMS, the dashboard is an Application on top of the underlying Skills, Memory, and Routines, not the Agentic OS itself.
Implications: agent-routines fills the scheduling layer that the earlier four-level page named but did not develop. For this vault, ARMS supports keeping governance and memory structures ahead of visual interfaces while treating applications as optional delivery surfaces.
Level 1: The Workflow Audit
Skills cannot be written before knowing which outputs are actually needed repeatedly. The source gives three elicitation methods, and they are not equivalent in reliability:
- Manual codification — perform the task by hand, confirm it works, then tell the agent to turn that validated run into a skill. Guards against codifying a workflow the agent was never good at.
- Session-history mining — have the agent read the last 10–20 sessions, extract repeated tasks, and return a table of task / expected output / proposed skill. Preferred, because it runs on observed behavior rather than self-report: “this isn’t a guessing game of what you think you should do.”
- Stream-of-consciousness interview — describe the work aloud and have the agent probe for blind spots before extracting candidate skills.
The stated mental model is onboarding a personal assistant: enumerate what you do, then hand over step-by-step instructions. The gap the source identifies is not that people cannot name their repeated tasks — it is that even people who can name them have not converted them into skills or automations.
Implications: Session-history mining is the method most transferable to this vault, and it is cheap: agent session logs already exist and already record what was actually asked for, repeatedly, without anyone having designed a survey. It is the same evidence-over-assertion instinct the vault applies to sources, turned on the operator’s own behavior.
Level 2: The Map, Not the Folders
The source describes the andrej-karpathy vault layout — unstructured input, structured wiki, deliverable outputs — and then argues the folders themselves are close to arbitrary. What matters is that an index.md exists at every level of the hierarchy, so that in each directory the agent enters there is a canonical place to learn what it is looking at and where to go next.
The argument given for this is economic rather than aesthetic. A flat, unlinked folder of many files forces the agent to search: “it’s going to use more tokens and ultimately it’s going to cost you more money.” A well-indexed vault is faster and cheaper for the same question. The source explicitly says the raw/wiki/outputs split is substitutable — “you just need a map for Claude Code that makes sense,” and that map will be unique to each person’s data. A CLAUDE.md recording vault conventions and the navigation path is named as the companion piece.
Obsidian is likewise framed as convenience, not requirement: everything demonstrated works in a plain database or plain folders, and a coherent file structure alone gets “99% of the way there.” See second-brain for the same layer viewed from the user’s side, and llm-wiki for the structured layer itself.
Implications: This reframes navigation files from bookkeeping into a cost control. This vault’s index.md, log.md, and AGENTS.md are, under this reading, not documentation overhead — they are the mechanism that keeps query cost flat as page count grows. It also means the correct response to a growing subfolder is a local index, not a reorganization.
Headless Execution
Every button in an AIOS interface resolves to a headless agent invocation — claude -p — running the same skill the operator would run manually in a terminal, with no visible window. The interface is therefore a launcher, not a separate runtime; whatever levels 1 and 2 can do, the buttons can do, and nothing more.
The source also records a billing claim, hedged in the source itself and not verified here: that Anthropic stated claude -p would draw on a $200 API-tied credit rather than the subscription, then “sort of walked back from that,” and that as of recording (2026-06-25) it still billed against the Max plan. This is preserved as a dated, second-hand report, not as a fact about current billing.
Implications: If the UI is only a launcher over headless sessions, then no interface work can compensate for missing skills — which is the mechanical reason levels 1 and 2 dominate the value. It also means level-3 cost is entirely a function of how often buttons get pressed, and the billing question above decides whether that cost is bounded by a subscription or metered per call.
Level 4: Distribution as Floor-Raising
An AIOS can be handed to people who will not use a terminal. Web-app builds distribute easily — a GitHub repository or a zip — while an Obsidian plug-in build requires hands-on per-person setup. The claimed payoff is organizational: a non-technical teammate pressing a button gets the output of a skill without ever being onboarded onto the agent itself.
The source’s blunt version of the audience claim is that roughly 99% of people will not approach a terminal or even a desktop agent app regardless of the value offered, and that technical audiences systematically underestimate this. It closes by noting that the “dashboard effect” on non-technical users — that a visual wrapper changes how people interpret an otherwise identical tool — “genuinely needs to be studied.”
Implications: Distribution is where level 3 stops being decoration and becomes the actual product: for a solo operator a dashboard is optional polish, but for anyone handing capability to a team it is the only delivery mechanism available. That is a different justification than the one usually offered for building one.
Related
- llm-as-os — the abstract “LLM as kernel” framing that this four-level stack operationalizes
- claude-code — the agent the construct is built on
- second-brain — level 2 seen from the individual user’s perspective
- claude-skills — the unit of codification at level 1
- ralph-loop — a concrete loop-engineering pattern the source gestures at but does not develop
- agent-loop — the underlying observe/decide/execute cycle each automation runs
Open Questions
- Is the
claude -pbilling behavior described here (subscription vs API credit) still accurate after 2026-06-25? The source is second-hand and explicitly hedged. - Loop engineering is named as the top of level 1 but deferred to a separate video; what distinguishes a “self-improving” automation from a scheduled one is not defined in this source.
- The source asserts levels 1 and 2 are ~90% of the value without offering a measurement. Is that a defensible ratio or a rhetorical one?
- Does a single-operator vault like this one ever justify level 3, given that its only stated payoff — reaching non-terminal users — does not apply?