Skip to content

Overview

Kaio follows Hexagonal / Clean Architecture: a shared core crate holds all business logic, a thin server crate exposes it over HTTP, a cli crate consumes that HTTP API, and an mcp crate turns the same core into tools for AI agents.

flowchart TD
subgraph clients [Clients]
  web["React front<br/>Tailwind, Monaco"]
  term["CLI and TUI<br/>comfy-table"]
  agent["AI agent<br/>MCP client"]
end

web -- "REST, SSE, WS" --> server
term -- "REST, SSE" --> server
agent -- "MCP, streamable HTTP" --> server

server["server crate<br/>Axum, thin HTTP layer"] --> mcp
server --> core
mcp["mcp crate<br/>tools at /mcp"] --> core
core["core crate<br/>stacks, scans, events, models, db"] <--> db[("SQLite, WAL")]

core --> sync["Docker sync worker"]
core --> checker["Update checker<br/>skopeo"]
core --> files["./data/stacks<br/>compose and env"]

sync --> sock["/var/run/docker.sock"]
checker --> registries["Registries"]
Every client talks to the same server, which delegates every decision to core.

Kaio drives the daemon of the host it runs on, and nothing else. Operating several hosts means running Kaio on each and having them join one another.

flowchart TD
subgraph cp ["Control plane, the server you open"]
  cpcore["core"]
  cpdb[("SQLite<br/>cluster, nodes")]
  cpcore <--> cpdb
  cpsock["/var/run/docker.sock"]
  cpcore --> cpsock
end

subgraph n1 ["Node: prod-01"]
  n1core["core"]
  n1sock["/var/run/docker.sock"]
  n1core --> n1sock
end

subgraph n2 ["Node: prod-02"]
  n2core["core"]
  n2sock["/var/run/docker.sock"]
  n2core --> n2sock
end

n1core -- "1. join, with the token" --> cpcore
cpcore -- "2. GET /api/health" --> n1core
cpcore -- "3. /api/nodes/prod-01/proxy/..." --> n1core
n2core -- "4. leave: DELETE /api/nodes/prod-02" --> cpcore
A node enrols itself, the control plane probes it and then forwards to it on request, and a node that leaves asks to be dropped. Every arrow uses the API each Kaio already serves.

There is no agent. A node runs an ordinary Kaio, and the HTTP API it already exposes is the protocol the control plane speaks to it. Nothing is installed on a node that is not installed on a single-host install.

The token is checked at step 1 and never again. It is not a credential the two keep presenting to each other: what binds them afterwards is a row in the registry, and what ends it is that row going away.

Both are initiated by the node, from the machine that is joining or leaving, and both reach across:

The node does The control plane ends up with
join presents the token and its own identity a row: name, URL, and the probe that follows
leave asks for its row to be dropped nothing, and the route closes with it

A node that cannot reach its control plane still leaves, locally, and says what is left to do. Decommissioning a server does not depend on another one being up.

The control plane can also evict with node rm, which drops the row without telling the node. Either side can end it; only the node’s own command leaves both in agreement.

Membership lives in a one-row cluster table on both sides, and the registry in a nodes table on the control plane:

Table On the control plane On a node
cluster its own identity, and the join token it hands out its own identity, and which cluster it joined
nodes one row per joined server empty

Leaving clears the node’s three membership columns and deletes the control plane’s row. The identities on both sides are untouched: a server keeps the same one across joins, which is what makes it recognisable if it comes back.

A server’s identity is a UUID minted on first use and kept for the life of its data/. It is the id of the node row, which is what makes a moved server recognisable: rejoining from a new address corrects the row instead of adding a second one. It is also how a control plane refuses to enrol itself, and how a probe tells a node apart from a stranger that took over its address.

/api/nodes/{name}/proxy/{*path} forwards to {node url}/{path}. It is what lets one dashboard drive any node, and what kaio-cli node prod-01 run uses: the name is resolved against the registry, so a node that is not in it has no route. Responses are streamed rather than buffered, so a followed log arrives line by line, and a WebSocket upgrade is bridged socket to socket, so the container terminal works across the hop.

The registry says where each server is and whether it answers. It schedules nothing: there is no placement, no failover, no workload moved between hosts. And it authorises nothing: it decides what the control plane can route to, while each node’s own address stays as reachable as it ever was.

Crate Role
core Business logic: stack lifecycle orchestration, vulnerability scanning, event recording, strictly typed models, Docker integration, and all database access
server Axum REST API and SSE, delegating every decision to core
cli Command-line interface and TUI, a thin client over the server API

Centralising persistence in core is what lets the CLI reuse the same data layer, and what keeps aggregation (stack state, CVE counts) identical across the three interfaces.

  • Backend in Rust: Axum, SQLx, Bollard, Tokio
  • Frontend: React 19, TypeScript, Tailwind CSS 4, Monaco Editor and Xterm.js (both lazy-loaded)
  • Database: SQLite via SQLx, WAL mode
  • Typography: Courier Prime
  • Connection pooling: a single Docker client connection is reused across all API requests and background workers.
  • Smart metrics collection: the poller only fetches expensive stats (CPU/IO) for running or paused containers.
  • Parallel I/O: registry checks and metric collection run concurrently on Tokio; nothing blocking sits on a request path.
  • Frontend: code splitting for xterm.js and monaco-editor (≈50 % off the initial bundle), and useMemo on the heavy aggregations so hundreds of containers stay fluid.
  • Clippy with -D warnings, rustfmt enforced in CI
  • End-to-end type safety: Rust on the backend, TypeScript on the frontend
  • make verify runs the whole gate locally

Kaio, built by Régis Gaidot