Troubleshooting
Docker socket permissions
Section titled “Docker socket permissions”Permission-denied errors on the socket mean the user running the server is not
allowed to read it. Add it to the docker group:
sudo usermod -aG docker $USER# Log out and log back in for the change to take effectnewgrp docker applies it to the current shell without logging out.
For rootless Podman, the socket belongs to your user already. See Podman.
Database lock issues
Section titled “Database lock issues”If the database reports a lock, make sure only one instance of the server is running against it. WAL mode allows concurrent reads and writes, but a single background writer is required.
ps aux | grep kaio-serverA common cause is make dev left running in another terminal while a container
instance points at the same data/ directory.
The dashboard is empty
Section titled “The dashboard is empty”- Check the socket path:
KAIO_DOCKER_HOSTunset uses local defaults, which may not match a rootless Podman setup. - Check the server logs with
RUST_LOG=debug. - Confirm the API answers:
curl localhost:8080/api/health.
Updates are never detected
Section titled “Updates are never detected”Remote image update checks need Skopeo. It ships in the production image; on a source build, install it on the host. Without it, stacks simply never show the ★ indicator.
For a local registry over HTTP, detection is heuristic, so check the server logs for the inspection error if a known-updated image stays unflagged.
Scans stay pending
Section titled “Scans stay pending”The first Trivy scan downloads the vulnerability database and runs alone to warm the cache, so it takes noticeably longer than later ones. If scans never complete:
- the host must be able to pull
KAIO_TRIVY_IMAGE; - check the Docker volume holding the Trivy cache has space.
A node never leaves unknown
Section titled “A node never leaves unknown”unknown means it has not been probed yet, not that it failed. Probes run every
KAIO_NODE_HEALTH_INTERVAL_SECONDS (default 60), and 0 disables them
entirely, which leaves every row frozen on its last result. Force one:
kaio-cli node check prod-01kaio-cli node checkA node is offline right after joining
Section titled “A node is offline right after joining”Enrolling probes the node immediately, so a status that is already wrong at join time is almost always the advertised address rather than the node itself. The control plane reaches the node at the address the node gave it, which is not necessarily the address you typed to reach the control plane.
kaio-cli node show prod-01 # the URL actually storedcurl http://<that url>/api/health # from the control plane, not from your laptopRejoin with an explicit --advertise host:port when the node sits behind a port
mapping, a NAT, or a wildcard bind. See
Multiple servers.
A node is unauthorized
Section titled “A node is unauthorized”Something in front of that node asked for a credential the control plane does
not have, typically a reverse proxy with authentication added after the join.
Kaio itself never asks for one. Either expose the node’s API on a path the
control plane can reach unauthenticated, or drop it from the registry with
kaio-cli node rm prod-01: an unreachable entry is only a dead route.
Joining is refused
Section titled “Joining is refused”Invalid join token: the token was rotated. Read the current one on the control plane withkaio-cli cluster token, and note thatkaio-cli cluster token --rotateretires the previous one.- No token at all: a server that has never minted one has nothing to verify
against. Run
kaio-cli cluster tokenon it first. - A control plane cannot join itself: registering the server that already holds the registry is refused by design.
The service restarts every five seconds
Section titled “The service restarts every five seconds”kaio-server[1242]: thread 'main' panicked at crates/server/src/main.rs:31:10:kaio-server[1242]: Failed to connect to Dockerkaio.service: Scheduled restart job, restart counter is at 93.The host has no Docker daemon, or the one it has is unreachable by the kaio
user. kaio-server talks to the daemon at startup and exits when it cannot,
and Restart=always then loops. Install Docker, which is a prerequisite rather
than an option:
curl -fsSL https://get.docker.com | shsystemctl restart kaioThe installer names this as its first step when docker is not on PATH. If
Docker is installed and the loop persists, the daemon is elsewhere: point
KAIO_DOCKER_HOST at it in /etc/kaio.env, and check that the kaio user is
in the docker group with id kaio.
node ... run answers 404
Section titled “node ... run answers 404”The gateway only forwards to servers in the registry, so a 404 there means the
node is not in it, usually because it was removed or it left. That says nothing
about the node’s own health: its address stays exactly as reachable as before,
and KAIO_ADDR=<node> kaio-cli status still works. It rejoins by running
kaio-cli join again, from the node itself.
A deploy failed and the stack looks unchanged
Section titled “A deploy failed and the stack looks unchanged”That is intended: a failed compose up restores the previous compose.yml and
.env, and the submitted variables are not persisted. Read the error returned
by the deploy: it is the raw Compose output.
Bare cargo fails in a dev checkout
Section titled “Bare cargo fails in a dev checkout”The toolchain is pinned by Nix. Wrap the command:
nix-shell shell.nix --run "cargo test --workspace"Kaio, built by Régis Gaidot