Skip to content

Troubleshooting

Permission-denied errors on the socket mean the user running the server is not allowed to read it. Add it to the docker group:

Terminal window
sudo usermod -aG docker $USER
# Log out and log back in for the change to take effect

newgrp docker applies it to the current shell without logging out.

For rootless Podman, the socket belongs to your user already. See Podman.

If the database reports a lock, make sure only one instance of the server is running against it. WAL mode allows concurrent reads and writes, but a single background writer is required.

Terminal window
ps aux | grep kaio-server

A common cause is make dev left running in another terminal while a container instance points at the same data/ directory.

  • Check the socket path: KAIO_DOCKER_HOST unset uses local defaults, which may not match a rootless Podman setup.
  • Check the server logs with RUST_LOG=debug.
  • Confirm the API answers: curl localhost:8080/api/health.

Remote image update checks need Skopeo. It ships in the production image; on a source build, install it on the host. Without it, stacks simply never show the ★ indicator.

For a local registry over HTTP, detection is heuristic, so check the server logs for the inspection error if a known-updated image stays unflagged.

The first Trivy scan downloads the vulnerability database and runs alone to warm the cache, so it takes noticeably longer than later ones. If scans never complete:

  • the host must be able to pull KAIO_TRIVY_IMAGE;
  • check the Docker volume holding the Trivy cache has space.

unknown means it has not been probed yet, not that it failed. Probes run every KAIO_NODE_HEALTH_INTERVAL_SECONDS (default 60), and 0 disables them entirely, which leaves every row frozen on its last result. Force one:

Terminal window
kaio-cli node check prod-01
kaio-cli node check

Enrolling probes the node immediately, so a status that is already wrong at join time is almost always the advertised address rather than the node itself. The control plane reaches the node at the address the node gave it, which is not necessarily the address you typed to reach the control plane.

Terminal window
kaio-cli node show prod-01 # the URL actually stored
curl http://<that url>/api/health # from the control plane, not from your laptop

Rejoin with an explicit --advertise host:port when the node sits behind a port mapping, a NAT, or a wildcard bind. See Multiple servers.

Something in front of that node asked for a credential the control plane does not have, typically a reverse proxy with authentication added after the join. Kaio itself never asks for one. Either expose the node’s API on a path the control plane can reach unauthenticated, or drop it from the registry with kaio-cli node rm prod-01: an unreachable entry is only a dead route.

  • Invalid join token: the token was rotated. Read the current one on the control plane with kaio-cli cluster token, and note that kaio-cli cluster token --rotate retires the previous one.
  • No token at all: a server that has never minted one has nothing to verify against. Run kaio-cli cluster token on it first.
  • A control plane cannot join itself: registering the server that already holds the registry is refused by design.
kaio-server[1242]: thread 'main' panicked at crates/server/src/main.rs:31:10:
kaio-server[1242]: Failed to connect to Docker
kaio.service: Scheduled restart job, restart counter is at 93.

The host has no Docker daemon, or the one it has is unreachable by the kaio user. kaio-server talks to the daemon at startup and exits when it cannot, and Restart=always then loops. Install Docker, which is a prerequisite rather than an option:

Terminal window
curl -fsSL https://get.docker.com | sh
systemctl restart kaio

The installer names this as its first step when docker is not on PATH. If Docker is installed and the loop persists, the daemon is elsewhere: point KAIO_DOCKER_HOST at it in /etc/kaio.env, and check that the kaio user is in the docker group with id kaio.

The gateway only forwards to servers in the registry, so a 404 there means the node is not in it, usually because it was removed or it left. That says nothing about the node’s own health: its address stays exactly as reachable as before, and KAIO_ADDR=<node> kaio-cli status still works. It rejoins by running kaio-cli join again, from the node itself.

A deploy failed and the stack looks unchanged

Section titled “A deploy failed and the stack looks unchanged”

That is intended: a failed compose up restores the previous compose.yml and .env, and the submitted variables are not persisted. Read the error returned by the deploy: it is the raw Compose output.

The toolchain is pinned by Nix. Wrap the command:

Terminal window
nix-shell shell.nix --run "cargo test --workspace"

Kaio, built by Régis Gaidot