Skip to content

Prometheus and Grafana

Kaio exposes its own state as Prometheus metrics at /metrics, on the same port as the API and the web UI. Prometheus keeps the history Kaio deliberately does not, Grafana draws it, and Alertmanager tells you about it: “the media stack has been partial for ten minutes”, “prod-02 has not answered a probe”, “a critical CVE appeared on something in production”.

The endpoint is off until you turn it on.

Terminal window
KAIO_METRICS=on
KAIO_METRICS_TOKEN=a-long-random-string

The server logs what it settled on at startup:

INFO kaio_server: Serving Prometheus metrics at /metrics, behind a bearer token

The token is optional and strongly recommended: /metrics lists your stack names, your image names and your CVE counts, and the REST API in front of it has no authentication of its own. With the token set, a scrape must carry it:

Terminal window
curl -H "Authorization: Bearer a-long-random-string" http://kaio.example.net:8080/metrics

Container CPU and memory are the part cAdvisor already does better, and they are exposed here because the numbers were already in hand. The reason to scrape Kaio is the rest: what Kaio knows and nothing else on the host does.

Metric The question it answers
kaio_stack_status{stack,status} Which stacks are not fully up
kaio_stack_update_available{stack} Which stacks have an image update waiting, and for how long
kaio_stack_env_pending{stack} Which stacks hold environment changes they were never redeployed with
kaio_stack_vulnerabilities{stack,severity} Which stacks run a critical CVE
kaio_stack_version{stack} Which compose version is live, so a rollback shows up as a step down
kaio_node_up{node} Which servers in the cluster have stopped answering
kaio_docker_sync_timestamp_seconds Whether Kaio itself is still reading the daemon

Each Kaio exposes its own host, so point Prometheus at every server rather than at one of them. The gateway (/api/nodes/{name}/proxy/) is deliberately not the route to use: it would mix two hosts under one instance label and make the control plane a single point of failure for your monitoring.

scrape_configs:
- job_name: kaio
metrics_path: /metrics
authorization:
credentials: a-long-random-string
static_configs:
- targets:
- kaio-01:8080
- kaio-02:8080

Every metric carries Prometheus’ own instance label, so sum by (instance) splits any of them per server and the shipped dashboard has a Server picker built on it.

If Kaio is not running yet, or you want the whole thing on one host to see what it looks like, compose.full.yml carries Kaio, Prometheus, Alertmanager and Grafana together, already wired to each other:

Terminal window
make metrics-full-run

Nothing to fill in first. Kaio comes up with KAIO_METRICS=on, Prometheus scrapes it by service name over the compose network, the alert rules are loaded, and Grafana opens on :3000 with admin / admin and the Kaio dashboard in its list. make metrics-full-logs follows all four, make metrics-full-stop takes them down.

The values are in the file in plain text, which is the point of this path: it is meant to be read and run. A few knobs are worth knowing:

Variable Default What it does
KAIO_IMAGE ghcr.io/rgaidot/kaio:latest The Kaio image, so you can pin a version or run a local build
KAIO_GRAFANA_PASSWORD admin Grafana’s admin password
KAIO_PORT, KAIO_GRAFANA_PORT, KAIO_PROMETHEUS_PORT, KAIO_ALERTMANAGER_PORT 8080, 3000, 9090, 9093 The published ports
Terminal window
KAIO_IMAGE=ghcr.io/rgaidot/kaio:0.3.7 KAIO_GRAFANA_PASSWORD=something-else make metrics-full-run

The two stacks are separate Compose projects, kaio-stack here and kaio-monitoring below, and the monitoring one is passed its own monitoring.env rather than the .env that Compose would otherwise hand to both from the shared directory. So neither inherits the other’s volumes or ports, and they run side by side on one host once the port variables move one of them out of the way.

The scrape token lives in two files that must agree: KAIO_METRICS_TOKEN in compose.full.yml and credentials in prometheus.full.yml. Change one and the scrape starts answering 401, so change both or neither.

When Kaio already runs on one host or twenty, compose.yml next to it carries only the monitoring: a Prometheus, an Alertmanager, a Grafana, the same provisioned dashboard and the same alert rules, pointed at the servers you name. One command brings it up:

Terminal window
make metrics-run

The first run creates the four files you fill in, each copied from the example next to it and none of them tracked by git, so your token never lands in a commit and git pull never conflicts with your edits:

File What goes in it
packaging/grafana/monitoring.env KAIO_GRAFANA_PASSWORD, the ports, how long and how large Prometheus may grow
packaging/grafana/token The value of KAIO_METRICS_TOKEN, read by Prometheus as a credentials_file
packaging/grafana/targets.yml Your servers, one entry per Kaio
packaging/grafana/alertmanager.yml Where alerts go

The stack refuses to start while KAIO_GRAFANA_PASSWORD is empty, rather than coming up with a password everyone can read in this repository:

required variable KAIO_GRAFANA_PASSWORD is missing a value: set it in packaging/grafana/.env

Fill the four in and run it again. Grafana is on :3000 with the Kaio dashboard already in its list, Prometheus on :9090 and Alertmanager on :9093. make metrics-logs follows all three, make metrics-stop takes them down.

targets.yml is re-read every 30 seconds, so adding a server needs no restart. The dashboard is provisioned with allowUiUpdates, so a panel you adjust in the browser is saved rather than lost on the next restart, which is also why Grafana is left on its own home page: pointing it at the dashboard file would land you on a copy that stops matching the one you edit. The three images are pinned to exact versions, like the scanner image.

To load the dashboard into a Grafana you already run, import packaging/grafana/dashboard.json and pick your Prometheus.

Maintaining targets.yml by hand duplicates something Kaio already knows: its node registry. /metrics/targets answers it in Prometheus’ own HTTP service discovery format, behind the same flag and the same token as /metrics:

[{"targets":["10.0.0.2:8080"],"labels":{"server":"prod-02"}}]

Point Prometheus at one Kaio and it discovers the rest of the cluster, so enrolling a server with kaio-cli join also puts it under monitoring:

scrape_configs:
- job_name: kaio
metrics_path: /metrics
authorization:
credentials_file: /etc/prometheus/token
static_configs:
- targets: [kaio-01:8080]
http_sd_configs:
- url: http://kaio-01:8080/metrics/targets
authorization:
credentials_file: /etc/prometheus/token
refresh_interval: 60s

That block replaces the file_sd_configs one in prometheus.yml, which is the single tracked file this trades away: targets.yml then stops being read, and the servers come from Kaio instead. The all-in-one stack above needs no such edit, since it has exactly one Kaio to scrape.

The server you ask stays in a static_configs of its own: a registry never lists the machine holding it, so nothing can discover itself. Each discovered target carries a server label holding the name it joined under, deliberately not node: Prometheus gives a target’s own labels precedence, so a target labelled node would overwrite the node label on every kaio_node_* series the scrape returns and rename the real one to exported_node. A node registered over https also carries __scheme__, so it is scraped over TLS rather than answered on the wrong protocol.

A node that stopped answering is still listed, deliberately: Prometheus then reports it as down, where dropping it would make an outage look like a server that was never there.

packaging/grafana/rules.yml carries the nine alerts worth having on day one, and the shipped Prometheus loads them and hands them to the Alertmanager next to it. The shape of each one is the point, so read them as a starting set rather than a default:

- alert: KaioStackPartial
expr: kaio_stack_status{status="partial"} == 1
for: 10m
- alert: KaioStackCriticalCve
expr: kaio_stack_vulnerabilities{severity="critical"} > 0
for: 1h
- alert: KaioNodeDown
expr: kaio_node_up == 0
for: 10m
- alert: KaioPollStalled
expr: time() - kaio_docker_sync_timestamp_seconds > 300
for: 5m

Out of the box they are grouped and visible in the Alertmanager UI and go no further: the shipped alertmanager.yml has a receiver named none with nowhere to send. Replace it with your slack_configs, email_configs or webhook_configs when you want to be told without looking.

KaioPollStalled is the one that watches the watcher. The timestamp only moves on a read that listed and stored the containers, so a Docker socket that stopped answering stops it, and the alert fires even though Kaio itself is still serving pages.

A container the daemon lists but then refuses to describe is a different failure, and has its own counter rather than holding that timestamp back: a container removed between the two calls fails that way routinely, so freezing the health signal on it would cry wolf. Watch it separately, since a host where every container stops being describable empties the state Kaio reports:

- alert: KaioContainersUndescribable
expr: increase(kaio_docker_materialize_failures_total[15m]) > 5
for: 5m

Scraping changes nothing about what Kaio stores. The container_metrics table keeps its fixed 24 hour window and the charts in the web UI keep reading it, so they stay the short-horizon live view. Anything older lives in Prometheus from the moment it starts scraping, which is also why the sparklines and your Grafana panels will disagree about how far back they go.

Metric Type Labels
kaio_containers gauge state
kaio_container_info gauge container, stack, image, id
kaio_container_state gauge container, stack, state
kaio_container_up gauge container, stack
kaio_container_cpu_percent gauge container, stack
kaio_container_memory_bytes gauge container, stack
kaio_container_network_receive_bytes_per_second gauge container, stack
kaio_container_network_transmit_bytes_per_second gauge container, stack
kaio_container_updated_timestamp_seconds gauge container, stack
kaio_container_restarts_total counter container, stack

A container outside any stack carries stack="". kaio_container_info carries the container id, so a recreate retires one series and starts another: an often-redeployed fleet churns more series than the steady count below. kaio_container_state and kaio_stack_status each emit one series, the one holding the current state, so a transition makes the old series go absent rather than drop to zero.

Metric Type Labels
kaio_stacks gauge
kaio_stack_status gauge stack, status
kaio_stack_containers gauge stack, state
kaio_stack_cpu_percent gauge stack
kaio_stack_memory_bytes gauge stack
kaio_stack_network_receive_bytes_per_second gauge stack
kaio_stack_network_transmit_bytes_per_second gauge stack
kaio_stack_version gauge stack
kaio_stack_update_available gauge stack
kaio_stack_outdated_images gauge stack
kaio_stack_env_pending gauge stack
kaio_stack_update_check_timestamp_seconds gauge stack
kaio_stack_update_check_failed gauge stack
kaio_stack_vulnerabilities gauge stack, severity

status is one of running, partial, stopped, paused, empty, the same badge the dashboard shows. A compose project running on the host that Kaio did not deploy has no row of its own, so it reports a status and resource totals but no version, update or environment state.

Metric Type Labels
kaio_vulnerabilities gauge severity
kaio_image_vulnerabilities gauge image, severity

severity is one of critical, high, medium, low, unknown. There is no total: sum the severities you care about. Both cover images in use on the host, as of their last completed scan.

Metric Type Labels
kaio_nodes gauge
kaio_node_up gauge node
kaio_node_status gauge node, status
kaio_node_info gauge node, url, version
kaio_node_last_seen_timestamp_seconds gauge node

These describe the registry of the server being scraped, which does not include itself. In a cluster every server reports the others, so a node that went down is reported by each of its peers.

Metric Type Labels
kaio_build_info gauge version, server_id
kaio_docker_sync_total counter
kaio_docker_sync_failures_total counter
kaio_docker_materialize_failures_total counter
kaio_docker_sync_duration_seconds gauge
kaio_docker_sync_timestamp_seconds gauge
kaio_events_total counter kind, level
kaio_scrape_duration_seconds gauge

The counters start at zero on every restart, which Prometheus reads as a counter reset, so rate() and increase() stay correct across one. kaio_events_total counts events as they are written and is capped at 256 distinct kind/level pairs, which no real deployment reaches.

A scrape walks the current rows, never the metrics history, so it costs six small SQLite reads and no call to the Docker daemon. A read that fails fails the whole scrape, rather than answering with the zeros that would read as no CVEs and no nodes down: Prometheus then marks the target down, which is the answer you want. kaio_scrape_duration_seconds reports what it took, and on a host with a few dozen containers it sits in the low milliseconds. A 15 or 30 second interval is fine.

Series count grows with containers: nine per container, and twenty-five per stack Kaio deployed, against eighteen for a Compose project it only found on the host and therefore reports no version or update state for. Two hundred containers across forty stacks is roughly three thousand series per server, which is small for any Prometheus.

Kaio, built by Régis Gaidot