Prometheus and Grafana
Kaio exposes its own state as Prometheus metrics at /metrics, on the same port
as the API and the web UI. Prometheus keeps the history Kaio deliberately does
not, Grafana draws it, and Alertmanager tells you about it: “the media stack
has been partial for ten minutes”, “prod-02 has not answered a probe”,
“a critical CVE appeared on something in production”.
The endpoint is off until you turn it on.
Turning it on
Section titled “Turning it on”KAIO_METRICS=onKAIO_METRICS_TOKEN=a-long-random-stringThe server logs what it settled on at startup:
INFO kaio_server: Serving Prometheus metrics at /metrics, behind a bearer tokenThe token is optional and strongly recommended: /metrics lists your stack
names, your image names and your CVE counts, and the REST API in front of it has
no authentication of its own. With the token set, a scrape must carry it:
curl -H "Authorization: Bearer a-long-random-string" http://kaio.example.net:8080/metricsWhat it is for
Section titled “What it is for”Container CPU and memory are the part cAdvisor already does better, and they are exposed here because the numbers were already in hand. The reason to scrape Kaio is the rest: what Kaio knows and nothing else on the host does.
| Metric | The question it answers |
|---|---|
kaio_stack_status{stack,status} |
Which stacks are not fully up |
kaio_stack_update_available{stack} |
Which stacks have an image update waiting, and for how long |
kaio_stack_env_pending{stack} |
Which stacks hold environment changes they were never redeployed with |
kaio_stack_vulnerabilities{stack,severity} |
Which stacks run a critical CVE |
kaio_stack_version{stack} |
Which compose version is live, so a rollback shows up as a step down |
kaio_node_up{node} |
Which servers in the cluster have stopped answering |
kaio_docker_sync_timestamp_seconds |
Whether Kaio itself is still reading the daemon |
Scraping
Section titled “Scraping”Each Kaio exposes its own host, so point Prometheus at every server rather than
at one of them. The gateway (/api/nodes/{name}/proxy/) is deliberately not the
route to use: it would mix two hosts under one instance label and make the
control plane a single point of failure for your monitoring.
scrape_configs: - job_name: kaio metrics_path: /metrics authorization: credentials: a-long-random-string static_configs: - targets: - kaio-01:8080 - kaio-02:8080Every metric carries Prometheus’ own instance label, so sum by (instance)
splits any of them per server and the shipped dashboard has a Server picker
built on it.
Everything in one command
Section titled “Everything in one command”If Kaio is not running yet, or you want the whole thing on one host to see what
it looks like, compose.full.yml carries Kaio, Prometheus, Alertmanager and
Grafana together, already wired to each other:
make metrics-full-runNothing to fill in first. Kaio comes up with KAIO_METRICS=on, Prometheus
scrapes it by service name over the compose network, the alert rules are loaded,
and Grafana opens on :3000 with admin / admin and the Kaio dashboard in
its list. make metrics-full-logs follows all four, make metrics-full-stop
takes them down.
The values are in the file in plain text, which is the point of this path: it is meant to be read and run. A few knobs are worth knowing:
| Variable | Default | What it does |
|---|---|---|
KAIO_IMAGE |
ghcr.io/rgaidot/kaio:latest |
The Kaio image, so you can pin a version or run a local build |
KAIO_GRAFANA_PASSWORD |
admin |
Grafana’s admin password |
KAIO_PORT, KAIO_GRAFANA_PORT, KAIO_PROMETHEUS_PORT, KAIO_ALERTMANAGER_PORT |
8080, 3000, 9090, 9093 |
The published ports |
KAIO_IMAGE=ghcr.io/rgaidot/kaio:0.3.7 KAIO_GRAFANA_PASSWORD=something-else make metrics-full-runThe two stacks are separate Compose projects, kaio-stack here and
kaio-monitoring below, and the monitoring one is passed its own
monitoring.env rather than the .env that Compose would otherwise hand to
both from the shared directory. So neither inherits the other’s volumes or
ports, and they run side by side on one host once the port variables move one of
them out of the way.
The scrape token lives in two files that must agree:
KAIO_METRICS_TOKEN in compose.full.yml and credentials in
prometheus.full.yml. Change one and the scrape starts answering 401, so
change both or neither.
Monitoring servers you already run
Section titled “Monitoring servers you already run”When Kaio already runs on one host or twenty, compose.yml next to it carries
only the monitoring: a Prometheus, an Alertmanager, a Grafana, the same
provisioned dashboard and the same alert rules, pointed at the servers you name.
One command brings it up:
make metrics-runThe first run creates the four files you fill in, each copied from the example
next to it and none of them tracked by git, so your token never lands in a
commit and git pull never conflicts with your edits:
| File | What goes in it |
|---|---|
packaging/grafana/monitoring.env |
KAIO_GRAFANA_PASSWORD, the ports, how long and how large Prometheus may grow |
packaging/grafana/token |
The value of KAIO_METRICS_TOKEN, read by Prometheus as a credentials_file |
packaging/grafana/targets.yml |
Your servers, one entry per Kaio |
packaging/grafana/alertmanager.yml |
Where alerts go |
The stack refuses to start while KAIO_GRAFANA_PASSWORD is empty, rather
than coming up with a password everyone can read in this repository:
required variable KAIO_GRAFANA_PASSWORD is missing a value: set it in packaging/grafana/.envFill the four in and run it again. Grafana is on :3000 with the Kaio
dashboard already in its list, Prometheus on :9090 and Alertmanager on
:9093. make metrics-logs follows all three, make metrics-stop takes them
down.
targets.yml is re-read every 30 seconds, so adding a server needs no restart.
The dashboard is provisioned with allowUiUpdates, so a panel you adjust in the
browser is saved rather than lost on the next restart, which is also why Grafana
is left on its own home page: pointing it at the dashboard file would land
you on a copy that stops matching the one you edit. The three images are pinned
to exact versions, like the scanner image.
To load the dashboard into a Grafana you already run, import
packaging/grafana/dashboard.json and pick your Prometheus.
Letting Kaio name your servers
Section titled “Letting Kaio name your servers”Maintaining targets.yml by hand duplicates something Kaio already knows: its
node registry. /metrics/targets answers it in Prometheus’ own
HTTP service discovery
format, behind the same flag and the same token as /metrics:
[{"targets":["10.0.0.2:8080"],"labels":{"server":"prod-02"}}]Point Prometheus at one Kaio and it discovers the rest of the cluster, so
enrolling a server with kaio-cli join also puts it under monitoring:
scrape_configs: - job_name: kaio metrics_path: /metrics authorization: credentials_file: /etc/prometheus/token static_configs: - targets: [kaio-01:8080] http_sd_configs: - url: http://kaio-01:8080/metrics/targets authorization: credentials_file: /etc/prometheus/token refresh_interval: 60sThat block replaces the file_sd_configs one in prometheus.yml, which is the
single tracked file this trades away: targets.yml then stops being read, and
the servers come from Kaio instead. The all-in-one stack above needs no such
edit, since it has exactly one Kaio to scrape.
The server you ask stays in a static_configs of its own: a registry never
lists the machine holding it, so nothing can discover itself. Each discovered
target carries a server label holding the name it joined under, deliberately
not node: Prometheus gives a target’s own labels precedence, so a target
labelled node would overwrite the node label on every kaio_node_* series
the scrape returns and rename the real one to exported_node. A node registered
over https also carries __scheme__, so it is scraped over TLS rather than
answered on the wrong protocol.
A node that stopped answering is still listed, deliberately: Prometheus then
reports it as down, where dropping it would make an outage look like a server
that was never there.
Alerting
Section titled “Alerting”packaging/grafana/rules.yml carries the nine alerts worth having on day one,
and the shipped Prometheus loads them and hands them to the Alertmanager next to
it. The shape of each one is the point, so read them as a starting set rather
than a default:
- alert: KaioStackPartial expr: kaio_stack_status{status="partial"} == 1 for: 10m
- alert: KaioStackCriticalCve expr: kaio_stack_vulnerabilities{severity="critical"} > 0 for: 1h
- alert: KaioNodeDown expr: kaio_node_up == 0 for: 10m
- alert: KaioPollStalled expr: time() - kaio_docker_sync_timestamp_seconds > 300 for: 5mOut of the box they are grouped and visible in the Alertmanager UI and go no
further: the shipped alertmanager.yml has a receiver named none with nowhere
to send. Replace it with your slack_configs, email_configs or
webhook_configs when you want to be told without looking.
KaioPollStalled is the one that watches the watcher. The timestamp only moves
on a read that listed and stored the containers, so a Docker socket that stopped
answering stops it, and the alert fires even though Kaio itself is still serving
pages.
A container the daemon lists but then refuses to describe is a different failure, and has its own counter rather than holding that timestamp back: a container removed between the two calls fails that way routinely, so freezing the health signal on it would cry wolf. Watch it separately, since a host where every container stops being describable empties the state Kaio reports:
- alert: KaioContainersUndescribable expr: increase(kaio_docker_materialize_failures_total[15m]) > 5 for: 5mRetention
Section titled “Retention”Scraping changes nothing about what Kaio stores. The container_metrics table
keeps its fixed 24 hour window and the charts in the
web UI keep reading it, so they stay the short-horizon live view. Anything older
lives in Prometheus from the moment it starts scraping, which is also why the
sparklines and your Grafana panels will disagree about how far back they go.
Every metric
Section titled “Every metric”Containers
Section titled “Containers”| Metric | Type | Labels |
|---|---|---|
kaio_containers |
gauge | state |
kaio_container_info |
gauge | container, stack, image, id |
kaio_container_state |
gauge | container, stack, state |
kaio_container_up |
gauge | container, stack |
kaio_container_cpu_percent |
gauge | container, stack |
kaio_container_memory_bytes |
gauge | container, stack |
kaio_container_network_receive_bytes_per_second |
gauge | container, stack |
kaio_container_network_transmit_bytes_per_second |
gauge | container, stack |
kaio_container_updated_timestamp_seconds |
gauge | container, stack |
kaio_container_restarts_total |
counter | container, stack |
A container outside any stack carries stack="". kaio_container_info carries
the container id, so a recreate retires one series and starts another: an
often-redeployed fleet churns more series than the steady count below. kaio_container_state and
kaio_stack_status each emit one series, the one holding the current state,
so a transition makes the old series go absent rather than drop to zero.
Stacks
Section titled “Stacks”| Metric | Type | Labels |
|---|---|---|
kaio_stacks |
gauge | |
kaio_stack_status |
gauge | stack, status |
kaio_stack_containers |
gauge | stack, state |
kaio_stack_cpu_percent |
gauge | stack |
kaio_stack_memory_bytes |
gauge | stack |
kaio_stack_network_receive_bytes_per_second |
gauge | stack |
kaio_stack_network_transmit_bytes_per_second |
gauge | stack |
kaio_stack_version |
gauge | stack |
kaio_stack_update_available |
gauge | stack |
kaio_stack_outdated_images |
gauge | stack |
kaio_stack_env_pending |
gauge | stack |
kaio_stack_update_check_timestamp_seconds |
gauge | stack |
kaio_stack_update_check_failed |
gauge | stack |
kaio_stack_vulnerabilities |
gauge | stack, severity |
status is one of running, partial, stopped, paused, empty, the same
badge the dashboard shows. A compose project running on the host that Kaio did
not deploy has no row of its own, so it reports a status and resource totals but
no version, update or environment state.
Vulnerabilities
Section titled “Vulnerabilities”| Metric | Type | Labels |
|---|---|---|
kaio_vulnerabilities |
gauge | severity |
kaio_image_vulnerabilities |
gauge | image, severity |
severity is one of critical, high, medium, low, unknown. There is no
total: sum the severities you care about. Both cover images in use on the
host, as of their last completed scan.
Cluster
Section titled “Cluster”| Metric | Type | Labels |
|---|---|---|
kaio_nodes |
gauge | |
kaio_node_up |
gauge | node |
kaio_node_status |
gauge | node, status |
kaio_node_info |
gauge | node, url, version |
kaio_node_last_seen_timestamp_seconds |
gauge | node |
These describe the registry of the server being scraped, which does not include itself. In a cluster every server reports the others, so a node that went down is reported by each of its peers.
Kaio itself
Section titled “Kaio itself”| Metric | Type | Labels |
|---|---|---|
kaio_build_info |
gauge | version, server_id |
kaio_docker_sync_total |
counter | |
kaio_docker_sync_failures_total |
counter | |
kaio_docker_materialize_failures_total |
counter | |
kaio_docker_sync_duration_seconds |
gauge | |
kaio_docker_sync_timestamp_seconds |
gauge | |
kaio_events_total |
counter | kind, level |
kaio_scrape_duration_seconds |
gauge |
The counters start at zero on every restart, which Prometheus reads as a counter
reset, so rate() and increase() stay correct across one. kaio_events_total
counts events as they are written and is capped at 256 distinct kind/level
pairs, which no real deployment reaches.
A scrape walks the current rows, never the metrics history, so it costs six
small SQLite reads and no call to the Docker daemon. A read that fails fails
the whole scrape, rather than answering with the zeros that would read as no
CVEs and no nodes down: Prometheus then marks the target down, which is the
answer you want. kaio_scrape_duration_seconds
reports what it took, and on a host with a few dozen containers it sits in the
low milliseconds. A 15 or 30 second interval is fine.
Series count grows with containers: nine per container, and twenty-five per stack Kaio deployed, against eighteen for a Compose project it only found on the host and therefore reports no version or update state for. Two hundred containers across forty stacks is roughly three thousand series per server, which is small for any Prometheus.
Kaio, built by Régis Gaidot