Health Checks
Three command line tools check the health of a Mina service. Each tool prints one result and sets its exit code, so you can use it as a Kubernetes exec probe, a Docker HEALTHCHECK, or a step in a script.
| Tool | Checks | Reads from | Installed with |
|---|---|---|---|
mina-healthcheck | Daemon | Daemon GraphQL API | mina-mainnet, mina-devnet |
mina-archive-healthcheck | Archive node | Archive PostgreSQL database | mina-archive-mainnet, mina-archive-devnet |
rosetta-healthcheck | Rosetta server | Rosetta API | mina-rosetta-mainnet, mina-rosetta-devnet |
The tools are installed in /usr/local/bin. The mina-daemon, mina-archive, and mina-rosetta Docker images also contain them.
All three tools use the same output rules:
- Exit code
0: the check passed. Exit code1: the check failed. Exit code2(archive and Rosetta tools only): a flag value is not valid, and no request is sent. --json(-j) prints one JSON object on stdout. Without--json, the tool prints one line.- The
waitsubcommand polls until the service is ready or the time limit expires. It writes progress to stderr, so stdout contains only the result.
Daemon: mina-healthcheck
mina-healthcheck queries the daemon GraphQL API. The default endpoint is http://127.0.0.1:3085/graphql.
| Subcommand | Passes when |
|---|---|
daemon-status | The daemon answers. Prints the full daemon status. |
sync-status | The sync status is SYNCED. |
peer-count | The peer count is at least --min-peers. |
chain-length | The local chain length is equal to the highest block length received. |
ready | sync-status, peer-count, and chain-length all pass. |
wait | ready passes before --timeout expires. |
| Flag | Default | Description |
|---|---|---|
--graphql-uri, -u | http://127.0.0.1:3085/graphql | Daemon GraphQL endpoint. |
--min-peers, -n | 2 | Minimum peer count. |
--timeout, -t | 600 | Maximum time in seconds (wait only). |
--interval, -i | 10 | Time between polls in seconds (wait only). |
--json, -j | off | Print JSON. |
mina-healthcheck ready --min-peers 2
mina-healthcheck sync-status --graphql-uri http://my-node:3085/graphql --json
mina-healthcheck wait --timeout 1800 --interval 10
During wait, each poll writes one progress line to stderr:
[ 10s] Bootstrap (peers: 0, chain: ?/?)
[ 20s] Bootstrap (peers: 2, chain: 50/1741)
[ 30s] Catchup (peers: 4, chain: 800/1741)
[ 40s] Synced (peers: 4, chain: 1741/1741)
READY
Do not use sync-status or ready as a liveness probe. A node is not synced during bootstrap and catchup, which can take 30 minutes or more. A liveness probe that fails during this time restarts the node before it can sync. Use daemon-status for liveness and ready for readiness.
Archive node: mina-archive-healthcheck
mina-archive-healthcheck reads the archive database directly. It does not need a running daemon or the Archive Node API. All subcommands except server-ready require --postgres-uri. server-ready requires --server-port.
| Subcommand | Passes when |
|---|---|
db-ready | The database is reachable and the blocks table can be queried. |
server-ready | The archive process accepts TCP connections on --server-port. |
block-height | The database is reachable. Prints the highest block height. |
block-recency | The newest block is not older than --max-delay seconds. |
missing-blocks | The number of missing heights in the last --window blocks is not more than --max-missing. |
unparented-blocks | The number of blocks without a parent is not more than --max-unparented. |
ready | All database checks pass. |
wait | ready passes before --timeout expires. |
| Flag | Default | Description |
|---|---|---|
--postgres-uri, -p | PostgreSQL connection URI. Required, except for server-ready. | |
--max-delay | 360 | Maximum age of the newest block in seconds. |
--max-missing | 10 | Maximum number of missing blocks. |
--max-unparented | 5 | Maximum number of blocks without a parent. |
--window | 2000 | Number of blocks that missing-blocks examines. |
--server-port | Archive server port. Required for server-ready. Optional for wait: wait then also requires that the port accepts connections. | |
--server-host | 127.0.0.1 | Archive server host for server-ready and wait. |
--db-only | off | wait only: wait for the database and skip the recency, missing, and unparented checks. |
--timeout, -t | 600 | Maximum time in seconds (wait only). |
--interval, -i | 10 | Time between polls in seconds (wait only). |
--json, -j | off | Print JSON. |
mina-archive-healthcheck ready --postgres-uri postgres://user@localhost:5432/archive --max-delay 360
mina-archive-healthcheck missing-blocks --postgres-uri postgres://user@localhost:5432/archive --max-missing 10
The database answers queries before the archive process listens for blocks. A block that the daemon sends before the archive process listens is lost. To wait until the archive can receive blocks, check the database and the server port together:
mina-archive-healthcheck wait --db-only --server-port 3086 \
--postgres-uri postgres://user@localhost:5432/archive
A failed check still reports the values that it measured. The error field gives the reason: Db_unreachable, Db_query_failed, No_blocks, Invalid_timestamp, Thresholds_exceeded, or Timed_out.
{ "healthy": false, "missing_blocks": 41, "max_missing": 10, "window": 2000,
"error": { "sexp": [ "Thresholds_exceeded", [ "missing blocks 41 > 10" ] ] } }
These probes are fast checks of the current state. For a full check of archive data integrity, use mina-missing-blocks-auditor. See Backfilling Missing Blocks.
Rosetta: rosetta-healthcheck
rosetta-healthcheck calls the Rosetta API. The default endpoint is http://localhost:3087.
| Subcommand | Rosetta endpoint | Passes when |
|---|---|---|
connectivity | /network/list | The server lists the network given in --network. |
tip-recency | /network/status | The tip block is not older than --max-age seconds. |
ready | All of the above and /network/options | All three checks pass. |
wait | Polls ready | ready passes before --deadline expires. |
| Flag | Default | Environment variable | Description |
|---|---|---|---|
--rosetta-uri | http://localhost:3087 | MINA_ROSETTA_URI | Rosetta server URL. |
--network | testnet | MINA_ROSETTA_NETWORK | Network name that the server must list. |
--blockchain | mina | MINA_ROSETTA_BLOCKCHAIN | Blockchain name. |
--max-age | 360 | Maximum age of the tip block in seconds. | |
--timeout | 5 | Maximum time in seconds for one request. | |
--deadline, -d | 600 | Maximum time in seconds (wait only). | |
--interval, -i | 10 | Time between polls in seconds (wait only). | |
--json, -j | off | Print JSON. |
A flag has priority over its environment variable.
Set --network to mainnet or devnet. The Rosetta server lists the network of its daemon, so the default value testnet does not match Mainnet or Devnet.
rosetta-healthcheck ready --network mainnet --max-age 360 --json
rosetta-healthcheck wait --network devnet --deadline 600 --interval 10
The same packages contain rosetta-client, which sends one Rosetta API request and prints the response. Use it to examine a Rosetta server manually, for example rosetta-client network status.
Examples
The probes run inside the container that they check. The daemon binds its GraphQL API to 127.0.0.1:3085, so mina-healthcheck in the same container needs no flags.
A Kubernetes exec probe runs its command without a shell, so it does not expand environment variables. To pass a value from the environment, for example a database password, run the probe with sh -c.
Set timeoutSeconds on each probe to more than the time that the probe can take. The default is 1 second.
Node operator
Kubernetes container for a daemon or a block producer. The startup probe allows up to 15 minutes for the daemon to start its GraphQL API. Liveness checks only that the daemon answers. Readiness checks that the node is synced and connected, so a Service sends traffic only to synced nodes.
containers:
- name: mina
image: minaprotocol/mina-daemon:<VERSION>-bookworm-mainnet
args: ["daemon", "--peer-list-url", "https://bootnodes.minaprotocol.com/networks/mainnet.txt"]
startupProbe:
exec:
command: ["mina-healthcheck", "daemon-status"]
periodSeconds: 10
failureThreshold: 90
timeoutSeconds: 10
livenessProbe:
exec:
command: ["mina-healthcheck", "daemon-status"]
periodSeconds: 30
failureThreshold: 5
timeoutSeconds: 10
readinessProbe:
exec:
command: ["mina-healthcheck", "ready", "--min-peers", "5"]
periodSeconds: 30
timeoutSeconds: 10
On a host with systemd, check the node from a timer or a cron job and send an alert when the check fails:
#!/bin/sh
# /usr/local/bin/mina-check: alert when the node is not synced and connected
if ! out=$(mina-healthcheck ready --min-peers 5 2>&1); then
echo "mina node not ready: $out" | mail -s "mina alert $(hostname)" ops@example.com
fi
Wait until the node is synced before a step that needs it, for example after an upgrade or before the next node in a rolling restart:
sudo systemctl restart mina # or: systemctl --user restart mina
mina-healthcheck wait --timeout 3600 --interval 30
With Docker Compose:
services:
mina-node:
image: minaprotocol/mina-daemon:<VERSION>-bookworm-mainnet
healthcheck:
test: ["CMD", "mina-healthcheck", "ready", "--min-peers", "5"]
interval: 30s
timeout: 10s
retries: 3
start_period: 60m
start_period gives the node time to sync before Docker marks it as unhealthy.
Archive node operator
The archive process must accept connections before the daemon sends it blocks: a block that the daemon sends before that is lost. An init container in the daemon pod waits for the archive server. Probes in the archive pod check the database and the age of the newest block.
# Archive pod
containers:
- name: archive
image: minaprotocol/mina-archive:<VERSION>-bookworm-mainnet
env:
- name: ARCHIVE_URI
valueFrom:
secretKeyRef: { name: archive-db, key: uri } # postgres://user:password@postgres:5432/archive
livenessProbe:
exec:
command: ["sh", "-c", "mina-archive-healthcheck db-ready --postgres-uri \"$ARCHIVE_URI\""]
periodSeconds: 30
timeoutSeconds: 10
readinessProbe:
exec:
command: ["sh", "-c", "mina-archive-healthcheck ready --postgres-uri \"$ARCHIVE_URI\" --max-delay 600"]
periodSeconds: 60
timeoutSeconds: 30
# Daemon pod: start the daemon only when the archive server listens
initContainers:
- name: wait-for-archive
image: minaprotocol/mina-archive:<VERSION>-bookworm-mainnet
env:
- name: ARCHIVE_URI
valueFrom:
secretKeyRef: { name: archive-db, key: uri }
command: ["sh", "-c", "mina-archive-healthcheck wait --db-only --server-host archive --server-port 3086 --postgres-uri \"$ARCHIVE_URI\" --timeout 600"]
Check the completeness of the archive at regular intervals and alert when blocks are missing:
apiVersion: batch/v1
kind: CronJob
metadata:
name: archive-missing-blocks
spec:
schedule: "*/10 * * * *"
jobTemplate:
spec:
template:
spec:
restartPolicy: Never
containers:
- name: check
image: minaprotocol/mina-archive:<VERSION>-bookworm-mainnet
env:
- name: ARCHIVE_URI
valueFrom:
secretKeyRef: { name: archive-db, key: uri }
command: ["sh", "-c", "mina-archive-healthcheck missing-blocks --postgres-uri \"$ARCHIVE_URI\" --max-missing 0 --json"]
A failed job means that blocks are missing. The JSON output gives the number in missing_blocks. To fill the gaps, see Backfilling Missing Blocks.
Exchange operator
An exchange that uses Rosetta depends on the daemon, the archive node, and the Rosetta server. Stop deposit and withdrawal processing when Rosetta is not ready or its newest block is old, and start it again when the check passes:
#!/bin/sh
export MINA_ROSETTA_URI=http://localhost:3087
export MINA_ROSETTA_NETWORK=mainnet
if ! rosetta-healthcheck ready --max-age 600 --json > /tmp/rosetta-health.json; then
echo "Rosetta not ready, pausing deposits and withdrawals" >&2
cat /tmp/rosetta-health.json >&2
exit 1
fi
# Height of the newest block that Rosetta serves
rosetta-client network status --compact | jq '.current_block_identifier.index'
MINA_ROSETTA_URI and MINA_ROSETTA_NETWORK apply to rosetta-healthcheck and rosetta-client.
To check each component separately, run mina-healthcheck ready for the daemon and mina-archive-healthcheck ready for the archive database. A Rosetta failure then shows which component causes it.
Kubernetes probes for the Rosetta container:
livenessProbe:
exec:
command: ["rosetta-healthcheck", "connectivity", "--network", "mainnet"]
periodSeconds: 30
timeoutSeconds: 10
readinessProbe:
exec:
command: ["rosetta-healthcheck", "ready", "--network", "mainnet", "--max-age", "600"]
periodSeconds: 30
timeoutSeconds: 10