Skip to main content

Health Checks

Three command line tools check the health of a Mina service. Each tool prints one result and sets its exit code, so you can use it as a Kubernetes exec probe, a Docker HEALTHCHECK, or a step in a script.

ToolChecksReads fromInstalled with
mina-healthcheckDaemonDaemon GraphQL APImina-mainnet, mina-devnet
mina-archive-healthcheckArchive nodeArchive PostgreSQL databasemina-archive-mainnet, mina-archive-devnet
rosetta-healthcheckRosetta serverRosetta APImina-rosetta-mainnet, mina-rosetta-devnet

The tools are installed in /usr/local/bin. The mina-daemon, mina-archive, and mina-rosetta Docker images also contain them.

All three tools use the same output rules:

  • Exit code 0: the check passed. Exit code 1: the check failed. Exit code 2 (archive and Rosetta tools only): a flag value is not valid, and no request is sent.
  • --json (-j) prints one JSON object on stdout. Without --json, the tool prints one line.
  • The wait subcommand polls until the service is ready or the time limit expires. It writes progress to stderr, so stdout contains only the result.

Daemon: mina-healthcheck​

mina-healthcheck queries the daemon GraphQL API. The default endpoint is http://127.0.0.1:3085/graphql.

SubcommandPasses when
daemon-statusThe daemon answers. Prints the full daemon status.
sync-statusThe sync status is SYNCED.
peer-countThe peer count is at least --min-peers.
chain-lengthThe local chain length is equal to the highest block length received.
readysync-status, peer-count, and chain-length all pass.
waitready passes before --timeout expires.
FlagDefaultDescription
--graphql-uri, -uhttp://127.0.0.1:3085/graphqlDaemon GraphQL endpoint.
--min-peers, -n2Minimum peer count.
--timeout, -t600Maximum time in seconds (wait only).
--interval, -i10Time between polls in seconds (wait only).
--json, -joffPrint JSON.
mina-healthcheck ready --min-peers 2
mina-healthcheck sync-status --graphql-uri http://my-node:3085/graphql --json
mina-healthcheck wait --timeout 1800 --interval 10

During wait, each poll writes one progress line to stderr:

[  10s] Bootstrap   (peers: 0, chain: ?/?)
[ 20s] Bootstrap (peers: 2, chain: 50/1741)
[ 30s] Catchup (peers: 4, chain: 800/1741)
[ 40s] Synced (peers: 4, chain: 1741/1741)
READY
Liveness probes

Do not use sync-status or ready as a liveness probe. A node is not synced during bootstrap and catchup, which can take 30 minutes or more. A liveness probe that fails during this time restarts the node before it can sync. Use daemon-status for liveness and ready for readiness.

Archive node: mina-archive-healthcheck​

mina-archive-healthcheck reads the archive database directly. It does not need a running daemon or the Archive Node API. All subcommands except server-ready require --postgres-uri. server-ready requires --server-port.

SubcommandPasses when
db-readyThe database is reachable and the blocks table can be queried.
server-readyThe archive process accepts TCP connections on --server-port.
block-heightThe database is reachable. Prints the highest block height.
block-recencyThe newest block is not older than --max-delay seconds.
missing-blocksThe number of missing heights in the last --window blocks is not more than --max-missing.
unparented-blocksThe number of blocks without a parent is not more than --max-unparented.
readyAll database checks pass.
waitready passes before --timeout expires.
FlagDefaultDescription
--postgres-uri, -pPostgreSQL connection URI. Required, except for server-ready.
--max-delay360Maximum age of the newest block in seconds.
--max-missing10Maximum number of missing blocks.
--max-unparented5Maximum number of blocks without a parent.
--window2000Number of blocks that missing-blocks examines.
--server-portArchive server port. Required for server-ready. Optional for wait: wait then also requires that the port accepts connections.
--server-host127.0.0.1Archive server host for server-ready and wait.
--db-onlyoffwait only: wait for the database and skip the recency, missing, and unparented checks.
--timeout, -t600Maximum time in seconds (wait only).
--interval, -i10Time between polls in seconds (wait only).
--json, -joffPrint JSON.
mina-archive-healthcheck ready --postgres-uri postgres://user@localhost:5432/archive --max-delay 360
mina-archive-healthcheck missing-blocks --postgres-uri postgres://user@localhost:5432/archive --max-missing 10

The database answers queries before the archive process listens for blocks. A block that the daemon sends before the archive process listens is lost. To wait until the archive can receive blocks, check the database and the server port together:

mina-archive-healthcheck wait --db-only --server-port 3086 \
--postgres-uri postgres://user@localhost:5432/archive

A failed check still reports the values that it measured. The error field gives the reason: Db_unreachable, Db_query_failed, No_blocks, Invalid_timestamp, Thresholds_exceeded, or Timed_out.

{ "healthy": false, "missing_blocks": 41, "max_missing": 10, "window": 2000,
"error": { "sexp": [ "Thresholds_exceeded", [ "missing blocks 41 > 10" ] ] } }

These probes are fast checks of the current state. For a full check of archive data integrity, use mina-missing-blocks-auditor. See Backfilling Missing Blocks.

Rosetta: rosetta-healthcheck​

rosetta-healthcheck calls the Rosetta API. The default endpoint is http://localhost:3087.

SubcommandRosetta endpointPasses when
connectivity/network/listThe server lists the network given in --network.
tip-recency/network/statusThe tip block is not older than --max-age seconds.
readyAll of the above and /network/optionsAll three checks pass.
waitPolls readyready passes before --deadline expires.
FlagDefaultEnvironment variableDescription
--rosetta-urihttp://localhost:3087MINA_ROSETTA_URIRosetta server URL.
--networktestnetMINA_ROSETTA_NETWORKNetwork name that the server must list.
--blockchainminaMINA_ROSETTA_BLOCKCHAINBlockchain name.
--max-age360Maximum age of the tip block in seconds.
--timeout5Maximum time in seconds for one request.
--deadline, -d600Maximum time in seconds (wait only).
--interval, -i10Time between polls in seconds (wait only).
--json, -joffPrint JSON.

A flag has priority over its environment variable.

note

Set --network to mainnet or devnet. The Rosetta server lists the network of its daemon, so the default value testnet does not match Mainnet or Devnet.

rosetta-healthcheck ready --network mainnet --max-age 360 --json
rosetta-healthcheck wait --network devnet --deadline 600 --interval 10

The same packages contain rosetta-client, which sends one Rosetta API request and prints the response. Use it to examine a Rosetta server manually, for example rosetta-client network status.

Examples​

The probes run inside the container that they check. The daemon binds its GraphQL API to 127.0.0.1:3085, so mina-healthcheck in the same container needs no flags.

A Kubernetes exec probe runs its command without a shell, so it does not expand environment variables. To pass a value from the environment, for example a database password, run the probe with sh -c.

Set timeoutSeconds on each probe to more than the time that the probe can take. The default is 1 second.

Node operator​

Kubernetes container for a daemon or a block producer. The startup probe allows up to 15 minutes for the daemon to start its GraphQL API. Liveness checks only that the daemon answers. Readiness checks that the node is synced and connected, so a Service sends traffic only to synced nodes.

containers:
- name: mina
image: minaprotocol/mina-daemon:<VERSION>-bookworm-mainnet
args: ["daemon", "--peer-list-url", "https://bootnodes.minaprotocol.com/networks/mainnet.txt"]
startupProbe:
exec:
command: ["mina-healthcheck", "daemon-status"]
periodSeconds: 10
failureThreshold: 90
timeoutSeconds: 10
livenessProbe:
exec:
command: ["mina-healthcheck", "daemon-status"]
periodSeconds: 30
failureThreshold: 5
timeoutSeconds: 10
readinessProbe:
exec:
command: ["mina-healthcheck", "ready", "--min-peers", "5"]
periodSeconds: 30
timeoutSeconds: 10

On a host with systemd, check the node from a timer or a cron job and send an alert when the check fails:

#!/bin/sh
# /usr/local/bin/mina-check: alert when the node is not synced and connected
if ! out=$(mina-healthcheck ready --min-peers 5 2>&1); then
echo "mina node not ready: $out" | mail -s "mina alert $(hostname)" ops@example.com
fi

Wait until the node is synced before a step that needs it, for example after an upgrade or before the next node in a rolling restart:

sudo systemctl restart mina   # or: systemctl --user restart mina
mina-healthcheck wait --timeout 3600 --interval 30

With Docker Compose:

services:
mina-node:
image: minaprotocol/mina-daemon:<VERSION>-bookworm-mainnet
healthcheck:
test: ["CMD", "mina-healthcheck", "ready", "--min-peers", "5"]
interval: 30s
timeout: 10s
retries: 3
start_period: 60m

start_period gives the node time to sync before Docker marks it as unhealthy.

Archive node operator​

The archive process must accept connections before the daemon sends it blocks: a block that the daemon sends before that is lost. An init container in the daemon pod waits for the archive server. Probes in the archive pod check the database and the age of the newest block.

# Archive pod
containers:
- name: archive
image: minaprotocol/mina-archive:<VERSION>-bookworm-mainnet
env:
- name: ARCHIVE_URI
valueFrom:
secretKeyRef: { name: archive-db, key: uri } # postgres://user:password@postgres:5432/archive
livenessProbe:
exec:
command: ["sh", "-c", "mina-archive-healthcheck db-ready --postgres-uri \"$ARCHIVE_URI\""]
periodSeconds: 30
timeoutSeconds: 10
readinessProbe:
exec:
command: ["sh", "-c", "mina-archive-healthcheck ready --postgres-uri \"$ARCHIVE_URI\" --max-delay 600"]
periodSeconds: 60
timeoutSeconds: 30
# Daemon pod: start the daemon only when the archive server listens
initContainers:
- name: wait-for-archive
image: minaprotocol/mina-archive:<VERSION>-bookworm-mainnet
env:
- name: ARCHIVE_URI
valueFrom:
secretKeyRef: { name: archive-db, key: uri }
command: ["sh", "-c", "mina-archive-healthcheck wait --db-only --server-host archive --server-port 3086 --postgres-uri \"$ARCHIVE_URI\" --timeout 600"]

Check the completeness of the archive at regular intervals and alert when blocks are missing:

apiVersion: batch/v1
kind: CronJob
metadata:
name: archive-missing-blocks
spec:
schedule: "*/10 * * * *"
jobTemplate:
spec:
template:
spec:
restartPolicy: Never
containers:
- name: check
image: minaprotocol/mina-archive:<VERSION>-bookworm-mainnet
env:
- name: ARCHIVE_URI
valueFrom:
secretKeyRef: { name: archive-db, key: uri }
command: ["sh", "-c", "mina-archive-healthcheck missing-blocks --postgres-uri \"$ARCHIVE_URI\" --max-missing 0 --json"]

A failed job means that blocks are missing. The JSON output gives the number in missing_blocks. To fill the gaps, see Backfilling Missing Blocks.

Exchange operator​

An exchange that uses Rosetta depends on the daemon, the archive node, and the Rosetta server. Stop deposit and withdrawal processing when Rosetta is not ready or its newest block is old, and start it again when the check passes:

#!/bin/sh
export MINA_ROSETTA_URI=http://localhost:3087
export MINA_ROSETTA_NETWORK=mainnet

if ! rosetta-healthcheck ready --max-age 600 --json > /tmp/rosetta-health.json; then
echo "Rosetta not ready, pausing deposits and withdrawals" >&2
cat /tmp/rosetta-health.json >&2
exit 1
fi

# Height of the newest block that Rosetta serves
rosetta-client network status --compact | jq '.current_block_identifier.index'

MINA_ROSETTA_URI and MINA_ROSETTA_NETWORK apply to rosetta-healthcheck and rosetta-client.

To check each component separately, run mina-healthcheck ready for the daemon and mina-archive-healthcheck ready for the archive database. A Rosetta failure then shows which component causes it.

Kubernetes probes for the Rosetta container:

livenessProbe:
exec:
command: ["rosetta-healthcheck", "connectivity", "--network", "mainnet"]
periodSeconds: 30
timeoutSeconds: 10
readinessProbe:
exec:
command: ["rosetta-healthcheck", "ready", "--network", "mainnet", "--max-age", "600"]
periodSeconds: 30
timeoutSeconds: 10