feat: supervise the daemon, and restart one that has lost every peer
Build docker image and push to registry.bitdeals.org / main-build-job (push) Successful in 7m12s

A PyBitmessage daemon that has dropped to zero network connections does not find
its way back. It can sit there for days — testnet1 did, and every BitDeals deal
completed on that stand since 2026-08-30 left its escrow unspent, because the
guarantor never received a CH3 the node could no longer publish. A daemon that
has just started, by contrast, dials hard and reconnects within seconds. The
cure was already known; what was missing was anything to notice and apply it.

Docker will not: it reacts to a process exiting, never to a healthcheck, and
health-driven restarts exist only under Swarm. Doing it from outside means
handing a container the docker socket, which is root on the host and a poor
trade for a relay with a published port. So the container supervises itself.

`run.sh` no longer ends at `exec gosu bitmessage pybitmessage -d`. That made the
daemon PID 1, and it is a bad PID 1: daemonize() double-forks and parks the
grandfather in `while True: time.sleep(1)`, the final child SIGTERMs it to say
"ready", and PID 1 drops that signal for want of a handler. The grandfather slept
for ever; `docker stop` therefore reached the real daemon only as the SIGKILL ten
seconds later, cutting a startup VACUUM in half — the way a node gets trapped
retrying one it can never finish; and a daemon that died on its own left the
container Up around a corpse, because what PID 1 was doing had nothing to do with
whether the daemon lived.

Away from PID 1 that grandfather does die on the ready signal — measured here:
the start call returns at once with status 143 and leaves exactly one
pybitmessage process behind. That makes starting the daemon an ordinary blocking
call, and the supervisor ordinary shell: a trap that stops the daemon through its
own API, a restart when it is gone, and the peer rule.

What the supervisor does not do is act on a daemon whose API is not answering at
all. That is the trapped-VACUUM node; a restart does not cure it and cuts the
next VACUUM short as well. watchdog.py reports that case as its own exit code so
the loop can leave it to a person. The healthcheck is untouched: it reports, and
does not act.

Four settings, on by default: BITMESSAGE_WATCHDOG, and _PERIOD, _AFTER,
_COOLDOWN. All validated at start, where a typo is visible, rather than hours
later as a supervisor that spins or one that never acts.

Callers must raise the stop grace period — 90s in compose, or --stop-timeout 90.
A clean shutdown took 17.5 s on a small database and grows with it, so under
Docker's default ten the daemon is killed mid-write anyway and the supervisor
buys nothing. An image cannot set this for itself.
This commit is contained in:
2026-09-04 09:53:07 +00:00
parent 58e3f22982
commit 88b5b896e1
5 changed files with 309 additions and 1 deletions
+1
View File
@@ -59,6 +59,7 @@ ENV BITMESSAGE_HOME=${HOME}
COPY --from=0 /usr/local/ /usr/local/
COPY ./docker/healthy_check.py /usr/local/bin/
COPY ./docker/seed_addr_gen.py /usr/local/bin/
COPY ./docker/watchdog.py /usr/local/bin/
COPY ./docker/run.sh /usr/local/bin/
# Install dependencies
+173 -1
View File
@@ -13,6 +13,16 @@ export BITMESSAGE_MAXTOTALCONNECTIONS="${BITMESSAGE_MAXTOTALCONNECTIONS:-200}"
export BITMESSAGE_TRUSTED_PEER="${BITMESSAGE_TRUSTED_PEER:-}"
export BITMESSAGE_SEND_OUTGOING="${BITMESSAGE_SEND_OUTGOING:-True}"
export BITMESSAGE_KNOWN_NODES="${BITMESSAGE_KNOWN_NODES:-}"
# The watchdog at the end of this file. A daemon that has lost every peer does
# not find its way back on its own, while one that has just started dials hard
# and does -- so a restart is the cure, and noticing is the whole difference.
# On by default; False leaves the supervisor holding the daemon and stops it
# acting. PERIOD is seconds between checks, AFTER how many peerless checks in a
# row it takes to act, COOLDOWN the floor between two restarts.
export BITMESSAGE_WATCHDOG="${BITMESSAGE_WATCHDOG:-True}"
export BITMESSAGE_WATCHDOG_PERIOD="${BITMESSAGE_WATCHDOG_PERIOD:-60}"
export BITMESSAGE_WATCHDOG_AFTER="${BITMESSAGE_WATCHDOG_AFTER:-5}"
export BITMESSAGE_WATCHDOG_COOLDOWN="${BITMESSAGE_WATCHDOG_COOLDOWN:-900}"
# Reject anything but a plain number: this value is written into keys.dat, and
# unlike the credentials below it has no business containing characters that
@@ -38,6 +48,40 @@ case "$BITMESSAGE_SEND_OUTGOING" in
;;
esac
# The same two rules again, for the watchdog. Checked here rather than in the
# loop because a typo would otherwise surface hours later as a supervisor that
# spins, or one that never acts -- and both look like a working container.
case "$BITMESSAGE_WATCHDOG" in
[Tt]rue) BITMESSAGE_WATCHDOG=True ;;
[Ff]alse) BITMESSAGE_WATCHDOG=False ;;
*)
echo "BITMESSAGE_WATCHDOG must be True or False" >&2
exit 1
;;
esac
case "$BITMESSAGE_WATCHDOG_PERIOD" in
'' | *[!0-9]* | 0)
echo "BITMESSAGE_WATCHDOG_PERIOD must be a positive integer" >&2
exit 1
;;
esac
case "$BITMESSAGE_WATCHDOG_AFTER" in
'' | *[!0-9]* | 0)
echo "BITMESSAGE_WATCHDOG_AFTER must be a positive integer" >&2
exit 1
;;
esac
# Zero is allowed here and means "no floor": restart on every verdict.
case "$BITMESSAGE_WATCHDOG_COOLDOWN" in
'' | *[!0-9]*)
echo "BITMESSAGE_WATCHDOG_COOLDOWN must be a non-negative integer" >&2
exit 1
;;
esac
# host:port with a numeric port -- the form both consumers need. PyBitmessage
# does check trustedpeer itself, but by sys.exit() from a constructor deep in
# the network thread: the container dies with the reason buried in the daemon
@@ -175,4 +219,132 @@ then
done &
fi
exec gosu bitmessage pybitmessage -d
# --- the daemon, and the supervisor that owns it --------------------------
#
# This file used to end at `exec gosu bitmessage pybitmessage -d`, which made
# the daemon PID 1, and it is a poor PID 1. daemonize() double-forks and parks
# the grandfather in `while True: time.sleep(1)`; the final child then SIGTERMs
# it to say "ready", and PID 1 drops that signal for want of a handler. Three
# things followed. The grandfather slept for ever. `docker stop` reached the
# real daemon only as the SIGKILL ten seconds later -- which is how a startup
# VACUUM gets cut in half and the node is then trapped retrying it. And a daemon
# that died on its own left the container Up around a corpse, because what PID 1
# was doing had nothing to do with whether the daemon was alive.
#
# Away from PID 1 that grandfather does die on the ready signal. Measured in
# this image: the call returns immediately with status 143 and leaves exactly
# one pybitmessage process behind. So starting the daemon is an ordinary
# blocking call, and everything below is ordinary shell.
#
# What the supervisor does NOT do is act on a daemon whose API is not answering
# at all. That is the trapped-VACUUM node, a restart does not cure it, and
# restarting anyway drops the next VACUUM half-done too. watchdog.py reports
# that case as its own exit code so this loop can leave it alone.
#: How long to wait for a daemon to go away before insisting, in seconds. Twice
#: this is the worst case for a stop, which is what `stop_grace_period` has to
#: cover -- see the note in the README: a clean PyBitmessage shutdown does not
#: fit in Docker's default ten seconds, so a compose file that does not raise
#: the grace period gets the SIGKILL this supervisor exists to avoid.
STOP_TIMEOUT=30
daemon_running() {
pgrep -f pybitmessage >/dev/null 2>&1
}
start_daemon() {
set +e
gosu bitmessage pybitmessage -d
rc=$?
set -e
# 143 is the ready signal reaching the grandfather, which is this call's
# ordinary end. Anything but that or a plain 0 never daemonized.
if [ "$rc" -ne 143 ] && [ "$rc" -ne 0 ]
then
echo "watchdog: the daemon did not start (status $rc)" >&2
return 1
fi
echo "watchdog: daemon started"
}
stop_daemon() {
daemon_running || return 0
# Through the daemon's own API, which runs doCleanShutdown: the database is
# closed instead of being cut off mid-write.
gosu bitmessage python /usr/local/bin/watchdog.py shutdown || true
waited=0
while daemon_running && [ "$waited" -lt "$STOP_TIMEOUT" ]
do
sleep 1
waited=$((waited + 1))
done
daemon_running || return 0
# The API would not answer. TERM, and never KILL: the daemon installs a
# handler for TERM (setSignalHandler) and shuts down properly on it.
echo "watchdog: the API did not stop the daemon, sending TERM" >&2
pkill -TERM -f pybitmessage || true
waited=0
while daemon_running && [ "$waited" -lt "$STOP_TIMEOUT" ]
do
sleep 1
waited=$((waited + 1))
done
}
on_signal() {
echo "watchdog: stopping"
stop_daemon
exit 0
}
trap on_signal TERM INT
start_daemon || exit 1
streak=0
last_restart=0
while :
do
sleep "$BITMESSAGE_WATCHDOG_PERIOD"
if ! daemon_running
then
echo "watchdog: the daemon is gone, starting it again" >&2
start_daemon || exit 1
streak=0
last_restart="$(date +%s)"
continue
fi
[ "$BITMESSAGE_WATCHDOG" = True ] || continue
set +e
gosu bitmessage python /usr/local/bin/watchdog.py peers
verdict=$?
set -e
# 1 is "answered, and has no peers" -- the only verdict worth acting on.
# 2 is "did not answer", and the streak starts over rather than carrying an
# interrupted observation forward.
if [ "$verdict" -ne 1 ]
then
streak=0
continue
fi
streak=$((streak + 1))
[ "$streak" -ge "$BITMESSAGE_WATCHDOG_AFTER" ] || continue
now="$(date +%s)"
if [ "$((now - last_restart))" -lt "$BITMESSAGE_WATCHDOG_COOLDOWN" ]
then
continue
fi
echo "watchdog: no peers for $streak checks, restarting the daemon" >&2
stop_daemon
start_daemon || exit 1
streak=0
last_restart="$(date +%s)"
done
+75
View File
@@ -0,0 +1,75 @@
#!/usr/bin/python
"""Two errands the supervisor in run.sh cannot do in shell: ask, and stop.
`peers` reports how many network connections the daemon has, as an exit code the
shell can branch on. `shutdown` asks the daemon to stop through its own API,
which runs doCleanShutdown -- the database is closed rather than cut off.
The streak, the cooldown and the restart itself are deliberately not here. This
process is short-lived, so a counter would need a file to live in, and then two
places would decide one thing. run.sh counts; this only reports.
Exit codes are the interface:
0 the daemon answered and has peers
1 the daemon answered and has none (`peers`)
or the shutdown was asked for (`shutdown`)
2 the daemon did not answer at all
2 is not 1 on purpose. A daemon whose API is refusing connections is the trapped
node of the startup-VACUUM failure, and restarting it neither cures that nor is
harmless: it drops the next VACUUM half-done as well. The supervisor must be
able to tell "no peers" from "no answer", and only act on the first.
"""
import json
import os
import sys
import urllib
import xmlrpclib
# Credentials go into a URL, so they must be percent-encoded: '@' splits the
# userinfo, '#' truncates the rest, '/' and ':' change what is parsed as host
# and port. healthy_check.py carries the same six lines for the same reason --
# these are one-shot scripts with no package around them to share a module.
API_USER = os.getenv('BITMESSAGE_API_USER', 'bitmessage_api_user')
API_PASSWORD = os.getenv('BITMESSAGE_API_PASSWORD', 'bitmessage_api_password')
API_PORT = os.getenv('BITMESSAGE_API_PORT', '8442')
API_LINK = "http://{}:{}@127.0.0.1:{}/".format(
urllib.quote(API_USER, safe=''), urllib.quote(API_PASSWORD, safe=''), API_PORT)
NO_ANSWER = 2
def peers(api):
"""How many peers the daemon holds, printed for the container log."""
count = json.loads(api.clientStatus())['networkConnections']
print "networkConnections:", count
return 0 if count > 0 else 1
def shutdown(api):
"""Ask the daemon to stop itself cleanly (api.py: HandleShutdown)."""
api.shutdown()
print "shutdown requested"
return 1
COMMANDS = {'peers': peers, 'shutdown': shutdown}
def main(argv):
if len(argv) != 2 or argv[1] not in COMMANDS:
print >> sys.stderr, "usage: watchdog.py peers|shutdown"
return NO_ANSWER
try:
return COMMANDS[argv[1]](xmlrpclib.ServerProxy(API_LINK))
except Exception as exc: # the API is down, or answering something else
print "no answer from the daemon:", exc
return NO_ANSWER
if __name__ == '__main__':
sys.exit(main(sys.argv))