From 88b5b896e1c7c9687314d9e5bdd0b9e376927860 Mon Sep 17 00:00:00 2001 From: bitdeals Date: Fri, 4 Sep 2026 09:53:07 +0000 Subject: [PATCH] feat: supervise the daemon, and restart one that has lost every peer MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A PyBitmessage daemon that has dropped to zero network connections does not find its way back. It can sit there for days — testnet1 did, and every BitDeals deal completed on that stand since 2026-08-30 left its escrow unspent, because the guarantor never received a CH3 the node could no longer publish. A daemon that has just started, by contrast, dials hard and reconnects within seconds. The cure was already known; what was missing was anything to notice and apply it. Docker will not: it reacts to a process exiting, never to a healthcheck, and health-driven restarts exist only under Swarm. Doing it from outside means handing a container the docker socket, which is root on the host and a poor trade for a relay with a published port. So the container supervises itself. `run.sh` no longer ends at `exec gosu bitmessage pybitmessage -d`. That made the daemon PID 1, and it is a bad PID 1: daemonize() double-forks and parks the grandfather in `while True: time.sleep(1)`, the final child SIGTERMs it to say "ready", and PID 1 drops that signal for want of a handler. The grandfather slept for ever; `docker stop` therefore reached the real daemon only as the SIGKILL ten seconds later, cutting a startup VACUUM in half — the way a node gets trapped retrying one it can never finish; and a daemon that died on its own left the container Up around a corpse, because what PID 1 was doing had nothing to do with whether the daemon lived. Away from PID 1 that grandfather does die on the ready signal — measured here: the start call returns at once with status 143 and leaves exactly one pybitmessage process behind. That makes starting the daemon an ordinary blocking call, and the supervisor ordinary shell: a trap that stops the daemon through its own API, a restart when it is gone, and the peer rule. What the supervisor does not do is act on a daemon whose API is not answering at all. That is the trapped-VACUUM node; a restart does not cure it and cuts the next VACUUM short as well. watchdog.py reports that case as its own exit code so the loop can leave it to a person. The healthcheck is untouched: it reports, and does not act. Four settings, on by default: BITMESSAGE_WATCHDOG, and _PERIOD, _AFTER, _COOLDOWN. All validated at start, where a typo is visible, rather than hours later as a supervisor that spins or one that never acts. Callers must raise the stop grace period — 90s in compose, or --stop-timeout 90. A clean shutdown took 17.5 s on a small database and grows with it, so under Docker's default ten the daemon is killed mid-write anyway and the supervisor buys nothing. An image cannot set this for itself. --- README.md | 31 ++++++++ README.ru-RU.md | 29 ++++++++ docker/Dockerfile | 1 + docker/run.sh | 174 ++++++++++++++++++++++++++++++++++++++++++++- docker/watchdog.py | 75 +++++++++++++++++++ 5 files changed, 309 insertions(+), 1 deletion(-) create mode 100644 docker/watchdog.py diff --git a/README.md b/README.md index a3de94a..36e694c 100644 --- a/README.md +++ b/README.md @@ -84,6 +84,10 @@ Container images are configured using parameters passed at runtime. |-e BITMESSAGE_TRUSTED_PEER|`host:port` of the one peer this node may connect out to; it dials nothing else. Default: empty — the node chooses its own peers. See Notes| |-e BITMESSAGE_SEND_OUTGOING|Whether the node dials out at all, `True` or `False`. Default: `True`. `False` gives a node that only accepts inbound connections — the hub of a private contour| |-e BITMESSAGE_KNOWN_NODES|Comma-separated `host:port` list, written into `knownnodes.dat` on every start in place of whatever was there. Default: empty — the file is left as it is. Also switches the DNS bootstrap off, see Notes| +|-e BITMESSAGE_WATCHDOG|Whether to restart a daemon that has lost every peer, `True` or `False`. Default: `True`. See Notes| +|-e BITMESSAGE_WATCHDOG_PERIOD|Seconds between peer checks. Default: `60`| +|-e BITMESSAGE_WATCHDOG_AFTER|How many peerless checks in a row it takes to act. Default: `5`, so five minutes at the default period| +|-e BITMESSAGE_WATCHDOG_COOLDOWN|Floor between two restarts, in seconds. Default: `900`. `0` removes it| # Notes @@ -96,6 +100,33 @@ Container images are configured using parameters passed at runtime. - The container turns healthy once the daemon has a network connection, which on a new node takes a few minutes. The start period also covers the startup `VACUUM` of `messages.dat`. +- **A peerless daemon is restarted; a silent one is not.** A daemon that has + lost every peer does not find its way back — it can sit at zero connections + for as long as you leave it — while one that has just started dials hard and + does. So the container supervises its own daemon: it asks the API for the + connection count every `BITMESSAGE_WATCHDOG_PERIOD`, and after + `BITMESSAGE_WATCHDOG_AFTER` answers of zero in a row it stops the daemon + through its own API and starts it again. The *container* is not restarted and + nothing outside it is involved, so a node that depends on this one keeps its + addresses and only sees the API blink. + + A daemon whose API does not answer at all is left alone, deliberately. That is + the node trapped retrying a startup `VACUUM` it cannot finish, restarting does + not cure it, and restarting anyway cuts the next `VACUUM` in half too. The + healthcheck reports both cases as unhealthy; only one of them is something to + do anything about, and a person has to look at the other. + + **Give the container `stop_grace_period: 90s`** (compose) or + `--stop-timeout 90` (`docker run`). Closing `messages.dat` properly takes + longer than Docker's default ten seconds — measured at about thirteen on a + small database, and it grows with the file — so without this the daemon is + SIGKILLed mid-write, which is the very thing the supervisor is there to + prevent. An image cannot set this for itself; only the caller can. + + `BITMESSAGE_WATCHDOG=False` turns the peer rule off. The supervisor stays + either way — it is also what stops the daemon cleanly on `docker stop`, and + what restarts one that died outright rather than leaving the container up + around a corpse. - **The P2P port must be published as `8444:8444`.** The daemon tells peers the port from its own config (`port` in `keys.dat`), not the port you mapped it to, so `8555:8444` advertises a port nobody can reach. A different host port diff --git a/README.ru-RU.md b/README.ru-RU.md index b88ddf6..e4451cc 100644 --- a/README.ru-RU.md +++ b/README.ru-RU.md @@ -84,6 +84,10 @@ docker run -d \ |-e BITMESSAGE_TRUSTED_PEER|`host:port` единственного пира, к которому узел подключается наружу; больше ни к кому. По умолчанию: пусто — узел выбирает пиров сам. См. «Замечания»| |-e BITMESSAGE_SEND_OUTGOING|Подключается ли узел наружу вообще: `True` или `False`. По умолчанию: `True`. При `False` получается узел, который только принимает входящие, — центр приватного контура| |-e BITMESSAGE_KNOWN_NODES|Список `host:port` через запятую; записывается в `knownnodes.dat` при каждом старте вместо того, что там было. По умолчанию: пусто — файл не трогается. Заодно выключает бутстрап по DNS, см. «Замечания»| +|-e BITMESSAGE_WATCHDOG|Перезапускать ли демона, потерявшего всех пиров: `True` или `False`. По умолчанию: `True`. См. «Замечания»| +|-e BITMESSAGE_WATCHDOG_PERIOD|Секунд между проверками числа пиров. По умолчанию: `60`| +|-e BITMESSAGE_WATCHDOG_AFTER|Сколько проверок подряд без пиров нужно, чтобы вмешаться. По умолчанию: `5` — то есть пять минут при периоде по умолчанию| +|-e BITMESSAGE_WATCHDOG_COOLDOWN|Минимальный промежуток между двумя перезапусками, в секундах. По умолчанию: `900`. `0` снимает ограничение| # Замечания @@ -96,6 +100,31 @@ docker run -d \ - Контейнер становится здоровым, когда у демона появилось сетевое соединение, — на новом узле это несколько минут. Начальный период ожидания заодно покрывает стартовый `VACUUM` файла `messages.dat`. +- **Демона без пиров перезапускаем, молчащего — нет.** Демон, потерявший всех + пиров, сам обратно дорогу не находит: он может простоять на нуле соединений + сколько угодно, — а только что запущенный ищет пиров напористо и находит. + Поэтому контейнер присматривает за собственным демоном: раз в + `BITMESSAGE_WATCHDOG_PERIOD` спрашивает у API число соединений и после + `BITMESSAGE_WATCHDOG_AFTER` нулей подряд останавливает демона через его же API + и поднимает заново. *Контейнер* при этом не пересоздаётся и никто снаружи не + участвует, так что соседний узел сохраняет адреса и видит лишь моргнувший API. + + Демона, у которого API не отвечает вовсе, сторож не трогает намеренно. Это + узел, застрявший на стартовом `VACUUM`, которого он не может закончить; + перезапуск его не лечит, а лишь обрывает следующий `VACUUM` на середине. + Проверка здоровья показывает `unhealthy` в обоих случаях, но вмешиваться стоит + только в первый, а во втором нужен человек. + + **Задайте контейнеру `stop_grace_period: 90s`** (compose) или + `--stop-timeout 90` (`docker run`). Корректное закрытие `messages.dat` не + укладывается в десять секунд, отпущенные Docker по умолчанию, — замерено около + тринадцати на небольшой базе, и растёт вместе с файлом, — так что без этого + демона убьёт SIGKILL на середине записи, ровно то, ради предотвращения чего + супервизор и заведён. Сам образ это задать не может, только вызывающая сторона. + + `BITMESSAGE_WATCHDOG=False` выключает правило о пирах. Супервизор остаётся в + любом случае: он же корректно останавливает демона по `docker stop` и + поднимает того, кто умер сам, вместо контейнера, стоящего вокруг трупа. - **P2P-порт публикуется только как `8444:8444`.** Пирам демон сообщает порт из собственной конфигурации (`port` в `keys.dat`), а не тот, в который вы его отобразили, поэтому `8555:8444` объявляет сети порт, где никого нет. Для diff --git a/docker/Dockerfile b/docker/Dockerfile index 87f973d..c4249c7 100644 --- a/docker/Dockerfile +++ b/docker/Dockerfile @@ -59,6 +59,7 @@ ENV BITMESSAGE_HOME=${HOME} COPY --from=0 /usr/local/ /usr/local/ COPY ./docker/healthy_check.py /usr/local/bin/ COPY ./docker/seed_addr_gen.py /usr/local/bin/ +COPY ./docker/watchdog.py /usr/local/bin/ COPY ./docker/run.sh /usr/local/bin/ # Install dependencies diff --git a/docker/run.sh b/docker/run.sh index 3876e02..0677048 100644 --- a/docker/run.sh +++ b/docker/run.sh @@ -13,6 +13,16 @@ export BITMESSAGE_MAXTOTALCONNECTIONS="${BITMESSAGE_MAXTOTALCONNECTIONS:-200}" export BITMESSAGE_TRUSTED_PEER="${BITMESSAGE_TRUSTED_PEER:-}" export BITMESSAGE_SEND_OUTGOING="${BITMESSAGE_SEND_OUTGOING:-True}" export BITMESSAGE_KNOWN_NODES="${BITMESSAGE_KNOWN_NODES:-}" +# The watchdog at the end of this file. A daemon that has lost every peer does +# not find its way back on its own, while one that has just started dials hard +# and does -- so a restart is the cure, and noticing is the whole difference. +# On by default; False leaves the supervisor holding the daemon and stops it +# acting. PERIOD is seconds between checks, AFTER how many peerless checks in a +# row it takes to act, COOLDOWN the floor between two restarts. +export BITMESSAGE_WATCHDOG="${BITMESSAGE_WATCHDOG:-True}" +export BITMESSAGE_WATCHDOG_PERIOD="${BITMESSAGE_WATCHDOG_PERIOD:-60}" +export BITMESSAGE_WATCHDOG_AFTER="${BITMESSAGE_WATCHDOG_AFTER:-5}" +export BITMESSAGE_WATCHDOG_COOLDOWN="${BITMESSAGE_WATCHDOG_COOLDOWN:-900}" # Reject anything but a plain number: this value is written into keys.dat, and # unlike the credentials below it has no business containing characters that @@ -38,6 +48,40 @@ case "$BITMESSAGE_SEND_OUTGOING" in ;; esac +# The same two rules again, for the watchdog. Checked here rather than in the +# loop because a typo would otherwise surface hours later as a supervisor that +# spins, or one that never acts -- and both look like a working container. +case "$BITMESSAGE_WATCHDOG" in + [Tt]rue) BITMESSAGE_WATCHDOG=True ;; + [Ff]alse) BITMESSAGE_WATCHDOG=False ;; + *) + echo "BITMESSAGE_WATCHDOG must be True or False" >&2 + exit 1 + ;; +esac + +case "$BITMESSAGE_WATCHDOG_PERIOD" in + '' | *[!0-9]* | 0) + echo "BITMESSAGE_WATCHDOG_PERIOD must be a positive integer" >&2 + exit 1 + ;; +esac + +case "$BITMESSAGE_WATCHDOG_AFTER" in + '' | *[!0-9]* | 0) + echo "BITMESSAGE_WATCHDOG_AFTER must be a positive integer" >&2 + exit 1 + ;; +esac + +# Zero is allowed here and means "no floor": restart on every verdict. +case "$BITMESSAGE_WATCHDOG_COOLDOWN" in + '' | *[!0-9]*) + echo "BITMESSAGE_WATCHDOG_COOLDOWN must be a non-negative integer" >&2 + exit 1 + ;; +esac + # host:port with a numeric port -- the form both consumers need. PyBitmessage # does check trustedpeer itself, but by sys.exit() from a constructor deep in # the network thread: the container dies with the reason buried in the daemon @@ -175,4 +219,132 @@ then done & fi -exec gosu bitmessage pybitmessage -d +# --- the daemon, and the supervisor that owns it -------------------------- +# +# This file used to end at `exec gosu bitmessage pybitmessage -d`, which made +# the daemon PID 1, and it is a poor PID 1. daemonize() double-forks and parks +# the grandfather in `while True: time.sleep(1)`; the final child then SIGTERMs +# it to say "ready", and PID 1 drops that signal for want of a handler. Three +# things followed. The grandfather slept for ever. `docker stop` reached the +# real daemon only as the SIGKILL ten seconds later -- which is how a startup +# VACUUM gets cut in half and the node is then trapped retrying it. And a daemon +# that died on its own left the container Up around a corpse, because what PID 1 +# was doing had nothing to do with whether the daemon was alive. +# +# Away from PID 1 that grandfather does die on the ready signal. Measured in +# this image: the call returns immediately with status 143 and leaves exactly +# one pybitmessage process behind. So starting the daemon is an ordinary +# blocking call, and everything below is ordinary shell. +# +# What the supervisor does NOT do is act on a daemon whose API is not answering +# at all. That is the trapped-VACUUM node, a restart does not cure it, and +# restarting anyway drops the next VACUUM half-done too. watchdog.py reports +# that case as its own exit code so this loop can leave it alone. + +#: How long to wait for a daemon to go away before insisting, in seconds. Twice +#: this is the worst case for a stop, which is what `stop_grace_period` has to +#: cover -- see the note in the README: a clean PyBitmessage shutdown does not +#: fit in Docker's default ten seconds, so a compose file that does not raise +#: the grace period gets the SIGKILL this supervisor exists to avoid. +STOP_TIMEOUT=30 + +daemon_running() { + pgrep -f pybitmessage >/dev/null 2>&1 +} + +start_daemon() { + set +e + gosu bitmessage pybitmessage -d + rc=$? + set -e + # 143 is the ready signal reaching the grandfather, which is this call's + # ordinary end. Anything but that or a plain 0 never daemonized. + if [ "$rc" -ne 143 ] && [ "$rc" -ne 0 ] + then + echo "watchdog: the daemon did not start (status $rc)" >&2 + return 1 + fi + echo "watchdog: daemon started" +} + +stop_daemon() { + daemon_running || return 0 + # Through the daemon's own API, which runs doCleanShutdown: the database is + # closed instead of being cut off mid-write. + gosu bitmessage python /usr/local/bin/watchdog.py shutdown || true + waited=0 + while daemon_running && [ "$waited" -lt "$STOP_TIMEOUT" ] + do + sleep 1 + waited=$((waited + 1)) + done + daemon_running || return 0 + # The API would not answer. TERM, and never KILL: the daemon installs a + # handler for TERM (setSignalHandler) and shuts down properly on it. + echo "watchdog: the API did not stop the daemon, sending TERM" >&2 + pkill -TERM -f pybitmessage || true + waited=0 + while daemon_running && [ "$waited" -lt "$STOP_TIMEOUT" ] + do + sleep 1 + waited=$((waited + 1)) + done +} + +on_signal() { + echo "watchdog: stopping" + stop_daemon + exit 0 +} + +trap on_signal TERM INT + +start_daemon || exit 1 + +streak=0 +last_restart=0 + +while : +do + sleep "$BITMESSAGE_WATCHDOG_PERIOD" + + if ! daemon_running + then + echo "watchdog: the daemon is gone, starting it again" >&2 + start_daemon || exit 1 + streak=0 + last_restart="$(date +%s)" + continue + fi + + [ "$BITMESSAGE_WATCHDOG" = True ] || continue + + set +e + gosu bitmessage python /usr/local/bin/watchdog.py peers + verdict=$? + set -e + + # 1 is "answered, and has no peers" -- the only verdict worth acting on. + # 2 is "did not answer", and the streak starts over rather than carrying an + # interrupted observation forward. + if [ "$verdict" -ne 1 ] + then + streak=0 + continue + fi + + streak=$((streak + 1)) + [ "$streak" -ge "$BITMESSAGE_WATCHDOG_AFTER" ] || continue + + now="$(date +%s)" + if [ "$((now - last_restart))" -lt "$BITMESSAGE_WATCHDOG_COOLDOWN" ] + then + continue + fi + + echo "watchdog: no peers for $streak checks, restarting the daemon" >&2 + stop_daemon + start_daemon || exit 1 + streak=0 + last_restart="$(date +%s)" +done diff --git a/docker/watchdog.py b/docker/watchdog.py new file mode 100644 index 0000000..0f7bdda --- /dev/null +++ b/docker/watchdog.py @@ -0,0 +1,75 @@ +#!/usr/bin/python + +"""Two errands the supervisor in run.sh cannot do in shell: ask, and stop. + +`peers` reports how many network connections the daemon has, as an exit code the +shell can branch on. `shutdown` asks the daemon to stop through its own API, +which runs doCleanShutdown -- the database is closed rather than cut off. + +The streak, the cooldown and the restart itself are deliberately not here. This +process is short-lived, so a counter would need a file to live in, and then two +places would decide one thing. run.sh counts; this only reports. + +Exit codes are the interface: + + 0 the daemon answered and has peers + 1 the daemon answered and has none (`peers`) + or the shutdown was asked for (`shutdown`) + 2 the daemon did not answer at all + +2 is not 1 on purpose. A daemon whose API is refusing connections is the trapped +node of the startup-VACUUM failure, and restarting it neither cures that nor is +harmless: it drops the next VACUUM half-done as well. The supervisor must be +able to tell "no peers" from "no answer", and only act on the first. +""" + +import json +import os +import sys +import urllib +import xmlrpclib + +# Credentials go into a URL, so they must be percent-encoded: '@' splits the +# userinfo, '#' truncates the rest, '/' and ':' change what is parsed as host +# and port. healthy_check.py carries the same six lines for the same reason -- +# these are one-shot scripts with no package around them to share a module. +API_USER = os.getenv('BITMESSAGE_API_USER', 'bitmessage_api_user') +API_PASSWORD = os.getenv('BITMESSAGE_API_PASSWORD', 'bitmessage_api_password') +API_PORT = os.getenv('BITMESSAGE_API_PORT', '8442') + +API_LINK = "http://{}:{}@127.0.0.1:{}/".format( + urllib.quote(API_USER, safe=''), urllib.quote(API_PASSWORD, safe=''), API_PORT) + +NO_ANSWER = 2 + + +def peers(api): + """How many peers the daemon holds, printed for the container log.""" + count = json.loads(api.clientStatus())['networkConnections'] + print "networkConnections:", count + return 0 if count > 0 else 1 + + +def shutdown(api): + """Ask the daemon to stop itself cleanly (api.py: HandleShutdown).""" + api.shutdown() + print "shutdown requested" + return 1 + + +COMMANDS = {'peers': peers, 'shutdown': shutdown} + + +def main(argv): + if len(argv) != 2 or argv[1] not in COMMANDS: + print >> sys.stderr, "usage: watchdog.py peers|shutdown" + return NO_ANSWER + try: + return COMMANDS[argv[1]](xmlrpclib.ServerProxy(API_LINK)) + except Exception as exc: # the API is down, or answering something else + print "no answer from the daemon:", exc + return NO_ANSWER + + +if __name__ == '__main__': + sys.exit(main(sys.argv))