bitdeals 88b5b896e1
Build docker image and push to registry.bitdeals.org / main-build-job (push) Successful in 7m12s
feat: supervise the daemon, and restart one that has lost every peer
A PyBitmessage daemon that has dropped to zero network connections does not find
its way back. It can sit there for days — testnet1 did, and every BitDeals deal
completed on that stand since 2026-08-30 left its escrow unspent, because the
guarantor never received a CH3 the node could no longer publish. A daemon that
has just started, by contrast, dials hard and reconnects within seconds. The
cure was already known; what was missing was anything to notice and apply it.

Docker will not: it reacts to a process exiting, never to a healthcheck, and
health-driven restarts exist only under Swarm. Doing it from outside means
handing a container the docker socket, which is root on the host and a poor
trade for a relay with a published port. So the container supervises itself.

`run.sh` no longer ends at `exec gosu bitmessage pybitmessage -d`. That made the
daemon PID 1, and it is a bad PID 1: daemonize() double-forks and parks the
grandfather in `while True: time.sleep(1)`, the final child SIGTERMs it to say
"ready", and PID 1 drops that signal for want of a handler. The grandfather slept
for ever; `docker stop` therefore reached the real daemon only as the SIGKILL ten
seconds later, cutting a startup VACUUM in half — the way a node gets trapped
retrying one it can never finish; and a daemon that died on its own left the
container Up around a corpse, because what PID 1 was doing had nothing to do with
whether the daemon lived.

Away from PID 1 that grandfather does die on the ready signal — measured here:
the start call returns at once with status 143 and leaves exactly one
pybitmessage process behind. That makes starting the daemon an ordinary blocking
call, and the supervisor ordinary shell: a trap that stops the daemon through its
own API, a restart when it is gone, and the peer rule.

What the supervisor does not do is act on a daemon whose API is not answering at
all. That is the trapped-VACUUM node; a restart does not cure it and cuts the
next VACUUM short as well. watchdog.py reports that case as its own exit code so
the loop can leave it to a person. The healthcheck is untouched: it reports, and
does not act.

Four settings, on by default: BITMESSAGE_WATCHDOG, and _PERIOD, _AFTER,
_COOLDOWN. All validated at start, where a typo is visible, rather than hours
later as a supervisor that spins or one that never acts.

Callers must raise the stop grace period — 90s in compose, or --stop-timeout 90.
A clean shutdown took 17.5 s on a small database and grows with it, so under
Docker's default ten the daemon is killed mid-write anyway and the supervisor
buys nothing. An image cannot set this for itself.
2026-09-04 09:53:07 +00:00
2026-03-18 11:42:32 +03:00

Intro

Русская версия: README.ru-RU.md

PyBitmessage is a client of the Bitmessages P2P communication protocol used to send encrypted messages to another person or to many subscribers.

PyBitmessage client running as a daemon in docker container with XML-RPC API enabled.

This repository covers the docker deployment only.

Usage

The container generates a Bitmessage Deterministic Addresses based on a BITMESSAGE_SEED_PHRASE variable.

Here are some example snippets to help you get started creating a container.

The container has two ports and they are not interchangeable. 8442 is the XML-RPC API: it controls the daemon completely and has no TLS, so set your own credentials and keep it on loopback. 8444 is the Bitmessage P2P port: publish it to let other nodes connect in, leave it unpublished to stay outbound-only. The daemon listens on both inside the container either way.

docker-compose

services:
  pybitmessage:
    build:
      context: https://git.bitdeals.org/private/bitmessage.git
      dockerfile: ./docker/Dockerfile
    image: registry.bitdeals.org/bitmessage
    environment:
      - BITMESSAGE_API_USER=CHANGE_ME
      - BITMESSAGE_API_PASSWORD=CHANGE_ME
      - BITMESSAGE_SEED_PHRASE=bitmessage_seed_phrase
      - BITMESSAGE_SEED_ADDRESSES=1
      - BITMESSAGE_TTL=172800
      - BITMESSAGE_STOPRESENDINGAFTERXDAYS=60
      - BITMESSAGE_MAXTOTALCONNECTIONS=40
    ports:
      - 127.0.0.1:8442:8442   # API — loopback only
      - 8444:8444             # P2P — omit this line to stay outbound-only
    volumes:
      - bitmessage:/home/bitmessage

volumes:
  bitmessage:

docker cli

docker run -d \
  -e BITMESSAGE_API_USER=CHANGE_ME \
  -e BITMESSAGE_API_PASSWORD=CHANGE_ME \
  -e BITMESSAGE_SEED_PHRASE=bitmessage_seed_phrase \
  -e BITMESSAGE_SEED_ADDRESSES=1 \
  -e BITMESSAGE_TTL=172800 \
  -e BITMESSAGE_STOPRESENDINGAFTERXDAYS=60 \
  -e BITMESSAGE_MAXTOTALCONNECTIONS=40 \
  -p 127.0.0.1:8442:8442 \
  -p 8444:8444 \
  -v bitmessage:/home/bitmessage \
  registry.bitdeals.org/bitmessage

Parameters

Container images are configured using parameters passed at runtime.

Parameter Function
-p 127.0.0.1:8442 API port. The daemon always binds 0.0.0.0 inside the container, so what you publish decides who reaches it
-p 8444 Bitmessage P2P port. Optional: without it the node still connects out to peers, it just cannot be connected to. Must be published as 8444:8444 — see Notes
-v /home/bitmessage Data directory: keys.dat (identity, settings) and messages.dat. Without it the node is a new node after every update
-e BITMESSAGE_API_USER XML-RPC API user. Default: bitmessage_api_user — change it
-e BITMESSAGE_API_PASSWORD XML-RPC API password. Default: bitmessage_api_password — change it, see Notes
-e BITMESSAGE_SEED_PHRASE Create Deterministic Addresses password. Default: regenerated on every start, giving different addresses each time. Only used when BITMESSAGE_SEED_ADDRESSES is above 0
-e BITMESSAGE_SEED_ADDRESSES Number of Deterministic Addresses to generate. Default: 0
-e BITMESSAGE_TTL The expiration of newly send messages, in seconds. Default: 172800
-e BITMESSAGE_STOPRESENDINGAFTERXDAYS Stop resending unreceived message after X days. Default: 30
-e BITMESSAGE_APIVARIANT provides xml or json-RPC API. Default: legacy
-e BITMESSAGE_MAXTOTALCONNECTIONS Cap on all connections at once, inbound and outbound together (maxoutboundconnections is 8, so this minus 8 is the inbound headroom). Default: 200, the PyBitmessage stock value — lower it when the P2P port is published
-e BITMESSAGE_TRUSTED_PEER host:port of the one peer this node may connect out to; it dials nothing else. Default: empty — the node chooses its own peers. See Notes
-e BITMESSAGE_SEND_OUTGOING Whether the node dials out at all, True or False. Default: True. False gives a node that only accepts inbound connections — the hub of a private contour
-e BITMESSAGE_KNOWN_NODES Comma-separated host:port list, written into knownnodes.dat on every start in place of whatever was there. Default: empty — the file is left as it is. Also switches the DNS bootstrap off, see Notes
-e BITMESSAGE_WATCHDOG Whether to restart a daemon that has lost every peer, True or False. Default: True. See Notes
-e BITMESSAGE_WATCHDOG_PERIOD Seconds between peer checks. Default: 60
-e BITMESSAGE_WATCHDOG_AFTER How many peerless checks in a row it takes to act. Default: 5, so five minutes at the default period
-e BITMESSAGE_WATCHDOG_COOLDOWN Floor between two restarts, in seconds. Default: 900. 0 removes it

Notes

  • % cannot be used in the API password: it breaks PyBitmessage's own config reader, and every API call then returns 500 while keys.dat looks correct. Other characters are fine, the container escapes them.

  • Clients must percent-encode the credentials — they go into http://user:password@host:port/, where @, #, / and : change how the URL parses. The scripts in this image do; yours has to as well.

  • The container turns healthy once the daemon has a network connection, which on a new node takes a few minutes. The start period also covers the startup VACUUM of messages.dat.

  • A peerless daemon is restarted; a silent one is not. A daemon that has lost every peer does not find its way back — it can sit at zero connections for as long as you leave it — while one that has just started dials hard and does. So the container supervises its own daemon: it asks the API for the connection count every BITMESSAGE_WATCHDOG_PERIOD, and after BITMESSAGE_WATCHDOG_AFTER answers of zero in a row it stops the daemon through its own API and starts it again. The container is not restarted and nothing outside it is involved, so a node that depends on this one keeps its addresses and only sees the API blink.

    A daemon whose API does not answer at all is left alone, deliberately. That is the node trapped retrying a startup VACUUM it cannot finish, restarting does not cure it, and restarting anyway cuts the next VACUUM in half too. The healthcheck reports both cases as unhealthy; only one of them is something to do anything about, and a person has to look at the other.

    Give the container stop_grace_period: 90s (compose) or --stop-timeout 90 (docker run). Closing messages.dat properly takes longer than Docker's default ten seconds — measured at about thirteen on a small database, and it grows with the file — so without this the daemon is SIGKILLed mid-write, which is the very thing the supervisor is there to prevent. An image cannot set this for itself; only the caller can.

    BITMESSAGE_WATCHDOG=False turns the peer rule off. The supervisor stays either way — it is also what stops the daemon cleanly on docker stop, and what restarts one that died outright rather than leaving the container up around a corpse.

  • The P2P port must be published as 8444:8444. The daemon tells peers the port from its own config (port in keys.dat), not the port you mapped it to, so 8555:8444 advertises a port nobody can reach. A different host port needs extport in keys.dat, which this container does not template.

  • Your own address is never configured. The version message carries a hardcoded 127.0.0.1 that every peer discards in favour of the IP it sees on the socket, and it takes the port from that same message. So a node with 8444 published is found by the network on its own, as soon as it connects out — there is no host IP or DNS name to set anywhere. The one exception is a Tor hidden service, which needs an explicit onionhostname.

  • A private contour needs its peers pinned, and a star to pin them into. PyBitmessage refuses a candidate whose network group (the /16 for IPv4) is already represented among its outbound connections. Every container of one compose project lives in a single /16, so each node keeps exactly one outbound connection, to a peer it picked at random — which as often as not leaves the contour split into components. Give one node BITMESSAGE_SEND_OUTGOING=False so it becomes a hub that only accepts (the check looks at outbound connections only, so inbound are not capped by it), point the rest at it with BITMESSAGE_TRUSTED_PEER=<hub-ip>:8444, and objects travel spoke → hub → spokes. Use IP addresses, not service names: the sybil check parses the host as an IP, and the addr exchange between nodes carries IPs anyway.

  • BITMESSAGE_KNOWN_NODES is what keeps a private contour private. A node whose knownnodes.dat names a peer outside PyBitmessage's built-in default list stops asking bootstrap8080.bitmessage.org for more; without it even a pinned node resolves the public bootstrap host on every start. The file is rewritten on each start, so the variable, not the container's history, is what the node believes on boot.

  • Publishing 8444 is a deliberate security trade-off. The daemon runs on Python 2 and already parses untrusted data from its outbound peers, so an open port does not create that exposure — it changes who may connect, when, and how many. The pre-handshake parser becomes reachable by anyone, and the practical risk is resource exhaustion rather than code execution. Keep BITMESSAGE_MAXTOTALCONNECTIONS low and put memory and CPU limits on the container.

S
Description
PyBitmessage dockerfile
Readme
219 KiB
Languages
Shell 58%
Dockerfile 22.8%
Python 19.2%