A private Bitmessage contour cannot be assembled by letting the nodes find
each other. The sybil check in connectionpool refuses a candidate whose /16 is
already among the outbound connections, and every container of a compose
project shares one /16 -- so each node keeps a single outbound connection to a
randomly chosen peer, and the contour splits into components on some runs and
not on others.
Three variables make the topology explicit instead:
BITMESSAGE_TRUSTED_PEER trustedpeer = host:port
BITMESSAGE_SEND_OUTGOING sendoutgoingconnections = True/False
BITMESSAGE_KNOWN_NODES host:port,... -> knownnodes.dat
With them a star is one line of config per node: the hub takes
SEND_OUTGOING=False and only accepts, the spokes take TRUSTED_PEER=<hub>:8444.
Details worth knowing:
- trustedpeer is absent from the stock keys.dat, so a substitution alone would
be a silent no-op. The key is added the same way maxtotalconnections is,
and it is added even when the value is empty -- that is how a node that was
pinned before can be unpinned. Its anchors stop at "=" rather than "= ",
because an empty value leaves no trailing space to match.
- knownnodes.dat is rewritten on every start, not only when missing. "Only
when missing" would never have fired: the image ships one, built by the
`pybitmessage -t` run in the Dockerfile, and a named volume inherits it.
Seeding it is also what stops the DNS bootstrap -- deserialising any peer
that is neither a DEFAULT_NODE nor "self" raises knownNodesActual, and
startBootstrappers only runs while that flag is down.
- Both peer variables are validated here. PyBitmessage does check trustedpeer,
but with a sys.exit() from a constructor in the network thread, which reads
as a container that died for no stated reason.
The daemon has always listened on 8444 and nothing published it, so every
node built from this image was outbound-only. Publishing it is now a
documented choice rather than an omission -- including the two things that
are not obvious from the config: peers are told the port from `port` in
keys.dat rather than the one you mapped it to (so only 8444:8444 works),
and the node's own address is never configured at all, because every peer
replaces the hardcoded 127.0.0.1 in the version message with the IP it
sees on the socket.
An open port needs a brake, hence BITMESSAGE_MAXTOTALCONNECTIONS. It
defaults to the PyBitmessage stock 200, so nothing changes for existing
users of the image. The key is inserted when the stock config lacks it
instead of trusting the substitution: the PyBitmessage clone is unpinned,
and a silent no-op would ship a node that looks capped and is not.
The API port in the example compose moves to loopback, which is what the
README already prescribed.
Same pairing the bitdeals-ng repositories use: each file points at the
other under its first heading, and both examples are byte-identical across
the two so they cannot drift.
The examples were not safe to copy. They published the API on every
interface next to the default credentials it documents, carried no volume
-- so keys.dat, the node's Bitmessage identity, was lost on every image
update -- and the docker run one was missing a line continuation and did
not run at all. Ports are now on loopback, the volume is there, and the
run was checked verbatim on a host.
Also: stopresendingafterxdays is documented as 30, which is what run.sh
actually defaults to, not 60; the seed phrase is noted as regenerated per
start; and a short Notes section records what cost time to find out -- '%'
is unusable in the API password, clients have to percent-encode
credentials, and the health check waits for a network connection.
The sed pass that configures the daemon punished anyone setting a strong
password: '&' in a replacement means the whole match, so a password
containing one was silently rewritten into something else, and the '|'
delimiter made sed exit with 'unknown option to s'. Every setting rode in
one invocation and the script had no set -e, so that failure applied none
of them -- not apipassword, not apienabled, not apiinterface -- and the
daemon came up on whatever it had before without saying so.
esc() now escapes backslash, '&' and the delimiter via printf. The
expressions are anchored to the start of the line and name their key in the
replacement instead of using \1, so no backreference is involved and
nothing in another section can match. keys.dat also holds privsigningkey
and privencryptionkey for this node's Bitmessage identities; editing in
place leaves every byte outside [bitmessagesettings] untouched, which is
why this stays a line edit rather than a parse-and-rewrite.
The clients needed the other half of this: credentials go into an XML-RPC
URL, where '@' splits the userinfo and '#' truncates the rest, so a strong
password wrote correctly and still failed to connect. Both now
percent-encode. A '%' in the password remains unusable -- it breaks
PyBitmessage's own config reader and every API call returns 500.
Also here, all found while making the above safe:
- set -eu, with the keys.dat chown guarded. A misconfiguration now stops
the container instead of passing unnoticed.
- The seed-address retry loop used bash brace expansion under CMD ["sh"],
where /bin/sh is dash and {1..4} is a literal, so it ran once, not four
times.
- apt-get update shared a layer with nothing, letting a cached update feed
install months-stale package lists.
- HEALTHCHECK gained a start period; the startup VACUUM takes tens of
seconds on a large messages.dat and the container reported unhealthy for
all of it.
- Removed the AppImage systemd unit, AppArmor profile and updater script.
Nothing referenced them -- not the image, the deployment or the ansible
roles -- and the updater fetched a binary with no signature or checksum
check.
The startup VACUUM of a large messages.dat exceeds the stock 60 s window,
and PyBitmessage then kills the daemon (os._exit). Since it dies mid-VACUUM
lastvacuumtime is never updated, so every later start repeats it and the
node stays down with its API not listening.