mirror of https://github.com/OpenIdentityPlatform/OpenDJ.git

Valery Kharseko
22 hours ago 947c0c9a191cc5b771411e0d0f03d7a56aeedf1e
refs
author Valery Kharseko <vharseko@3a-systems.ru>
Tuesday, October 6, 2026 08:58 +0200
committer GitHub <noreply@github.com>
Tuesday, October 6, 2026 08:58 +0200
commit947c0c9a191cc5b771411e0d0f03d7a56aeedf1e
tree d602f32c087f2cc0059df2e04ce02eeb82371ec8 tree | zip | gz
parent 85b28b3d0fc05bc49d350a98c07459f4bcb21f43 view | diff
[#1086] Join replication in the background on every start of the Docker image (#1115)

Fixes #1086.

The container joined its replication topology once, during the first
bootstrap only, through a single `MASTER_SERVER` it recognised by an
unanchored `grep` of `/etc/hosts`, and a join that failed turned into a
**healthy, unreplicated** server on the next restart, because a restart
wrote the health marker right after `upgrade -n`. A StatefulSet cannot
rely on any of that (#1086, discussion #1079).

The join becomes a background step next to the server, which stays PID 1
of the container (#1085):

- **`bootstrap/join.sh` (new)** runs in the background on every start
and serves `OPENDJ_REPLICATION_TYPE=simple`.
- **Peers come from a list**: `REPLICATION_PEERS=host1,host2,…` (DNS
names), which a chart derives from the StatefulSet ordinals;
`MASTER_SERVER` keeps working as a one-element list. The server
recognises itself by a name equal to `hostname -f` or to `hostname -f`
cut at a dot (a pod listed as `<sts>-N.<svc>` has the FQDN
`<sts>-N.<svc>.<ns>.svc.…`), and by its own addresses and the names
`/etc/hosts` gives them — compared whole. The `grep` of `/etc/hosts` it
replaces took `opendj-1` for `opendj-10`, a replica for its master when
an `--add-host` named the master, and a master that lost its volume for
a replica of itself.
- **Membership is decided from what is there**, never from an exit code:
exit 5 of `dsreplication enable` covers "already replicated" and
"`BASE_DN` not found on one of the servers" alike, so the step instead
checks the replication domain for `BASE_DN` in `cn=config` and the
registration in `cn=admin data`, and until both are there it tries
again, a bounded, configurable number of times, each `enable` under its
own `timeout` (`REPLICATION_ATTEMPT_TIMEOUT`), since it can hang on a
peer that stops mid-operation. The retry and timeout knobs are read as
decimal whole numbers (a leading zero does not turn them octal); a value
that is not one falls back to its default. `dsreplication initialize` —
a full import — has a bound of its own,
`REPLICATION_INITIALIZE_TIMEOUT`, none by default.
- **Joins that involve the same server take turns**: two `dsreplication
enable` runs through the same server at once break each other — both
create the replication server a seed does not have yet, the loser fails
half-way (exit 17) and every later enable exits 5 without ever
completing its membership. An enable therefore holds a lock on the peer
it runs through and on its own server — the entry `cn=Docker Join
Lock,cn=config`, which only one add creates — and a server that takes
its own replication down holds its own. Neither lock is waited for, and
rounds wait a random part of `REPLICATION_RETRY_INTERVAL` on top of it,
so two joins that each hold the other's turn do not keep meeting. A lock
left by a killed join is broken once it is older than
`REPLICATION_ATTEMPT_TIMEOUT` plus a minute, or at once by a later start
of the server that left it, by a delete that asserts the value it read.
- **Whose data the topology carries follows what each volume went
through**: every fresh volume holds entries
(`ADD_BASE_ENTRY`/`SAMPLE_DATA`), and two freshly bootstrapped volumes
even share a generation ID, so "does `BASE_DN` have entries" decides
nothing. `run.sh` marks a volume before it bootstraps it, and the join
publishes that to the peers in a local, non-replicated entry, `cn=Docker
Join,cn=config`: `pending` until the volume received the data of the
topology, `ready` afterwards — published before the marker goes, so the
two never tell different stories — and `rejoining` while a volume that
holds that data has its replication taken down to enable it anew. Every
server joins only through a ready peer, and a pending one also
initializes only from one — however many restarts that takes — and
cross-checks the generation IDs. So servers may also start together
(Compose, `podManagementPolicy: Parallel`) without one taking another's
bootstrap data for the topology's, a bootstrap that failed or was killed
half-way is initialized from the topology rather than taken for its
data, and two servers whose replication is down — rejoining, pending or
still bootstrapping — never register only each other and serve a
topology of their own beside the survivors. Where only `MASTER_SERVER`
is set, a peer that publishes no state (an older image) counts as ready.
- **The seed is a rule**: only the first entry of `REPLICATION_PEERS`
may declare that there is no topology and seed it with its own data — at
once when every other peer answers and is pending, otherwise only once
its retries are exhausted while no other peer answers that may hold the
data: a ready or rejoining one, one that publishes no state but
replicates `BASE_DN` (a server of an earlier image), or one that refuses
the bind with `ROOT_PASSWORD` (only a server past its bootstrap has
another root password). Anyone else stays unhealthy, which surfaces a
lost topology instead of forking it. A `-0` that lost its volume
therefore rejoins through the survivors and takes the data of the
topology back, instead of bootstrapping an empty, unreplicated server
behind the same Service. The residual risk — every other server down
*and* the first peer's volume lost — is documented in the README.
- **Leaving is handled by the survivors**: with `REPLICATION_PEERS` set
explicitly, a joined server removes every server registered in `cn=admin
data` but no longer listed, and prunes it from every
`replication-server` list it holds — that of its replication server and
those of the domains of `BASE_DN`, `cn=schema` and `cn=admin data`. A
server registered by an address (a master that `MASTER_SERVER=<address>`
named) is left alone, as nothing tells it from a listed name. It also
adds every listed, registered peer its lists lack, so joins that ran at
the same time cannot leave two replication servers unaware of each
other. `dsreplication disable` in a `preStop` hook could not do this: it
fires on every termination — rolling update, drain — and changes only
the servers it can reach, while OpenDJ 4 has no cleanup subcommand for a
dead one. With only `MASTER_SERVER` set nothing is removed or added, so
servers joined by hand stay.
- **Coming back is handled by the returning server**: scaled up again on
the volume it kept, a server the survivors removed still replicates, and
`dsreplication enable` between two servers whose `cn=admin data` is
replicated registers nobody. When at least one answering peer that
replicates `BASE_DN` lists the registrations and none of them registers
it — a search that fails decides nothing — it takes its own replication
configuration down (`dsreplication disable --disableAll`) and enables
again, which registers it. From the first change until that enable
registered it, writes it took would replicate nowhere. So before
anything changes, the backend of `BASE_DN` refuses the writes of clients
(`writability-mode: internal-only`; replication and `dsreplication`
still write), the container stops reporting itself healthy, and the
server publishes `rejoining`, so no peer joins or initializes through it
— all of it across restarts too (`.replication-rejoin-pending` on the
volume, which names the backend). Once it rejoined, the backend takes
writes again, unless it was not `enabled` before; letting them in and
publishing `ready` are tried again for as long as the retries last, so
one failed `dsconfig` does not end a join that already got this far. A
container started on such a volume without the join logs that the
backend may still refuse writes, as no join runs to let them in. The
health status alone could not keep clients out: it turns unhealthy only
after the probe's retries. The check runs at the start of every round,
so a reset that could not take its own lock is tried again. A pending
volume is checked the same way, and stays `pending` meanwhile. An enable
that stopped half-way is taken down the same way. In the CI run behind
this (37201543789) its initialize of `cn=admin data` failed, which left
`BASE_DN` replicated and the server registered at its peers, but not in
its own `cn=admin data`; every later enable exited 5, "already
replicated", without registering it there, for all 30 attempts. A search
of its own `cn=admin data` that fails otherwise than with `noSuchObject`
decides nothing.
- **`run.sh`**: on a restart the health marker follows the upgrade
unless the volume is still pending — a server whose volume holds the
data of the topology is ready as soon as it serves. Gating it on its
peers would deadlock a whole-cluster restart under `OrderedReady`, and
gating it on the join would keep a seed that no peer joined yet (it has
no replication domain) unready once its root password was changed, since
the join binds with `ROOT_PASSWORD`. A volume whose bootstrap, join or
initialize never completed still carries the marker and waits for the
join, so none of these turns into a healthy server with bootstrap data
only; nor does a volume whose replication the join took down to enable
it anew.
- **`bootstrap/replicate.sh`** keeps the one-shot `srs`, `sdsr` and `rg`
paths as they are (deprecated in the README), with
`$ADMIN_PORT`/`$REPLICATION_PORT` in place of the hardcoded
`4444`/`8989` — so, like `simple`, they need one `ADMIN_PORT` on every
server.
- **`Dockerfile-alpine`** installs `coreutils`: the `timeout` of BusyBox
signals only the shell script that starts java, that of coreutils the
whole process group.
- **CI**: both image jobs run
`.github/scripts/docker-test-replication.sh` with their own image (and
check out `.github/scripts` for it). It covers the seed decision, a seed
after a cut-off reset letting client writes in again (its volume started
once without the join first: healthy, still refusing client writes, and
saying which backend), the retry while the peer is unreachable, the
initialize (entries the seed only ever imported), a member ready again
while its peer is down and not resetting its replication on a plain
restart, a seed ready again after its root password changed, a bootstrap
that failed after its import, an initialize that failed, a failed join
staying unhealthy across a restart, the seed losing its volume, a
scale-up, a scale-down after which the survivors have dropped the
removed server from every list, the removed server scaled up again on
its volume (with its own join lock held by hand it keeps its
replication; with both peers' locks held it waits for each, stays
unready, publishes `rejoining` and refuses client writes with 53, across
a restart too; with its own lock held while both peers are free it still
waits; a lock left under its own name is broken at once), two removed
servers scaled up again together (while the survivor's lock is held,
neither enables through the other; then both publish `ready` and
register with the survivor, and a backend the operator made read-only
stays so), a server whose own `cn=admin data` lacks its entry and one
that lacks `cn=Servers` altogether while their peers register them (each
takes its replication down and rejoins), a first peer giving up rather
than seeding next to a member that publishes no state, and next to a
peer that refuses the bind with `ROOT_PASSWORD` (with a retry interval
written `08`), three fresh servers started at once (and no join lock
left behind), a master named by its own address (kept by its replica
moved to a `REPLICATION_PEERS` of names), the shared-`/dev/shm` hygiene
of a Kubernetes pod, and the deprecated `sdsr` path with its retry on
exit 8. When a check fails, the job prints the log of every container
and, for each one still running, the tail of the server's errors log and
every detailed `dsreplication` log left in the instance.
- **README**: a Replication section documents the contract (all servers
share `BASE_DN`, `ROOT_USER_DN`, `ROOT_PASSWORD`, `ADMIN_PORT`,
`REPLICATION_PORT`), how a server recognises itself, the retry and
timeout knobs, the published state, the join lock, the seed rule with
its residual risk, and the cleanup.

No password reaches a command line (#1084): the tools read it from a
file on `/dev/shm` (or on `/tmp` under a name `run.sh` removes), and
`run.sh` removes what a killed join or `replicate.sh` leaves there, by
the `ADMIN_PORT` in the name, keeping the files of the other containers
of the pod.

Verified locally by running `.github/scripts/docker-test-replication.sh`
in full: at round 2 over Debian and Alpine images built the way the
image jobs build them, from a server package of the branch at
ff54b086c2; at rounds 3 to 5 over the Debian one, with the scripts of
the head laid over it (round 4: 2411 s; round 5: 8070 s on a slow host,
with every wait of the script given three times its budget, an enable
300 s and the health probe a 120 s timeout in the local copy only;
Alpine could not run locally at rounds 4 and 5). Every scenario passed.
At round 6 only the new half-enabled case ran locally, on its own, over
the Debian image with the scripts of the head laid over it: it passed at
the head (1644 s) and failed with `join.sh` of round 5. Round 7 changes
only the CI script and was not run locally. The docker jobs of this head
run the full script on both images over the current base (85b28b3d0f).
6 files modified
2 files added
2148 ■■■■■ changed files
.github/scripts/docker-test-replication.sh 718 ●●●●● diff | view | raw | blame | history
.github/workflows/build.yml 152 ●●●●● diff | view | raw | blame | history
opendj-packages/opendj-docker/Dockerfile 10 ●●●●● diff | view | raw | blame | history
opendj-packages/opendj-docker/Dockerfile-alpine 12 ●●●●● diff | view | raw | blame | history
opendj-packages/opendj-docker/README.md 142 ●●●●● diff | view | raw | blame | history
opendj-packages/opendj-docker/bootstrap/join.sh 976 ●●●●● diff | view | raw | blame | history
opendj-packages/opendj-docker/bootstrap/replicate.sh 62 ●●●●● diff | view | raw | blame | history
opendj-packages/opendj-docker/run.sh 76 ●●●●● diff | view | raw | blame | history