Skip to content

[#1086] Join replication in the background on every start of the Docker image - #1115

Merged
vharseko merged 9 commits into
OpenIdentityPlatform:masterfrom
vharseko:feature/1086-replication-join
Oct 6, 2026
Merged

vharseko merged 9 commits into
OpenIdentityPlatform:masterfrom
vharseko:feature/1086-replication-join

Conversation

@vharseko

@vharseko vharseko commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

Fixes #1086.

The container joined its replication topology once, during the first bootstrap only, through a single MASTER_SERVER it recognised by an unanchored grep of /etc/hosts, and a join that failed turned into a healthy, unreplicated server on the next restart, because a restart wrote the health marker right after upgrade -n. A StatefulSet cannot rely on any of that (#1086, discussion #1079).

The join becomes a background step next to the server, which stays PID 1 of the container (#1085):

  • bootstrap/join.sh (new) runs in the background on every start and serves OPENDJ_REPLICATION_TYPE=simple.
    • Peers come from a list: REPLICATION_PEERS=host1,host2,… (DNS names), which a chart derives from the StatefulSet ordinals; MASTER_SERVER keeps working as a one-element list. The server recognises itself by a name equal to hostname -f or to hostname -f cut at a dot (a pod listed as <sts>-N.<svc> has the FQDN <sts>-N.<svc>.<ns>.svc.…), and by its own addresses and the names /etc/hosts gives them — compared whole. The grep of /etc/hosts it replaces took opendj-1 for opendj-10, a replica for its master when an --add-host named the master, and a master that lost its volume for a replica of itself.
    • Membership is decided from what is there, never from an exit code: exit 5 of dsreplication enable covers "already replicated" and "BASE_DN not found on one of the servers" alike, so the step instead checks the replication domain for BASE_DN in cn=config and the registration in cn=admin data, and until both are there it tries again, a bounded, configurable number of times, each enable under its own timeout (REPLICATION_ATTEMPT_TIMEOUT), since it can hang on a peer that stops mid-operation. The retry and timeout knobs are read as decimal whole numbers (a leading zero does not turn them octal); a value that is not one falls back to its default. dsreplication initialize — a full import — has a bound of its own, REPLICATION_INITIALIZE_TIMEOUT, none by default.
    • Joins that involve the same server take turns: two dsreplication enable runs through the same server at once break each other — both create the replication server a seed does not have yet, the loser fails half-way (exit 17) and every later enable exits 5 without ever completing its membership. An enable therefore holds a lock on the peer it runs through and on its own server — the entry cn=Docker Join Lock,cn=config, which only one add creates — and a server that takes its own replication down holds its own. Neither lock is waited for, and rounds wait a random part of REPLICATION_RETRY_INTERVAL on top of it, so two joins that each hold the other's turn do not keep meeting. A lock left by a killed join is broken once it is older than REPLICATION_ATTEMPT_TIMEOUT plus a minute, or at once by a later start of the server that left it, by a delete that asserts the value it read.
    • Whose data the topology carries follows what each volume went through: every fresh volume holds entries (ADD_BASE_ENTRY/SAMPLE_DATA), and two freshly bootstrapped volumes even share a generation ID, so "does BASE_DN have entries" decides nothing. run.sh marks a volume before it bootstraps it, and the join publishes that to the peers in a local, non-replicated entry, cn=Docker Join,cn=config: pending until the volume received the data of the topology, ready afterwards — published before the marker goes, so the two never tell different stories — and rejoining while a volume that holds that data has its replication taken down to enable it anew. Every server joins only through a ready peer, and a pending one also initializes only from one — however many restarts that takes — and cross-checks the generation IDs. So servers may also start together (Compose, podManagementPolicy: Parallel) without one taking another's bootstrap data for the topology's, a bootstrap that failed or was killed half-way is initialized from the topology rather than taken for its data, and two servers whose replication is down — rejoining, pending or still bootstrapping — never register only each other and serve a topology of their own beside the survivors. Where only MASTER_SERVER is set, a peer that publishes no state (an older image) counts as ready.
    • The seed is a rule: only the first entry of REPLICATION_PEERS may declare that there is no topology and seed it with its own data — at once when every other peer answers and is pending, otherwise only once its retries are exhausted while no other peer answers that may hold the data: a ready or rejoining one, one that publishes no state but replicates BASE_DN (a server of an earlier image), or one that refuses the bind with ROOT_PASSWORD (only a server past its bootstrap has another root password). Anyone else stays unhealthy, which surfaces a lost topology instead of forking it. A -0 that lost its volume therefore rejoins through the survivors and takes the data of the topology back, instead of bootstrapping an empty, unreplicated server behind the same Service. The residual risk — every other server down and the first peer's volume lost — is documented in the README.
    • Leaving is handled by the survivors: with REPLICATION_PEERS set explicitly, a joined server removes every server registered in cn=admin data but no longer listed, and prunes it from every replication-server list it holds — that of its replication server and those of the domains of BASE_DN, cn=schema and cn=admin data. A server registered by an address (a master that MASTER_SERVER=<address> named) is left alone, as nothing tells it from a listed name. It also adds every listed, registered peer its lists lack, so joins that ran at the same time cannot leave two replication servers unaware of each other. dsreplication disable in a preStop hook could not do this: it fires on every termination — rolling update, drain — and changes only the servers it can reach, while OpenDJ 4 has no cleanup subcommand for a dead one. With only MASTER_SERVER set nothing is removed or added, so servers joined by hand stay.
    • Coming back is handled by the returning server: scaled up again on the volume it kept, a server the survivors removed still replicates, and dsreplication enable between two servers whose cn=admin data is replicated registers nobody. When at least one answering peer that replicates BASE_DN lists the registrations and none of them registers it — a search that fails decides nothing — it takes its own replication configuration down (dsreplication disable --disableAll) and enables again, which registers it. From the first change until that enable registered it, writes it took would replicate nowhere. So before anything changes, the backend of BASE_DN refuses the writes of clients (writability-mode: internal-only; replication and dsreplication still write), the container stops reporting itself healthy, and the server publishes rejoining, so no peer joins or initializes through it — all of it across restarts too (.replication-rejoin-pending on the volume, which names the backend). Once it rejoined, the backend takes writes again, unless it was not enabled before; letting them in and publishing ready are tried again for as long as the retries last, so one failed dsconfig does not end a join that already got this far. A container started on such a volume without the join logs that the backend may still refuse writes, as no join runs to let them in. The health status alone could not keep clients out: it turns unhealthy only after the probe's retries. The check runs at the start of every round, so a reset that could not take its own lock is tried again. A pending volume is checked the same way, and stays pending meanwhile. An enable that stopped half-way is taken down the same way. In the CI run behind this (37201543789) its initialize of cn=admin data failed, which left BASE_DN replicated and the server registered at its peers, but not in its own cn=admin data; every later enable exited 5, "already replicated", without registering it there, for all 30 attempts. A search of its own cn=admin data that fails otherwise than with noSuchObject decides nothing.
  • run.sh: on a restart the health marker follows the upgrade unless the volume is still pending — a server whose volume holds the data of the topology is ready as soon as it serves. Gating it on its peers would deadlock a whole-cluster restart under OrderedReady, and gating it on the join would keep a seed that no peer joined yet (it has no replication domain) unready once its root password was changed, since the join binds with ROOT_PASSWORD. A volume whose bootstrap, join or initialize never completed still carries the marker and waits for the join, so none of these turns into a healthy server with bootstrap data only; nor does a volume whose replication the join took down to enable it anew.
  • bootstrap/replicate.sh keeps the one-shot srs, sdsr and rg paths as they are (deprecated in the README), with $ADMIN_PORT/$REPLICATION_PORT in place of the hardcoded 4444/8989 — so, like simple, they need one ADMIN_PORT on every server.
  • Dockerfile-alpine installs coreutils: the timeout of BusyBox signals only the shell script that starts java, that of coreutils the whole process group.
  • CI: both image jobs run .github/scripts/docker-test-replication.sh with their own image (and check out .github/scripts for it). It covers the seed decision, a seed after a cut-off reset letting client writes in again (its volume started once without the join first: healthy, still refusing client writes, and saying which backend), the retry while the peer is unreachable, the initialize (entries the seed only ever imported), a member ready again while its peer is down and not resetting its replication on a plain restart, a seed ready again after its root password changed, a bootstrap that failed after its import, an initialize that failed, a failed join staying unhealthy across a restart, the seed losing its volume, a scale-up, a scale-down after which the survivors have dropped the removed server from every list, the removed server scaled up again on its volume (with its own join lock held by hand it keeps its replication; with both peers' locks held it waits for each, stays unready, publishes rejoining and refuses client writes with 53, across a restart too; with its own lock held while both peers are free it still waits; a lock left under its own name is broken at once), two removed servers scaled up again together (while the survivor's lock is held, neither enables through the other; then both publish ready and register with the survivor, and a backend the operator made read-only stays so), a server whose own cn=admin data lacks its entry and one that lacks cn=Servers altogether while their peers register them (each takes its replication down and rejoins), a first peer giving up rather than seeding next to a member that publishes no state, and next to a peer that refuses the bind with ROOT_PASSWORD (with a retry interval written 08), three fresh servers started at once (and no join lock left behind), a master named by its own address (kept by its replica moved to a REPLICATION_PEERS of names), the shared-/dev/shm hygiene of a Kubernetes pod, and the deprecated sdsr path with its retry on exit 8. When a check fails, the job prints the log of every container and, for each one still running, the tail of the server's errors log and every detailed dsreplication log left in the instance.
  • README: a Replication section documents the contract (all servers share BASE_DN, ROOT_USER_DN, ROOT_PASSWORD, ADMIN_PORT, REPLICATION_PORT), how a server recognises itself, the retry and timeout knobs, the published state, the join lock, the seed rule with its residual risk, and the cleanup.

No password reaches a command line (#1084): the tools read it from a file on /dev/shm (or on /tmp under a name run.sh removes), and run.sh removes what a killed join or replicate.sh leaves there, by the ADMIN_PORT in the name, keeping the files of the other containers of the pod.

Verified locally by running .github/scripts/docker-test-replication.sh in full: at round 2 over Debian and Alpine images built the way the image jobs build them, from a server package of the branch at ff54b08; at rounds 3 to 5 over the Debian one, with the scripts of the head laid over it (round 4: 2411 s; round 5: 8070 s on a slow host, with every wait of the script given three times its budget, an enable 300 s and the health probe a 120 s timeout in the local copy only; Alpine could not run locally at rounds 4 and 5). Every scenario passed. At round 6 only the new half-enabled case ran locally, on its own, over the Debian image with the scripts of the head laid over it: it passed at the head (1644 s) and failed with join.sh of round 5. Round 7 changes only the CI script and was not run locally. The docker jobs of this head run the full script on both images over the current base (85b28b3).

@vharseko vharseko added enhancement docker replication CI kubernetes Kubernetes / Helm / OpenShift deployment labels Sep 29, 2026

@maximthomas maximthomas left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

praise: The join reads membership from the server's own state instead of from dsreplication's exit codes, and the root password stays off every command line.

  • is_member (bootstrap/join.sh:135-140) needs both the BASE_DN domain in cn=config and the cn=admin data registration, so the two meanings of exit 5 no longer decide anything.
  • Every tool reads --bindPasswordFile from mktemp -p /dev/shm "opendj-join.$ADMIN_PORT.XXXXXX" (join.sh:100), run.sh:45 removes only the leftovers for this container's ADMIN_PORT, and build.yml:593 greps the scripts statically for a password flag.
  • The anchored self-recognition (join.sh:111-116) replaces the /etc/hosts grep that took opendj-1 for opendj-10.

issue (blocking): A failed dsreplication initialize is reported as success, so the container turns healthy without the topology's data.

opendj-packages/opendj-docker/bootstrap/join.sh:198-208

rc=$? right after if initialize_from …; then …; fi holds the exit status of the if itself, which is 0 when the condition failed and there is no else. After the last attempt, return $rc returns 0, and joined() then runs cleanup_departed and touches $BOOTSTRAP_COMPLETE. So a replica whose every initialize failed reports itself healthy while it holds only its bootstrap entries. This happens when the source is unreachable, or when each attempt is killed as described in the third comment. Every retry line also logs "exited with 0". A bash probe of the same shape prints attempt 1 rc=0, attempt 2 rc=0, g returned 0.

    initialize_from "$source"
    rc=$?
    if [ "$rc" -eq 0 ]; then
      rm -f "$INITIALIZE_PENDING"
      # ... generation ID cross-check as now ...
      return 0
    fi
    [ "$i" -eq "$REPLICATION_RETRY_COUNT" ] && return $rc

issue (blocking): On a restart, run.sh writes the health marker even when $INITIALIZE_PENDING is still on the volume.

opendj-packages/opendj-docker/run.sh:142

The gate only looks for ds-cfg-replication-domain in config.ldif. On a first start where the enable succeeded but the initialize failed or was killed, the container is correctly unhealthy. The next start is healthy immediately, with bootstrap-only or half-imported data, while join.sh re-initializes in the background. If no peer answers generation_id, ensure_initialized "" returns 1 (join.sh:182-195) and the server stays healthy but unreplicated. The README ("healthy only once the join succeeded - across restarts too") and the comment at run.sh:136-141 promise the opposite.

    if ! join_requested || { grep -q "ds-cfg-replication-domain" ./data/config/config.ldif && [ ! -f "$INITIALIZE_PENDING" ]; }; then

With this change, under OrderedReady a -0 that still carries the marker waits for a peer. That is the honest state for a server that never received the topology's data, and it deserves a line in the README.


issue (blocking): dsreplication initialize is killed after REPLICATION_ATTEMPT_TIMEOUT (120 s by default), so a directory whose total update takes longer than that can never be initialized.

opendj-packages/opendj-docker/bootstrap/join.sh:160-165, opendj-packages/opendj-docker/README.md:196

A total update is a full import of BASE_DN. If it needs more than 120 s, every attempt is killed before its task finishes, and the next attempt requests another full import, so no attempt can ever succeed. $INITIALIZE_PENDING is never cleared and every start imports again. Combined with the first comment, the container turns healthy after 30 failures; combined with the second, on the next restart. The README justifies the bound with a hanging dsreplication enable, while the base replicate.sh ran initialize without a bound. With the defaults, scaling up a StatefulSet over any non-trivial directory hits this.

initialize_from() {
  dsreplication initialize --baseDN "$BASE_DN" \
    --adminUID admin --adminPasswordFile "$PASSWORD_FILE" \
    --hostSource "$1" --portSource "$ADMIN_PORT" \
    --hostDestination "$MYHOSTNAME" --portDestination "$ADMIN_PORT" -X -n
}

Or: add a separate REPLICATION_INITIALIZE_TIMEOUT, unset by default, and make README:196 say which commands the attempt timeout covers.


issue (blocking): cleanup_departed leaves a departed server in the cn=schema (and cn=admin data) replication domains, so the scale-down check cannot pass.

opendj-packages/opendj-docker/bootstrap/join.sh:263-279, .github/workflows/build.yml:694

join.sh passes no --noSchemaReplication, so dsreplication enable also configures the cn=schema domain with the same replication-server list, on every server in the topology. On the first enable it does the same for the cn=admin data domain (ReplicationCliMain :4761-4767, :4850-4859, :6329-6391). The cleanup only prunes the replication-server entry and the BASE_DN domain. After the scale-down, dj-0 and dj-1 therefore keep dj-2:8989 in their cn=schema domain. The check at build.yml:694 greps every replication-domain entry for dj-2, so it loops until timeout 2m fails the step. Outside CI, the survivors keep dialling the removed server. The docker jobs have not run at this head yet, and scale-down is the one scenario missing from the description's "Verified locally" list.

  search localhost --baseDN "cn=config" --searchScope sub "(objectClass=ds-cfg-replication-domain)" cn \
    | awk '/^cn: /{print substr($0,5)}' \
    | while read -r domain; do
      search localhost --baseDN "cn=config" --searchScope sub \
        "(&(objectClass=ds-cfg-replication-domain)(cn=$domain))" ds-cfg-replication-server \
        | awk '/^ds-cfg-replication-server: /{print $2}' \
        | while read -r value; do
          host=${value%:*}
          if ! in_peers "$host" && ! is_self "$host"; then
            echo "join: removing departed replication server $value from replication domain $domain"
            dsconfig set-replication-domain-prop --provider-name "Multimaster Synchronization" \
              --domain-name "$domain" --remove "replication-server:$value" \
              --hostname localhost --port "$ADMIN_PORT" --bindDN "$ROOT_USER_DN" \
              --bindPasswordFile "$PASSWORD_FILE" --trustAll --no-prompt || true
          fi
        done
    done

issue (blocking): The build-docker-alpine "Docker test replication" step still exercises the simple path of replicate.sh, which run.sh no longer calls.

.github/workflows/build.yml:1115-1120, :1142

Only build-docker's step was rewritten. In build-docker-alpine, test_replica runs with MASTER_SERVER=dj-master OPENDJ_REPLICATION_TYPE=simple, and run.sh now hands that to join.sh. The step then waits 5 minutes for "Will sleep for a bit", which only replicate.sh:82 prints. The wait exits 124 on every run. The alpine image also gets no coverage of join.sh, and its password grep at :1098 does not list join.sh either. Port build.yml:581-708 into the alpine job, or move the step into a script under .github/scripts/ that both jobs call with their own image.


issue (non-blocking): is_member never finds the seed's cn=admin data entry when hostname -f is longer than the listed name, which is the case on Kubernetes.

opendj-packages/opendj-docker/bootstrap/join.sh:138-139

The seed never runs an enable of its own. A peer's enable registers it under the name that peer connected with (ServerDescriptor :648, HOST_NAME = conn.getHostPort().getHost()), which is <sts>-0.<svc>, while hostname -f in the pod is the FQDN. So on every restart -0 fails is_member, runs REPLICATION_RETRY_COUNT rounds of enables that exit 5, never runs cleanup_departed, and leaves through the seed branch. Health is not affected. CI hides the problem because --hostname dj-N equals the listed name.

is_member() {
  local host
  search localhost --baseDN "cn=config" --searchScope sub \
    "(&(objectClass=ds-cfg-replication-domain)(ds-cfg-base-dn=$BASE_DN))" 1.1 | grep -q "^dn:" || return 1
  for host in $(search localhost --baseDN "cn=Servers,cn=admin data" --searchScope one "(objectClass=*)" hostname \
      | awk '/^hostname: /{print $2}'); do
    is_self "$host" && return 0
  done
  return 1
}

issue (non-blocking): The top-level docker logs X 2>&1 | grep -q … checks run under pipefail and can fail even after a match.

.github/workflows/build.yml:617, :625, :628, :666, :667

shell: bash runs as bash -eo pipefail. grep -q exits on the first match, so the next frame docker logs writes gets SIGPIPE (141), and the pipeline is non-zero although the line was found. Checks :617, :625, :628 and :666 then fail spuriously, and the negative check at :667 passes silently exactly when the seed line is present. The polls inside timeout … bash -c are not affected, because that inner bash has no pipefail. The failure rate on the runner was not measured.

          docker logs dj-0 2>&1 | grep "seeding it with this server's data" >/dev/null || { echo "::error::dj-0 did not seed the topology"; false; }

issue (non-blocking): On the Alpine image, timeout is BusyBox's, and it kills only the sh launcher of dsreplication, not the JVM.

opendj-packages/opendj-docker/bootstrap/join.sh:143, :161, opendj-packages/opendj-docker/Dockerfile-alpine:60

Dockerfile-alpine installs bash "$JDK" and no coreutils. BusyBox timeout signals only the PID it execs, and _mixed-script.sh:69 starts java without exec. A timed-out attempt therefore leaves its JVM running under the container's memory limit while the next attempt starts, so the per-attempt bound does not hold on Alpine. This was not probed in the image.

 && apk add bash coreutils "$JDK" \

issue (non-blocking): generation_id can return another server's generation ID, because (connected-to=*) does not single out the domain's own monitor entry.

opendj-packages/opendj-docker/bootstrap/join.sh:152-158

DataServerHandler (:244) and ServerHandler (:559, :591) put connected-to, domain-name and generation-id on the replication server's entry for every connected directory server. awk … exit takes the first match, so the cross-checks at :174-178 and :200-204 may compare the wrong entry. replayed-updates is published only by ReplicationMonitor (:79). The comment at :152-153 needs the same correction.

    "(&(domain-name=$BASE_DN)(replayed-updates=*))" generation-id \

issue (non-blocking): The README says plain docker run passes the container names, but that only works when --hostname equals --name.

opendj-packages/opendj-docker/README.md:77-80, opendj-packages/opendj-docker/bootstrap/join.sh:65

MYHOSTNAME defaults to hostname -f, which is the container ID when --hostname is not given. So in a container named dj-0, is_self dj-0 is false: the container never takes the seed rule and tries to enable with itself under a second name. CI passes --hostname "$1" (build.yml:606). What that self-enable leaves behind was not run.

no registry; plain `docker run` passes the container names, each container started with
`--hostname` equal to its `--name` (or with `MYHOSTNAME` set to it). `MASTER_SERVER` keeps working

question (non-blocking): Is a parallel first start (Compose, podManagementPolicy: Parallel) meant to be supported?

opendj-packages/opendj-docker/bootstrap/join.sh:319-321, opendj-packages/opendj-docker/run.sh:179

run.sh marks every bootstrapped volume, including the first peer's, and only the seed branch (join.sh:335-337) removes that marker. When the peers start together, B's enable through A makes A a member. A then initializes itself from B while B initializes from A, and the topology ends up with whichever import lands last, not with the first peer's data as the README says. If the bootstrap data is identical everywhere, the only cost is duplicate total updates. If parallel starts are out of scope, this is minor, and the README should say the first start has to be ordered. If they are supported, it is a major issue.


question (non-blocking): Was dropping IP and alias recognition of MASTER_SERVER intended?

opendj-packages/opendj-docker/bootstrap/join.sh:111-116

The base replicate.sh treated any /etc/hosts line containing MASTER_SERVER (an IP, hostAliases, --add-host) as "I am the master". is_self compares only with hostname -f and its first label. A master started with OPENDJ_REPLICATION_TYPE=simple and its own IP or an alias as MASTER_SERVER now tries to enable against itself, cannot take the seed rule, and stays unhealthy, where the old image was healthy. If this was intended, a compatibility line in the README (list the hostname, or set MYHOSTNAME) covers it.


suggestion (non-blocking): Pin the initialize with data that never went through the changelog. The log line and the current data checks both pass without any initialize.

.github/workflows/build.yml:626-631, :666-668

"initializing from dj-N" is printed before the call (join.sh:197). ou=replicated and ou=replicated2 were written over LDAP, so dj-1's changelog replays them to a server that was never initialized, since two fresh volumes share a generation ID. A no-op initialize_from that returns 0 keeps the step green (not run).

          # dj-0 bootstraps with -e SAMPLE_DATA=10 in place of ADD_BASE_ENTRY (start_node takes the extra -e)
          wait_healthy dj-1
          docker exec dj-1 /opt/opendj/bin/ldapsearch --hostname localhost --port 1636 --bindDN "cn=Directory Manager" --bindPassword "$ROOT_PASSWORD" --useSsl --trustAll --baseDN "uid=user.0,ou=People,dc=example,dc=com" --searchScope base "(objectClass=*)" 1.1

Pin: dj-1 must serve an entry that the seed only ever imported; the no-op mutant fails the search.


suggestion (non-blocking): Hold the "dj-x is not healthy" assertions over a probe cycle. A single inspect right after the failure line reads "starting" whatever join.sh did.

.github/workflows/build.yml:653, :657

The health status changes only after a probe that starts once the marker exists (the 5 s start interval plus the probe's ldapsearch). So a join.sh that wrote $BOOTSTRAP_COMPLETE on its failure path would pass most runs (not run). The restart check does catch the base behaviour, which wrote the marker after the upgrade, but not a marker written at the moment of failure.

          for _ in $(seq 1 12); do
            [ "$(docker inspect --format='{{.State.Health.Status}}' dj-x)" != healthy ] || { echo "::error::dj-x reports itself healthy although its join failed"; false; }
            sleep 5
          done

suggestion (non-blocking): Keep the pins on run.sh's removal of opendj-replicate.$ADMIN_PORT.* and on replicate.sh's retry on exit 8. The rewritten step dropped both.

.github/workflows/build.yml:643-647, :700

The step now plants only opendj-join.* files, and dj-sdsr runs without a network disconnect. The alpine copy still pins both behaviours, but it fails earlier (see above). Deleting the opendj-replicate glob from run.sh:45, or breaking retry 8, would stay green.

          docker exec dj-shm sh -c ': >/dev/shm/opendj-join.4444.killed && : >/dev/shm/opendj-replicate.4444.killed && : >/dev/shm/opendj-join.5444.other'
          docker run -d --memory="512m" --network test_replication --ipc=container:dj-shm --name dj-shm-probe --hostname dj-shm-probe -e ROOT_PASSWORD="$ROOT_PASSWORD" "$IMAGE"
          timeout 1m bash -c 'while docker exec dj-shm sh -c "ls /dev/shm/opendj-*.4444.* >/dev/null 2>&1"; do sleep 2; done'

suggestion (non-blocking): No step in CI pins "membership from what is there, not the exit code", although the comment at build.yml:619-620 says this one does.

.github/workflows/build.yml:619-625, opendj-packages/opendj-docker/bootstrap/join.sh:317-319

A mutant enable_through "$peer" && is_member stays green. While dj-0 is off the network the enable fails without membership, afterwards it exits 0 with membership, and restarted members never call enable. A deterministic case of exit 5 with membership needs two enables of the same pair. Failing that, narrow the comment to the retry the step actually pins.

          # a joining server tries again while its peer is unreachable

suggestion (non-blocking): is_self and in_peers treat any two names with the same first label as one server.

opendj-packages/opendj-docker/bootstrap/join.sh:115, :219-226

With the same StatefulSet name in two namespaces (opendj-0.opendj.east…, opendj-0.opendj.west…), west-0 skips east-0 and counts the first peer as itself, so it may seed a second topology. With an IPv4 list and MYHOSTNAME set to an IP, every name collapses to "10" and every server seeds at once. The documented list shapes never collide, but a guard would stop the other shapes from forking silently.

  [ "$peer" = "$host" ] && return 0
  case $peer in *[!0-9.]*) ;; *) return 1 ;; esac   # an IPv4 address has no host label
  [ "${peer%%.*}" = "${host%%.*}" ]

nitpick (non-blocking): dj-sdsr still runs with --rm, so the ERR trap cannot print its log if its server exits.

.github/workflows/build.yml:700

cleanup() already removes the container.

          docker run -d --memory="512m" --network test_replication --name dj-sdsr --hostname dj-sdsr -e ROOT_PASSWORD="$ROOT_PASSWORD" -e MASTER_SERVER=dj-0 -e OPENDJ_REPLICATION_TYPE=sdsr "$IMAGE"

vharseko added a commit to vharseko/OpenDJ that referenced this pull request Sep 29, 2026
…gy's data, and clean every replication list

Review round 1 of OpenIdentityPlatform#1115:

- join.sh: a failed dsreplication initialize is no longer taken for a success (the
  exit status was that of an if without else), and it is bounded by its own
  REPLICATION_INITIALIZE_TIMEOUT, none by default; REPLICATION_ATTEMPT_TIMEOUT bounds
  the enable only, and the bound sits on the dsreplication call, as timeout cannot run
  a shell function.
- join.sh publishes pending/ready in the local entry cn=Docker Join,cn=config; a pending
  server enables and initializes only through a ready peer, the first peer seeds at
  once when every other one is pending and after the retries only while no peer that
  could hold the data answers. Fresh servers started together no longer take each
  other's bootstrap data for the topology's.
- The cleanup prunes departed servers from every replication list (replication server,
  BASE_DN, cn=schema, cn=admin data), and the lists get every listed, registered peer
  they lack; dsconfig inside the read loops reads /dev/null.
- Self-recognition: a name equal to hostname -f or cut from it at a dot, and the
  container's own addresses and their /etc/hosts names, compared whole; is_member
  checks every registered hostname; the generation ID comes from the domain's own
  monitor entry.
- run.sh stays unhealthy on a restart while $INITIALIZE_PENDING is on the volume.
- Dockerfile-alpine installs coreutils, whose timeout signals the whole process group.
- CI: both image jobs run .github/scripts/docker-test-replication.sh, which adds the
  initialize pin, held unhealthy checks, the full-list scale-down check, three fresh
  servers started at once, a master named by its own address, and the replicate.sh
  pins on the /dev/shm glob and the retry on exit 8.
- README documents all of the above.
@vharseko

Copy link
Copy Markdown
Member Author

Round 1 is in 89627af. All five blocking issues were real, and so were the non-blocking ones; two of them are fixed differently from the snippet, and both questions got an answer in code rather than in the README. Point by point:

Blocking

  1. Failed initialize reported as success — fixed. initialize_from is no longer the condition of an if: its exit status is taken directly, the marker is removed and the state flips to ready only on 0, and a failure is logged with its real code and tried again on the next round (join.sh, the pending loop).
  2. Health marker on a restart with $INITIALIZE_PENDING on the volume — fixed with your condition: run.sh writes the marker on a restart only when the replication domain is configured and the marker is gone. The README now says that a volume that joined but never received the data keeps waiting for a ready peer on every start, and what that means for a -0 under OrderedReady.
  3. Initialize killed after REPLICATION_ATTEMPT_TIMEOUT — fixed with the second variant: REPLICATION_ATTEMPT_TIMEOUT bounds dsreplication enable only, and a new REPLICATION_INITIALIZE_TIMEOUT bounds dsreplication initialize, 0 (no bound) by default. The comment next to it records why a bound is harmful there: a killed client leaves the task running on the server, and the next attempt asks for another full import. README table updated for both.
  4. cleanup_departed left the departed server in cn=schema and cn=admin data — fixed as you suggested, over every list the server holds: one search returns the replication server entry and every replication domain, and each value that is neither listed nor this server is removed from its own list. Verified: after the scale-down both survivors logged the removal of dj-2:8989 from the replication server list and from the domains of dc=example,dc=com, cn=schema and cn=admin data, and the check over all of them passes. The dsconfig calls inside these while read loops now read /dev/null: they could otherwise consume the lines the loop was still to read.
  5. build-docker-alpine still tested the removed simple path of replicate.sh — fixed by the second option: the step is now .github/scripts/docker-test-replication.sh <image>, and both jobs call it with their own image, so the two can no longer drift. The password grep lists join.sh for both images.

Non-blocking

  1. is_member on Kubernetes — fixed, together with 17 (below): is_member now walks the hostnames registered in cn=admin data and asks is_self of each, as in your snippet, and is_self accepts a name that is hostname -f cut at a dot, so the <sts>-0.<svc> a peer registered the seed under is recognised against the pod's FQDN.

  2. SIGPIPE under pipefail — fixed. Every check of a container log goes through logs_have, whose grep reads the whole input (>/dev/null, not -q); the waits are loops in the script itself rather than timeout … bash -c.

  3. BusyBox timeout on Alpine — fixed: Dockerfile-alpine installs coreutils, and the CI script checks that the image's timeout is the one of coreutils. While at it I found that timeout cannot run a shell function (it execs its argument; exit 127), so the bound now sits on the dsreplication call inside enable_through/initialize_from through a small bounded helper, which also implements "0 = no bound".

  4. generation_id could read a handler's entry — fixed with your filter, (&(domain-name=$BASE_DN)(replayed-updates=*)), and the comment now says why.

  5. README: docker run names — fixed with your wording (--hostname equal to --name, or MYHOSTNAME).

  6. Parallel first start — supported now; this became the largest change of the round. The pending marker was local, so a peer could not tell "fresh like me" from "holds the topology's data", and without that distinction the reborn-seed case and the parallel start look the same. Each server now publishes its state to its peers in a local, non-replicated entry, cn=Docker Join,cn=config (description: pending|ready, ds-cfg-branch + extensibleObject; it survives restarts and upgrade, and lives on the volume next to the marker). The rules:

    • a pending server enables and initializes only through a ready peer, so two fresh servers never exchange bootstrap data;
    • where the list is only MASTER_SERVER, a peer that publishes no state (an older image) counts as ready, as the old replicate.sh trusted its master;
    • the first peer seeds at once when every other peer answers and is pending, and otherwise only once its retries are exhausted and no peer that could hold the data answers — the latter is new too: before, it seeded after the retries even with a ready peer present;
    • joins that ran at the same time can leave two replication servers that know nothing of each other; with an explicit REPLICATION_PEERS each server therefore adds every listed, registered peer that its lists lack, under the name it is registered by (so nobody is listed twice), after it already reports itself healthy.

    CI starts three fresh volumes at once, the first with SAMPLE_DATA=10: dj-p0 seeds, dj-p1/dj-p2 initialize from dj-p0 only and hold uid=user.0, every replication server lists the two others, and a change on dj-p1 reaches both.

  7. IP and alias recognition of MASTER_SERVER — restored, anchored: is_self also accepts one of the container's own non-loopback addresses (hostname -i and the /etc/hosts lines that name it) and any name /etc/hosts gives those addresses, compared whole — so an --add-host naming another server no longer makes a replica a master. CI starts a master with MASTER_SERVER set to its own address and a replica joining through that address. The README asks for DNS names in an explicit REPLICATION_PEERS: servers register in cn=admin data by name, and a peer listed by address alone would be removed by the cleanup.

  8. Pin the initialize — done as suggested: dj-0 seeds with SAMPLE_DATA=10, and dj-1 must serve uid=user.0,ou=People,dc=example,dc=com; the reborn dj-0 and the parallel replicas are checked the same way.

  9. Hold the "not healthy" assertions — done: stays_unhealthy watches the container for 75 s, more than two probe cycles of the 30 s interval.

  10. Pins on opendj-replicate.$ADMIN_PORT.* and on the retry on exit 8 — restored: the shared /dev/shm holds a killed opendj-replicate.4444.* file next to the join's, and dj-sdsr loses dj-0 while it sleeps before its first try and must log exited with 8, trying again. The old "second replicate.sh exits 5" pin does not carry over to sdsr — a second enable of a replica without a replication server succeeds and re-initializes — so it now runs replicate.sh for a base DN neither server holds: exit 5, nothing retried, nothing initialized.

  11. "membership from what is there" pin — comment narrowed as you proposed: no deterministic case of an enable failing with membership present exists without a race.

  12. is_self/in_peers collapsing names with the same first label — fixed differently from the snippet: the first-label rule is gone. Two names are one server when they are equal or one of them is the other cut at a dot, and addresses only when equal; the snippet's guard would have kept opendj-0.opendj.east equal to opendj-0.opendj.west.

  13. dj-sdsr with --rm — removed.

Verified locally with the branch's bootstrap/, run.sh and healthcheck.sh layered over the published Debian and Alpine images (the round changes no server code): on Debian, every scenario of the script up to and including dj-sdsr joining through its retry on exit 8 passed with this join.sh - seed, retry while the peer is unreachable, the uid=user.0 initialize pin, restart of both members, /dev/shm hygiene, dj-x held unhealthy across a restart, the reborn seed, scale-up, and the scale-down with dj-2 gone from all four lists on both survivors. That run stopped at the second replicate.sh pin, which I then reworked (see 15), before reaching the master named by its address and the parallel start; and the Alpine run so far passed up to the restart of both members (its /dev/shm check then caught a fault of the check itself, fixed in this push: it matched the probe's own live setup password file). Full runs of the script on both images are in progress; I will post their outcome here, and the docker jobs of this PR run the same script.

@vharseko

Copy link
Copy Markdown
Member Author

Follow-up to the round above, in a82819f:

  • The CI script's ERR trap now acts in the main shell only. set -E hands the trap to every $(...), so when the reworked replicate.sh pin (15) captured the output of an enable that is meant to fail, the trap dumped the container logs into the captured text - where dj-sdsr's own "trying again" matched - and removed the containers under the running test. The Alpine run reached this pin with every earlier scenario green and failed exactly there, with the expected exit 5. 89627af still carries the fault, so its docker jobs will fail at that pin.
  • join.sh waits for the peers to register with one search a round instead of one per peer: every ldapsearch starts a JVM.

Verification status, stated plainly: the scenarios after that pin - the master named by its own address and the three servers started at once - have not completed a run yet on either image. The later local runs could not tell anything: the host was saturated by unrelated builds (load average up to ~300 on 8 cores), and the containers missed their timeouts from the first seed on. The docker jobs of this head are the first clean run of the full script; I will report on them here.

@maximthomas maximthomas left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

praise: Every round-1 blocking issue is fixed where it lived, and the pending/ready state gives the join a real answer to "whose data is this".

  • join.sh:551-552 reads the exit status of initialize_from directly, and REPLICATION_INITIALIZE_TIMEOUT (join.sh:88, used at :287) no longer shares the enable's 120 s bound.
  • replication_server_lists (join.sh:349-360) runs one search over every ds-cfg-replication-server and ds-cfg-replication-domain entry, so cleanup_departed prunes cn=schema and cn=admin data as well.
  • is_self / names_match compare names whole and addresses only for equality, replacing the unanchored /etc/hosts grep.

issue (blocking): Neither image job checks out .github/scripts, so the "Docker test replication" step exits 127 before any check runs.

.github/workflows/build.yml:585, :966, checkouts at :479-481, :859-861

Both jobs check out only sparse-checkout: .github/benchmark, and .github/scripts is outside that cone. In run 36587186866 at a82819f, job 109537652639 logs .github/scripts/docker-test-replication.sh: No such file or directory and Process completed with exit code 127, and the alpine job fails the same way. No scenario of the script has run in CI at this head, and merging would turn both cells red on master.

      - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
        with:
          sparse-checkout: |
            .github/benchmark
            .github/scripts

issue (blocking): A volume whose first bootstrap failed or was killed has no $INITIALIZE_PENDING, so on its next start it publishes ready, joins, and turns healthy with bootstrap-only data.

opendj-packages/opendj-docker/run.sh:170-182, opendj-packages/opendj-docker/bootstrap/join.sh:482-485, :314-321

run.sh touches the marker only when BOOTSTRAPPED=true. setup.sh can exit 1 after ./data/config already exists (create-backend, makeldif and import-ldif all end in || exit 1), and a container killed mid-import never reaches :181. The restart road finds no domain and starts join.sh. join.sh sees no marker, which it reads as "never bootstrapped by run.sh", so it publishes ready, enables through a peer, and joined() writes the health marker over a partial userRoot that was never initialized. A later pending peer listed ahead of the other ready ones then initializes from it in trusted_source. README:59-60 says a failed bootstrap never reports healthy. Not run: dsreplication enable over a partially imported backend (it needs the configured base DN, not the entries).

BOOTSTRAPPED=true
# marked first, so that a bootstrap that fails or is killed half-way never
# leaves a volume that counts as holding the data of the topology
if join_requested; then
  touch "$INITIALIZE_PENDING"
fi
if ! sh "${BOOTSTRAP}"; then

Then drop the touch at run.sh:180-182.


issue (blocking): On the non-pending road the first peer seeds once its retries run out, even while a ready peer answers.

opendj-packages/opendj-docker/bootstrap/join.sh:516-519, twin guard at :581

:516 runs if [ "$FIRST" = yes ]; then seed … with no any_other_trusted check. The pending road's last rule (:581) has that check, and so does the rule as the header (:54-57), :576-578 and README:114-116 state it. Take a -0 with an unmarked volume that is not a member (a restored volume, or the failed-bootstrap volume above) and a live, ready dj-1. If every enable through dj-1 fails for all rounds (killed at REPLICATION_ATTEMPT_TIMEOUT, a TLS or admin-port mismatch, or exit 5 on a userRoot the bootstrap never created), dj-0 seeds, writes .bootstrap-complete, and serves unreplicated data behind the same Service as the live topology.

  if [ "$FIRST" = yes ] && ! any_other_trusted; then
    seed "no topology found after $REPLICATION_RETRY_COUNT attempts"
    exit 0
  fi

issue (blocking): A seed that no peer joined has no replication domain, so every restart makes its health wait on join.sh binding with the bootstrap ROOT_PASSWORD. Once the root password is changed, it never turns healthy again.

opendj-packages/opendj-docker/run.sh:144, opendj-packages/opendj-docker/bootstrap/join.sh:463-467, :212-214, :470-476

seed() runs no dsreplication, so config.ldif never gets ds-cfg-replication-domain, and run.sh:144 leaves the marker to join.sh. join.sh's server_up binds as ROOT_USER_DN with the env password. After 150 failures it logs "the server did not come up, giving up" and exits 1, even though the server is up. This hits a StatefulSet at replicas: 1, or an upgraded old-image master that never had a replica, once its root password has been changed over LDAP. README:34-36 says ROOT_PASSWORD is only the initial password, which is why #1092 took the health check off it. At BASE the same restart was healthy right after upgrade -n.

# join.sh, seed()
touch "$SEEDED"
# run.sh
export SEEDED=${SEEDED:-/opt/opendj/data/.replication-seeded}
if ! join_requested || { { grep -q "ds-cfg-replication-domain" ./data/config/config.ldif || [ -f "$SEEDED" ]; } && [ ! -f "$INITIALIZE_PENDING" ]; }; then

issue (non-blocking): seed() ignores a failed publish_state, so a seed whose state never reached its peers still reports itself healthy.

opendj-packages/opendj-docker/bootstrap/join.sh:463-467, :219-224

If both the replace and the --defaultAdd on the local cn=config fail during the pending road's seed (:568-570), pending stays on record. trusted() rejects it, so every other peer skips the seed and exits 1 unhealthy until the seed restarts, while the seed itself is healthy.

seed() { # <why>
  echo "join: $1, seeding it with this server's data"
  publish_state ready || return 1
  rm -f "$INITIALIZE_PENDING"
  touch "$BOOTSTRAP_COMPLETE"
}

At the callers (:508, :517, :562, :569, :582), write seed "…" || exit 1.


issue (non-blocking): With REPLICATION_PEERS set, a fresh first peer seeds while peers still on the previous image answer as absent and hold the data.

opendj-packages/opendj-docker/bootstrap/join.sh:245-249, :581

any_other_trusted is built on trusted(), which rejects absent when the peers are explicit. The final seed rule therefore treats an answering old-image peer as "no peer that could hold the data", which README:113-116 says it is not. Distrusting absent as an initialize source is right, because a new-image peer answers absent while it bootstraps. Only the seed guard needs a separate predicate: "answers and is not pending". The window is narrow: pod-0 on a fresh volume while the others still run the old image.


issue (non-blocking): The deprecated rg path now calls MASTER_SERVER on the local $ADMIN_PORT, so an rg replica whose admin port differs from its master's never turns healthy.

opendj-packages/opendj-docker/bootstrap/replicate.sh:28-29 and the rg branch

At BASE the rg self call on 4444 was unchecked, and the script's exit came from the master call on 4444. Now that call goes to MASTER_SERVER:$ADMIN_PORT, fails 30 times, and run.sh sets BOOTSTRAPPED=false. README:133 says the one-shot types keep their previous behaviour. Either say there that they, like simple, need one ADMIN_PORT on every server, or keep 4444 for the master call.


issue (non-blocking): If /dev/shm is not writable, the || mktemp fallback writes the root password to /tmp/tmp.*, which run.sh's sweep never removes.

opendj-packages/opendj-docker/bootstrap/join.sh:116-118, opendj-packages/opendj-docker/run.sh:45-46

A join killed by a stop never runs its EXIT trap, so each killed start leaves one more password file in the container's writable layer. The comment at join.sh:113-115 says the opposite. setup.sh already uses a swept /tmp template.

PASSWORD_FILE=$(mktemp -p /dev/shm "opendj-join.$ADMIN_PORT.XXXXXX" 2>/dev/null || mktemp "/tmp/opendj-join.$ADMIN_PORT.XXXXXX")

Also add /tmp/opendj-join."$ADMIN_PORT".* to the rm -f in run.sh:45-46.


issue (non-blocking): Moving from MASTER_SERVER=<ip> to a name-based REPLICATION_PEERS makes cleanup_departed deregister the live master.

opendj-packages/opendj-docker/bootstrap/join.sh:382-402

Under MASTER_SERVER=10.0.0.5, the replica's enable registered the master as hostname=10.0.0.5 and listed 10.0.0.5:8989, and names_match never equates an address with a name. At the replica's first start with REPLICATION_PEERS=dj-0,dj-1, the replica deletes the master's cn=Servers entry (the delete replicates everywhere) and removes it from every list. It heals at the master's next start. README:90-92 warns only about the reverse case. A migration note would do: rename the master to its DNS name and restart it first.


issue (non-blocking): A server that the survivors pruned and that is scaled back up on its retained volume can pass is_member on its stale local cn=admin data.

opendj-packages/opendj-docker/bootstrap/join.sh:264-272

If the replicated delete of its cn=Servers entry arrives after join.sh's is_member search, the join logs "already a member" and repairs nothing. run.sh:144 has already marked the server healthy. It then stays unregistered and missing from the survivors' lists until its own next restart, when the enable re-registers it. Not run: the outcome depends on when the delete arrives relative to the search, and needs a docker run of scale-down, survivor restart, and scale-up on the retained volume.


issue (non-blocking): An empty replication_server_lists result still runs one dsconfig per registered peer with an empty domain name.

opendj-packages/opendj-docker/bootstrap/join.sh:436-452

The here-string feeds one empty line, and [ -n "$value" ] && continue does not skip it. After a failed local search, the result is a log line "adding replication server X to the list of " and one failing JVM per peer, swallowed by || true.

  lists=$(replication_server_lists)
  [ -n "$lists" ] || return 0

suggestion (non-blocking): Nothing pins "a member is ready again right after a restart, without waiting for its peers": both members restart together under a 480 s wait.

.github/scripts/docker-test-replication.sh:200-204

A mutant that gates the marker on any one peer answering stays green, at run.sh:144-145 and at join.sh's already-member exit (:486-490). That gate is the OrderedReady deadlock the run.sh:136-143 comment describes.

# a member is ready again right after a restart, without waiting for its peers
docker stop dj-1 >/dev/null
docker restart dj-0 >/dev/null
wait_until 120 "dj-0 is healthy while dj-1 is down" is_healthy dj-0
docker start dj-1 >/dev/null
wait_healthy dj-1

Pin: this kills the peer-gated mutants. run.sh:145 itself is pinned only by a restart after the root password was changed over LDAP.


suggestion (non-blocking): No case makes dsreplication initialize fail, so the round-1 fix (keep the marker, stay unhealthy on a failed initialize) is unpinned.

opendj-packages/opendj-docker/bootstrap/join.sh:551-559

Every initialize in the script succeeds on 10 sample entries. initialize_from "$source" || true; rc=0 runs identically in CI and stays green.

docker run -d --memory="512m" --network $NETWORK --name dj-fi --hostname dj-fi \
  -e ROOT_PASSWORD="$ROOT_PASSWORD" -e ADD_BASE_ENTRY="--addBaseEntry" \
  -e OPENDJ_REPLICATION_TYPE=simple -e REPLICATION_PEERS=dj-0,dj-fi \
  -e REPLICATION_INITIALIZE_TIMEOUT=1 \
  -e REPLICATION_RETRY_COUNT=2 -e REPLICATION_RETRY_INTERVAL=3 "$IMAGE" >/dev/null
wait_until 300 "the initialize of dj-fi is cut off" logs_have dj-fi "initialize from dj-0 exited with 124"
stays_unhealthy dj-fi "its initialize failed"
docker exec dj-fi test -f /opt/opendj/data/.replication-initialize-pending \
  || fail "dj-fi dropped its pending marker after a failed initialize"

Pin: a one-second bound cuts off the dsreplication JVM before it connects. The mutant then removes the marker and turns healthy, so stays_unhealthy and the test -f go red.


suggestion (non-blocking): The dj-sdsr exit-8 pin depends on a race of about 5 s.

.github/scripts/docker-test-replication.sh:289-291

The test detects "Will sleep for a bit" with a 2 s poll and only then disconnects dj-0. BASE polled every 0.2 s. replicate.sh:82-84 sleeps 5 s before its first enable. On a loaded runner the enable can reach dj-0 first, and then both cells go red even though the code is correct. Not measured. A deterministic order: disconnect dj-0 before starting dj-sdsr, wait for "exited with 8, trying again", then reconnect it.

vharseko added a commit to vharseko/OpenDJ that referenced this pull request Sep 30, 2026
…gy's data, and clean every replication list

Review round 1 of OpenIdentityPlatform#1115:

- join.sh: a failed dsreplication initialize is no longer taken for a success (the
  exit status was that of an if without else), and it is bounded by its own
  REPLICATION_INITIALIZE_TIMEOUT, none by default; REPLICATION_ATTEMPT_TIMEOUT bounds
  the enable only, and the bound sits on the dsreplication call, as timeout cannot run
  a shell function.
- join.sh publishes pending/ready in the local entry cn=Docker Join,cn=config; a pending
  server enables and initializes only through a ready peer, the first peer seeds at
  once when every other one is pending and after the retries only while no peer that
  could hold the data answers. Fresh servers started together no longer take each
  other's bootstrap data for the topology's.
- The cleanup prunes departed servers from every replication list (replication server,
  BASE_DN, cn=schema, cn=admin data), and the lists get every listed, registered peer
  they lack; dsconfig inside the read loops reads /dev/null.
- Self-recognition: a name equal to hostname -f or cut from it at a dot, and the
  container's own addresses and their /etc/hosts names, compared whole; is_member
  checks every registered hostname; the generation ID comes from the domain's own
  monitor entry.
- run.sh stays unhealthy on a restart while $INITIALIZE_PENDING is on the volume.
- Dockerfile-alpine installs coreutils, whose timeout signals the whole process group.
- CI: both image jobs run .github/scripts/docker-test-replication.sh, which adds the
  initialize pin, held unhealthy checks, the full-list scale-down check, three fresh
  servers started at once, a master named by its own address, and the replicate.sh
  pins on the /dev/shm glob and the retry on exit 8.
- README documents all of the above.
vharseko added a commit to vharseko/OpenDJ that referenced this pull request Sep 30, 2026
…p, keep a seed ready across restarts, and take turns to enable

Review round 2 of OpenIdentityPlatform#1115:
- both image jobs check out .github/scripts, so the replication test runs at all
- run.sh marks the volume pending before the bootstrap, so a bootstrap that failed or was
  killed is initialized from the topology rather than published ready
- on a restart the volume is ready unless it is pending; a seed that no peer joined has no
  replication domain, and the join cannot bind once the root password was changed
- the first peer seeds after its retries on either road only while no other peer answers
  that may hold the data: ready, or publishing no state but replicating BASE_DN
- the state reaches the peers before the pending marker goes, or the join fails
- a server the survivors removed while it was away takes its replication configuration
  down and enables anew: dsreplication enable does not register a server whose replicated
  cn=admin data already matches the peer's
- servers registered by an address are not removed as departed
- the /tmp fallback of the password files has a name run.sh removes
- an empty list of replication servers changes nothing
- joins through one peer take turns, with a lock entry in the peer's cn=config: two enable
  runs at once leave the loser a member only in part, for good
- README: one ADMIN_PORT for the one-shot types, address-registered servers, the lock
- CI: a member ready while its peer is down, a seed ready after its root password changed,
  a failed bootstrap and a failed initialize, a scale-up on a retained volume, no lock left
  after the parallel start, and sdsr's exit 8 without a race: the replica starts on a network
  of its own rather than dj-0 leaving the topology's, which left dj-0's replication server a
  dead connection of its own directory server to route the replica's initialize into
@vharseko
vharseko force-pushed the feature/1086-replication-join branch from a82819f to df083d9 Compare September 30, 2026 15:40
@vharseko

Copy link
Copy Markdown
Member Author

Round 2 is in df083d9, rebased onto the current master (3295ece). All four blocking issues were real, and so were the non-blocking ones; three are fixed differently from the snippet, and the local runs turned up two more faults, fixed here as well. Point by point:

Blocking

  1. .github/scripts not checked out — fixed with your snippet in both image jobs. The docker jobs of a82819f therefore never ran a scenario; the report I promised on them has nothing to report.
  2. A failed or killed first bootstrap published ready — fixed as you suggested: run.sh touches $INITIALIZE_PENDING before setup.sh runs, and the touch after it is gone. Such a volume is now initialized from the topology on its next start. Pinned by dj-bf, whose bootstrap runs setup.sh and then exits 1: after a restart it must log initializing from dj-0 and serve uid=user.0, which only that initialize brings.
  3. The first peer seeded on the non-pending road while a ready peer answered — fixed: both final seed rules now carry the same guard. That guard is not any_other_trusted any more, see 6.
  4. A seed that no peer joined stayed unready once its root password changed — fixed differently from the snippet. A $SEEDED marker is written by the join, and the join cannot run once the password changed, so it would not cover your second case (an old-image master that never had a replica). Since 2 marks every bootstrapped volume up front, "not pending" now means exactly "this volume holds the data of the topology", so run.sh gates the restart on that alone: if ! join_requested || [ ! -f "$INITIALIZE_PENDING" ]. The replication-domain grep is gone — a seed has none. The cost: an old-image volume whose replica never joined is healthy on restart, which is what the base did. Pinned by dj-solo: it seeds, its root password is changed with ldappasswordmodify, and after a restart it must turn healthy. join.sh now says "did not come up, or does not take ROOT_PASSWORD any more" when it gives up.

Non-blocking

  1. seed() ignored a failed publish_state — fixed as suggested (publish_state ready || return 1 before the marker goes, || exit 1 at every caller), and the same order after a successful initialize: the state is published before the marker is removed, and the join fails otherwise.
  2. Old-image peers answering absent did not hold the seed off — fixed with a narrower predicate than "answers and is not pending". A server of this image also answers absent while it bootstraps (setup has started it, the state entry is not there yet), so that predicate would make the first peer give up during a parallel start whenever a neighbour's bootstrap outlasts its retries, and the pending neighbours would then wait for a seed that never comes. any_other_may_hold_data counts a peer that is ready, or absent and replicating BASE_DN — an old-image member does, a bootstrapping server does not.
  3. rg and a per-server ADMIN_PORT — README, as you offered: the one-shot types, like simple, need one ADMIN_PORT and one REPLICATION_PORT on every server. At the base, srs and sdsr could not work with a non-default ADMIN_PORT at all (their self call went to 4444), and rg silently skipped the group-id of its own domain; going back to 4444 would restore that half-configuration.
  4. The /tmp fallback of the password file — fixed with your template, the same for replicate.sh, and run.sh removes /tmp/opendj-join.* and /tmp/opendj-replicate.*.
  5. MASTER_SERVER=<ip> → names deregisters the live master — fixed in code rather than with a note. A note cannot work: restarted with the new list, the master recognises itself in cn=admin data by its own address and takes the "already a member" road, so it never re-registers by name. cleanup_departed now leaves servers registered by an address alone, in cn=admin data and in the lists; the README says they are removed by hand once really gone.
  6. A pruned server scaled up again on its volume — confirmed, and the fix you implied is not enough. join.sh now asks an answering peer that replicates BASE_DN whether it still registers this server (registered_with_peers). But then dsreplication enable exits 0 and still registers nobody: with both cn=admin data replicated and the registries alike, updateConfiguration takes the "already replicated: nothing to do in terms of ADS" branch (ReplicationCliMain, the areEqual(registry1, registry2) case). So a server that replicates BASE_DN but that a peer no longer registers drops a stale copy of its own entry, runs dsreplication disable --disableAll on itself, and then enables, which registers it. Pinned: after the scale-down dj-2 comes back on the volume it kept, must reappear in cn=admin data of dj-0, and a change made on it must reach dj-0.
  7. Empty replication_server_lists — fixed with your guard.

Suggestions

  1. Restart without the peers — done with your sequence (dj-1 stopped, dj-0 restarted, healthy within 120 s, then dj-1 started).
  2. Failed initialize — done with your dj-fi case (one-second bound, exited with 124, stays_unhealthy, marker kept). It runs right after the reborn-seed case and is removed afterwards; its registration is pruned when the survivors restart with the three-server list.
  3. The dj-sdsr race — fixed, but not by taking dj-0 off the network first. Held off for the whole bootstrap of the replica, dj-0's directory server reconnected through another replication server while dj-0's own replication server kept a handler of that directory server on the connection that died with the interface. When the replica then connected to dj-0's replication server, the initialize of cn=admin data went to that handler (Domain cn=admin data: the server with serverId=… is unreachable every two minutes) and dsreplication enable hung at "Initializing registration information". It depends on which replication server the directory server comes back through, so short runs passed. Now the replica is what is cut off: it starts on a network of its own, where it resolves its own name but not dj-0, and moves to the network of the topology once it has logged exited with 8, trying again.

Found in the local runs

  • Two enables through one peer at once break each other. With dj-1 and dj-2 joining through dj-0 together, the loser failed with ManagedObjectAlreadyExistsException (exit 17, both created the replication server the seed lacked), and every later enable exited 5, "already replicated", without its membership ever completing. An enable now holds a lock on the peer it runs through: the entry cn=Docker Join Lock,cn=config, which only one add creates (a peer that refuses it for another reason than holding it does not keep the join out); a lock left by a killed join is broken once older than REPLICATION_ATTEMPT_TIMEOUT plus a minute, by a delete that asserts the value it read. The parallel start now also checks that no lock is left on dj-p0.
  • An overlay over the published 5.1.2 image hung in dsreplication initialize: the replica listener died on ieCtx being null in processErrorMsg, which master already guards. So the local runs no longer layer the scripts over a published image.

Verification: .github/scripts/docker-test-replication.sh passed in full over Debian and Alpine images built the way the image jobs build them, from a server package of the branch at ff54b08 (Debian 2367 s, Alpine 1948 s). The rebase onto 3295ece brings #1116, #1117 and #1118, which change no Docker file; the docker jobs of this head run the same script on that base.

@maximthomas maximthomas left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

praise: the round-2 fixes land where the bugs were, and the new CI cases pin them.

  • run.sh:177-179 marks the volume pending before setup.sh runs, so a bootstrap killed half-way is no longer healthy on restart (pinned by dj-bf).
  • One predicate, any_other_may_hold_data (join.sh:435), now guards both final seed rules (:663, :730).
  • dj-bf, dj-solo and dj-fi (.github/scripts/docker-test-replication.sh:219-289) run green in both build-docker and build-docker-alpine at df083d9.

issue (blocking): reset_replication takes a member's replication down while the container already reports itself healthy.

opendj-packages/opendj-docker/run.sh:147-148, opendj-packages/opendj-docker/bootstrap/join.sh:626-636, :502-513

On a restart without $INITIALIZE_PENDING, run.sh touches $BOOTSTRAP_COMPLETE before join.sh starts. When the first answering peer does not register this server, join.sh deletes its own cn=admin data entry and runs dsreplication disable --disableAll, and only then tries to enable again; nothing in join.sh removes the marker. The CI's own scale-up case (docker-test-replication.sh:328-335) is this road: dj-2 on its retained volume is healthy, so a Service routes writes to it while it has no BASE_DN domain. Those writes get no CSN and no changelog record and never reach dj-0/dj-1. If every enable then fails, :667 says the container "will not report itself healthy" and exits 1, while it stays healthy and unreplicated until its next restart. The reset runs only after localhost and a peer both took a bind with ROOT_PASSWORD, so the reason for healthy-at-once (a changed root password) does not apply on this road.

if replicates_base_dn localhost && ! registered_with_peers; then
  # no replication domain until the enable below registers this server again: writes taken
  # meanwhile would never replicate, so it is not healthy until joined() says so
  rm -f "$BOOTSTRAP_COMPLETE"
  reset_replication
elif is_member; then

issue (blocking): registered_with_peers treats a failed search of the peer's cn=admin data as "this server is not registered", so a transient failure resets a healthy member.

opendj-packages/opendj-docker/bootstrap/join.sh:260-263, :288-299

registered_hosts pipes search … 2>/dev/null | awk, so a failed ldapsearch gives an empty list. The loop then runs zero times and the function returns 1, the same answer as a peer that really dropped this server. Example: a two-server StatefulSet rolling update. dj-1 restarts on the ready road and is Ready at once, then the controller terminates dj-0. If dj-0 stops answering between replicates_base_dn "$peer" and registered_hosts "$peer", dj-1 deletes its own registration and runs disable --disableAll. It can re-enable only through dj-0, which is down, so every write dj-1 takes meanwhile is never replicated (see the comment above). A peer whose own cn=admin data still lags is a second trigger that this fix does not cover.

registered_hosts() { # [<host>, localhost by default]
  local out
  out=$(search "${1:-localhost}" --baseDN "cn=Servers,cn=admin data" --searchScope one "(objectClass=*)" hostname) || return 1
  printf '%s\n' "$out" | awk 'tolower($1) == "hostname:" { print $2 }'
}

registered_with_peers() {
  local peer host hosts
  for peer in "${PEERS[@]}"; do
    is_self "$peer" && continue
    replicates_base_dn "$peer" || continue
    hosts=$(registered_hosts "$peer") || continue
    for host in $hosts; do
      is_self "$host" && return 0
    done
    return 1
  done
  return 0
}

issue (blocking): peer_state reports a survivor that refuses the ROOT_PASSWORD bind (exit 49) as down, so a first peer that lost its volume seeds next to live survivors.

opendj-packages/opendj-docker/bootstrap/join.sh:236-249, :435-445, :730

Every exit other than 0 and 32 maps to down, including 49 from a survivor whose root password was changed over LDAP. That change lives in cn=config, is not replicated, and is supported (README: ROOT_PASSWORD is only the initial password). any_other_may_hold_data has only ready and absent arms. So when PEERS[0] loses its volume after such a change, every survivor reads as down, and after REPLICATION_RETRY_COUNT rounds :730 lets seed() publish ready. -0 then turns healthy with bootstrap data only, next to survivors that hold the real data: the fork this rule exists to prevent. The README names only "every other server down" as the residual risk. Only an initialized server can have a password other than the bootstrap's, so a peer that answers and refuses the bind may hold the data.

peer_state() {
  local out state
  out=$(search "$1" --baseDN "$STATE_DN" --searchScope base "(objectClass=*)" description)
  case $? in
    0)
      state=$(printf '%s\n' "$out" | awk 'tolower($1) == "description:" { print $2; exit }')
      case $state in
        ready | pending) echo "$state" ;;
        *) echo absent ;;
      esac
      ;;
    32) echo absent ;;
    # it answers but no longer takes ROOT_PASSWORD: a server past its bootstrap
    49) echo refused ;;
    *) echo down ;;
  esac
}

any_other_may_hold_data() {
  local peer
  for peer in "${PEERS[@]}"; do
    is_self "$peer" && continue
    case $(peer_state "$peer") in
      ready | refused) return 0 ;;
      absent) replicates_base_dn "$peer" && return 0 ;;
    esac
  done
  return 1
}

issue (non-blocking): the pending road still decides membership with is_member, which reads only the local copy of cn=admin data, while the ready road now asks the peers.

opendj-packages/opendj-docker/bootstrap/join.sh:676, :694

Take a pending volume that once enabled (its initialize never completed) and that the survivors pruned while it was away: it still lists itself in its own cn=admin data. :676 skips the enable, :694 initializes from a trusted source, and :703-706 publish ready and mark the container healthy. If the survivors' delete has not reached it yet, it turns healthy while registered only in its own registry, then drops out once the delete arrives. If the delete has already replayed, every enable "registers nobody" (the premise at :492-498), and this road has no reset, so the container stays unhealthy on every start. Not run: this needs a scale-down and a scale-up of a pending volume. That the initialize succeeds against a server no peer registers comes from reading ReplicationCliMain. Suggested change: decide membership here with member_of_topology as well, and reset a local domain the peers do not register before enabling, as the ready road does.


issue (non-blocking): the join lock serializes only enables through the same peer. A ready server that is not yet a member can enable or reset while a pending joiner enables through it.

opendj-packages/opendj-docker/bootstrap/join.sh:369, :629, :502-513

take_join_lock locks only host1. The ready road publishes ready (:629) before the server is a member and before reset_replication, which takes no lock, and the pending road trusts a peer on peer_state alone (:680). Two cases hit this: a seed restarted before its replicas joined, or a returning volume in reset. Either can enable through Y while a pending J enables through it, so both run the configuration of X at once, the overlap the lock was added for. The replication-server collision alone completes on retry. Whether the concurrent registration and initialize can leave a member registered only in part was not traced or run (that needs three containers timed against each other). Suggested change: also take the lock on this server, in enable_through next to host1's and around reset_replication. A non-blocking take_join_lock cannot deadlock.


suggestion (non-blocking): the absent) replicates_base_dn arm of any_other_may_hold_data is not pinned, so dropping it leaves the suite green.

opendj-packages/opendj-docker/bootstrap/join.sh:441, .github/scripts/docker-test-replication.sh

Only dj-0's initial seed reaches the FIRST-peer give-up rules (:663, :730), and it does so while dj-1 is not started (down). dj-sdsr, the one server that replicates without publishing a state, is in no REPLICATION_PEERS list. If a later edit drops the arm, round 2's "fresh first peer seeds next to an old-image member" comes back and CI stays green.

# once dj-sdsr has joined dj-0 and sits on the topology's network
start_node dj-f dj-f,dj-sdsr -e REPLICATION_RETRY_COUNT=2 -e REPLICATION_RETRY_INTERVAL=5
wait_until 300 "dj-f gives up" logs_have dj-f "could not join the replication topology after 2 attempts"
! logs_have dj-f "seeding it with this server's data" || fail "dj-f seeded next to a member that publishes no state"
stays_unhealthy dj-f "a peer that may hold the data answers"

Pin: with absent) replicates_base_dn "$peer" && return 0 ;; dropped, dj-f seeds and this case goes red.


suggestion (non-blocking): the is_address "$host" && continue skips in cleanup_departed are not pinned: the only address case runs with MASTER_SERVER alone, so cleanup_departed returns before it reaches them.

opendj-packages/opendj-docker/bootstrap/join.sh:527, :530, :539, .github/scripts/docker-test-replication.sh:377-388

dj-ipm and dj-ipr start with -e MASTER_SERVER=$MASTER_ADDRESS and no REPLICATION_PEERS, so :527 returns before either skip runs. Dropping either skip throws a live address-registered master out of cn=admin data and out of the replication server lists, and the suite stays green.

# dj-ipm and dj-ipr with volumes; after the existing dj-ipr checks, recreate dj-ipr on its volume
# with -e REPLICATION_PEERS=dj-ipm,dj-ipr and no MASTER_SERVER, then
wait_healthy dj-ipr
! logs_have dj-ipr "removing departed server $MASTER_ADDRESS" || fail "dj-ipr removed the master registered by its address"

Pin: dropping the skip at :530 makes dj-ipr log exactly that line.


suggestion (non-blocking): the join lock's mutual exclusion is not pinned. The only check is that no lock entry is left on dj-p0, and a take_join_lock that never adds the entry passes it too.

.github/scripts/docker-test-replication.sh:417-420, opendj-packages/opendj-docker/bootstrap/join.sh:329-357

Nothing checks for "is enabling replication through … waiting for it" or "breaking the lock", and whether the parallel race fires at all depends on timing. A mutant whose take_join_lock returns 0 without adding the entry brings back the ManagedObjectAlreadyExists / exit 5 failure the lock was added for, and CI stays green. On exit 75 the loop moves on to the next peer (:637-650), so a lock on dj-0 alone is not enough.

# before the scale-up, add the entry add_join_lock writes, on dj-0 and on dj-1,
# with "description: dj-held $(date +%s)"; then
wait_until 120 "dj-2 waits for dj-0" logs_have dj-2 "dj-held is enabling replication through dj-0, waiting for it"
wait_until 120 "dj-2 waits for dj-1" logs_have dj-2 "dj-held is enabling replication through dj-1, waiting for it"
sleep 6 # two retry intervals: start_node passes REPLICATION_RETRY_INTERVAL=3
admin_data_lacks dj-0 dj-2 || fail "dj-2 enabled while both peers were locked"
# delete both entries and keep the existing "dj-2 registers in the topology again" wait

Pin: a take_join_lock that ignores the entry never logs the wait lines.


suggestion (non-blocking): no case pins that a healthy member's plain restart does not reset its replication. The restart cases check only health and data flow, and a reset followed by a re-enable passes both.

.github/scripts/docker-test-replication.sh:205-214

Nowhere does the suite look for "the peers no longer register this server" (join.sh:504). A mutant whose registered_with_peers returns 1 for a registered member resets dj-1 on its restart and passes. Once the failed-search fix above is in, this is the one observable that tells a reset from a plain rejoin.

# after: wait_has dj-0 "ou=replicated2,dc=example,dc=com"
for n in dj-0 dj-1; do
  ! logs_have $n "the peers no longer register this server" || fail "$n reset its replication on a plain restart"
done

nitpick (non-blocking): on the ready road join.sh still logs that the container "will not report itself healthy", but run.sh has already marked it healthy there.

opendj-packages/opendj-docker/bootstrap/join.sh:667, :61-64

The [ ! -f "$INITIALIZE_PENDING" ] branch is reached only after run.sh:147-148 touched $BOOTSTRAP_COMPLETE. On this road an operator sees that line on a container the health check reports healthy, serving but not replicating. The README describes the behaviour correctly. With the first blocking fix above, the line is true again after a reset, but still false on a plain give-up.

echo "join: could not join the replication topology after $REPLICATION_RETRY_COUNT attempts, this server keeps serving its data unreplicated until a later start joins it"

Or narrow the header comment at :61-64 to the pending road.


note (non-blocking): documentation that differs from the code, with no stated reason:

  • opendj-packages/opendj-docker/README.md:124: the README breaks a stale join lock once it is "older than REPLICATION_ATTEMPT_TIMEOUT plus a minute". That holds only for a positive timeout. With 0 or a non-numeric value the TTL is a fixed 660 s (join.sh:314-318), and bounded() (:211-219) runs the enable with no bound, contrary to join.sh:83-84 "no enable runs without a bound".

note (non-blocking): Not checked:

  • Whether two enables through different peers can each read-modify-write a third server's replication-server list and lose one value. This is the untraced half of the join-lock comment above, and it does not change this review's action.

vharseko added a commit to vharseko/OpenDJ that referenced this pull request Oct 1, 2026
…gy's data, and clean every replication list

Review round 1 of OpenIdentityPlatform#1115:

- join.sh: a failed dsreplication initialize is no longer taken for a success (the
  exit status was that of an if without else), and it is bounded by its own
  REPLICATION_INITIALIZE_TIMEOUT, none by default; REPLICATION_ATTEMPT_TIMEOUT bounds
  the enable only, and the bound sits on the dsreplication call, as timeout cannot run
  a shell function.
- join.sh publishes pending/ready in the local entry cn=Docker Join,cn=config; a pending
  server enables and initializes only through a ready peer, the first peer seeds at
  once when every other one is pending and after the retries only while no peer that
  could hold the data answers. Fresh servers started together no longer take each
  other's bootstrap data for the topology's.
- The cleanup prunes departed servers from every replication list (replication server,
  BASE_DN, cn=schema, cn=admin data), and the lists get every listed, registered peer
  they lack; dsconfig inside the read loops reads /dev/null.
- Self-recognition: a name equal to hostname -f or cut from it at a dot, and the
  container's own addresses and their /etc/hosts names, compared whole; is_member
  checks every registered hostname; the generation ID comes from the domain's own
  monitor entry.
- run.sh stays unhealthy on a restart while $INITIALIZE_PENDING is on the volume.
- Dockerfile-alpine installs coreutils, whose timeout signals the whole process group.
- CI: both image jobs run .github/scripts/docker-test-replication.sh, which adds the
  initialize pin, held unhealthy checks, the full-list scale-down check, three fresh
  servers started at once, a master named by its own address, and the replicate.sh
  pins on the /dev/shm glob and the retry on exit 8.
- README documents all of the above.
vharseko added a commit to vharseko/OpenDJ that referenced this pull request Oct 1, 2026
…p, keep a seed ready across restarts, and take turns to enable

Review round 2 of OpenIdentityPlatform#1115:
- both image jobs check out .github/scripts, so the replication test runs at all
- run.sh marks the volume pending before the bootstrap, so a bootstrap that failed or was
  killed is initialized from the topology rather than published ready
- on a restart the volume is ready unless it is pending; a seed that no peer joined has no
  replication domain, and the join cannot bind once the root password was changed
- the first peer seeds after its retries on either road only while no other peer answers
  that may hold the data: ready, or publishing no state but replicating BASE_DN
- the state reaches the peers before the pending marker goes, or the join fails
- a server the survivors removed while it was away takes its replication configuration
  down and enables anew: dsreplication enable does not register a server whose replicated
  cn=admin data already matches the peer's
- servers registered by an address are not removed as departed
- the /tmp fallback of the password files has a name run.sh removes
- an empty list of replication servers changes nothing
- joins through one peer take turns, with a lock entry in the peer's cn=config: two enable
  runs at once leave the loser a member only in part, for good
- README: one ADMIN_PORT for the one-shot types, address-registered servers, the lock
- CI: a member ready while its peer is down, a seed ready after its root password changed,
  a failed bootstrap and a failed initialize, a scale-up on a retained volume, no lock left
  after the parallel start, and sdsr's exit 8 without a race: the replica starts on a network
  of its own rather than dj-0 leaving the topology's, which left dj-0's replication server a
  dead connection of its own directory server to route the replica's initialize into
@vharseko
vharseko force-pushed the feature/1086-replication-join branch from df083d9 to aca6ebe Compare October 1, 2026 13:37
@vharseko

vharseko commented Oct 1, 2026

Copy link
Copy Markdown
Member Author

Round 3 is in aca6ebe, rebased onto the current master (df0af76). All three blocking issues were real, and so were the non-blocking ones; two fixes go further than the snippet. Point by point:

Blocking

  1. reset_replication on a healthy container — fixed, and for restarts as well. Your snippet covers the start in which the reset runs. Once dsreplication disable has run, though, the volume has no BASE_DN domain, $INITIALIZE_PENDING is not there, and a restart takes run.sh's ready road again: healthy at once, enabling in the background, with the same unreplicated writes. So the reset now marks the volume first: reset_replication touches $REJOIN_PENDING (/opt/opendj/data/.replication-rejoin-pending) and removes .bootstrap-complete before anything changes. run.sh marks a restart healthy only if neither marker is there, and joined() and seed() remove it. The OrderedReady argument does not apply here: the reset needs a peer that answered, and a StatefulSet never prunes -0. Pinned in the scale-up case (see 7): with both survivors' locks held, dj-2 must have no .bootstrap-complete once it logged the reset. After docker restart dj-2 it must log was taken out of its replication topology, wait for dj-0 again, and still have none.
  2. A failed search read as "not registered" — fixed with your registered_hosts, and registered_with_peers now asks every peer instead of the first one. It concludes "not registered" only when at least one answering peer that replicates BASE_DN listed its registrations and none of them lists this server. This also covers the second trigger you named, a peer whose cn=admin data lags, provided another peer lists the server. The cost is on the other side. A pruned server whose survivors' delete has not reached every survivor yet counts as registered. That is the round-2 window: it stays unregistered until its next start. Of the two, that does less harm than resetting a member.
  3. Exit 49 read as down — fixed with your snippet. The header comment and the README name the refusing peer next to the ready one and the stateless member.

Non-blocking

  1. Pending road membership — fixed as suggested. Every round starts with the same check as the ready road: reset a local domain that the peers no longer register. Membership after an enable is member_of_topology.

  2. The join lock — enable_through now takes the lock on this server before the peer's, and releases it if the peer's is taken. reset_replication holds its own lock while it drops its entry and disables, and waits for the lock if a joiner holds it. Two additions:

    • Neither lock is waited for, so there is no deadlock, as you say. But two servers that each hold their own lock and ask for the other's give up together, and with one REPLICATION_RETRY_INTERVAL they also try again together, round after round. The pause between rounds is now the interval plus a random part of as much again (pause(), README).
    • A lock in a server's own cn=config that a killed join left behind would otherwise keep the next start out for the whole TTL. A lock that holds this server's own name is now broken at once, because a start runs one join.

    On the unchecked half (two enables through different peers and the list of a third server): ReplicationCliMain writes the whole list (setReplicationServer(set) then commit(), at :6229, :6279 and :6373), so a lost value is possible, and the locks on the two ends do not serialize it. repair_replication_servers adds a registered peer back on the third server's next start. I am recording this, not fixing it in this round.

Suggestions

  1. The absent arm — pinned with your dj-f case, placed after dj-sdsr has joined.
  2. The join lock's exclusion — pinned in the scale-up of dj-2, together with 1. The locks of dj-0 and dj-1 are held by hand before dj-2 starts. dj-2 must log the wait for both of them, and after two rounds dj-0 must still not register it. Then the restart described in 1, and the locks are dropped. dj-2 runs with REPLICATION_ATTEMPT_TIMEOUT=600, so it does not break the held locks as stale. It waits 12 s rather than 6 s, because a round can now take up to twice the interval.
  3. The is_address skips — pinned. dj-ipr keeps a volume and is recreated on it with REPLICATION_PEERS=dj-ipm,dj-ipr and no MASTER_SERVER. Once it logs that it is a member, neither removing departed server $MASTER_ADDRESS nor removing departed replication server $MASTER_ADDRESS: may be in its log, and the master must still be in its cn=admin data.
  4. No reset on a plain restart — pinned with your loop after the restart of both members.

Nitpick and notes

  1. The give-up message — give_up() now says what holds: a container that already reports itself healthy keeps serving its data unreplicated until a later start joins it, and one that does not stays unhealthy.
  2. REPLICATION_ATTEMPT_TIMEOUT 0 or not a number — falls back to 120 with a log line. So "no enable runs without a bound" stays true, the lock TTL is always that bound plus a minute, and the README table says so. REPLICATION_RETRY_INTERVAL is checked the same way, because pause() computes with it.

Verification: .github/scripts/docker-test-replication.sh passed in full over the Debian image, with the scripts of this head laid over an image built from a server package of the branch (ff54b08), in 2628 s. That includes every new case: the plain restart, dj-2 behind the held locks with its restart, dj-f and the moved dj-ipr. Three earlier runs were cut short by the load and the memory of the host, at steps this round does not touch. Alpine was not run locally this time, and the mutants named in the pins were not run either; the docker jobs of this head run the script on both images over the current master.

@vharseko
vharseko requested a review from maximthomas October 1, 2026 13:38

@maximthomas maximthomas left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

praise: The new commit fixes the round-3 failure roads where they start.

  • registered_hosts now fails when its search fails (|| return 1, join.sh:279-280). registered_with_peers skips that peer and returns as soon as any answering peer lists this server (:316-319), so a transient failure no longer resets a healthy member.
  • peer_state maps exit 49 to refused (join.sh:263) and any_other_may_hold_data counts it (:486). A first peer no longer seeds next to a survivor whose root password was changed.
  • The scale-up-again case now holds both survivors' join locks by hand and checks that dj-2 stays unready across a restart (.github/scripts/docker-test-replication.sh:355-371).

issue (blocking): Removing the health marker does not make the container unhealthy before reset_replication takes its replication down, so the reset and usually the whole re-enable run while it still reports itself healthy.

opendj-packages/opendj-docker/bootstrap/join.sh:548-575, :704-712, opendj-packages/opendj-docker/run.sh:154-155, opendj-packages/opendj-docker/Dockerfile:93

On the ready road, run.sh touches $BOOTSTRAP_COMPLETE before the server starts. One probe that passes before join.sh:556 makes the container healthy. The probes during the start period come every few seconds, while server_up, publish_state and the per-peer searches take several seconds. reset_replication then removes the marker and, within seconds, unregisters the server and runs dsreplication disable --disableAll. But --interval=30s --retries=3 turns the status unhealthy only after three failed probes in a row, which is 60-90 s after the marker is removed. A readinessProbe on healthcheck.sh lags by failureThreshold × period in the same way, and LDAP connections that are already open are not affected by either. A typical re-enable finishes inside that window, so the status never leaves healthy. Writes in that window get no change number and never reach the peers, which is round 3's divergence, now limited to the length of the lag. The comment at :548-550 ("stops reporting itself healthy before anything changes") promises more than a marker can deliver. The CI check at docker-test-replication.sh:361/:371 tests the marker file, not the health status, so it passes either way. Not run: it needs a pruned volume restarted on the ready road while a script watches the health status.

# reset_replication, right after own=$JOIN_LOCK_VALUE (join.sh:565)
  dsconfig_local set-backend-prop --backend-name userRoot --set writability-mode:internal-only || {
    echo "join: could not stop client writes, its replication configuration stays as it is"
    release_join_lock localhost "$own"
    return 1
  }
# joined(), before rm -f "$REJOIN_PENDING" (join.sh:661)
  [ ! -f "$REJOIN_PENDING" ] || dsconfig_local set-backend-prop --backend-name userRoot \
    --set writability-mode:enabled || echo "join: could not let client writes in again"

Change only the data backend (setup.sh creates it as userRoot): a global writability-mode:internal-only would also reject the writes dsreplication enable makes to cn=admin data and cn=config. With internal-only, replication can still write to the backend. The setting persists across a restart, and $REJOIN_PENDING already covers that road. Then update the comment at :548-550 to say what the marker does: new probes fail, but open connections and the probe's retries are not affected.


issue (non-blocking): A knob written with a leading zero passes the new decimal checks and then breaks the octal $(( )).

opendj-packages/opendj-docker/bootstrap/join.sh:91, :95, :342, :421

[ "$REPLICATION_RETRY_INTERVAL" -ge 0 ] accepts 08, but pause()'s $((REPLICATION_RETRY_INTERVAL + ...)) fails with "value too great for base". Outside posix mode, bash then abandons the whole top-level command it is in. On the ready road that is the if [ ! -f "$INITIALIZE_PENDING" ] block (:701-742), so a volume that holds the data falls through to :750 and runs publish_state pending. A first peer may then count it as pending and seed beside it. A probe on bash 3.2.57 printed the error, then "REACHED pending road". REPLICATION_ATTEMPT_TIMEOUT=08 passes the check at :91 and leaves JOIN_LOCK_TTL empty at :342, after which :372 never breaks a stale lock. REPLICATION_RETRY_COUNT is not checked at all.

seconds() { case $1 in ''|*[!0-9]*) return 1 ;; esac; echo $((10#$1)); }
REPLICATION_RETRY_INTERVAL=$(seconds "$REPLICATION_RETRY_INTERVAL") || {
  echo "join: REPLICATION_RETRY_INTERVAL is not a number of seconds, using 10"
  REPLICATION_RETRY_INTERVAL=10
}

Do the same for REPLICATION_ATTEMPT_TIMEOUT and REPLICATION_RETRY_COUNT, both of which must also be greater than 0.


issue (non-blocking): On the ready road, a reset_replication that could not take its own lock is ignored and never retried during that start.

opendj-packages/opendj-docker/bootstrap/join.sh:705-712, :557-562

The reset check runs once, before the round loop. When the lock loop gives up, the function has already touched $REJOIN_PENDING and removed the marker, and :706 ignores its return 1. Every remaining round then enables against the stale domain, which by :538-547 registers nobody. The server ends up unhealthy until its next start. The pending road re-checks in every round (:753-755). This affects liveness only, since nothing was changed. Moving the check into the round loop, as on the pending road, retries the reset in the next round.


issue (non-blocking): Once reset_replication releases its own lock (:574), the server still publishes ready (:704), so a pending joiner can enable through it while it has no replication domain.

opendj-packages/opendj-docker/bootstrap/join.sh:574, :704, :763

The pending road needs only trusted ready before it calls enable_through. The reset holds this server's lock only while it unregisters and disables, and the server holds no lock between its rounds or during pause(). A pending joiner Z can then run dsreplication enable --host1 R --host2 Z against R. Depending on what R's local registry still lists, Z either joins the survivors' mesh or a two-server registry forms next to theirs until R's own enable merges it. Not traced: what dsreplication disable --disableAll leaves in R's cn=admin data. Not run: it needs three containers timed so that the joiner arrives after the reset. One fix is to publish a state such as rejoining from the reset until joined(), which trusted() refuses and any_other_may_hold_data treats like ready.


suggestion (non-blocking): No test pins the lock a server takes on itself. Deleting take_join_lock localhost in enable_through (:398), the own-lock loop of reset_replication (:557-564), or the own-name break at :372 stays green.

.github/scripts/docker-test-replication.sh:355-356, opendj-packages/opendj-docker/bootstrap/join.sh:372, :398, :557-564

The scale-up-again case holds only dj-0's and dj-1's locks. No case holds dj-2's own lock or enables through a server that is itself enabling or resetting.

# before `docker rm -f dj-2` (docker-test-replication.sh:330): the entry stays in dj-2's cn=config on its volume
hold_lock dj-2
# after `start_node dj-2 dj-0,dj-1,dj-2 ...` (:357)
wait_until 120 "dj-2 waits for its own lock" logs_have dj-2 "dj-held is enabling replication through this server, waiting for it"
! logs_have dj-2 "dj-held is enabling replication through dj-0" || fail "dj-2 enabled before its own lock was free"
drop_lock dj-2

Pin: this kills the :557-564 mutant. The :398 and :372 mutants need a joiner that enables through a server whose own join holds its lock.


suggestion (non-blocking): No CI case exercises the fixes for round 3's refused-peer and failed-search roads. Mapping 49 to down (join.sh:263), dropping | refused (:486), or restoring the empty-list fallthrough of registered_hosts all stay green.

.github/scripts/docker-test-replication.sh:238-247, :398

In no case does a peer answer 49 to another server's peer_state. The only root-password change is on dj-solo, whose list contains only itself, and the dj-f case runs next to dj-sdsr, which is an absent peer.

# between docker-test-replication.sh:245 and :246 (dj-solo is up with its root password changed)
start_node dj-f dj-f,dj-solo -e REPLICATION_RETRY_COUNT=2 -e REPLICATION_RETRY_INTERVAL=5
wait_until 300 "dj-f gives up beside a peer that refuses the bind" logs_have dj-f "could not join the replication topology after 2 attempts"
if logs_have dj-f "seeding it with this server's data"; then fail "dj-f seeded next to a peer past its bootstrap"; fi
docker rm -f dj-f >/dev/null; docker volume rm vol-dj-f >/dev/null

Pin: this case fails when 49 is mapped to down. The failed-search road needs a peer whose cn=admin data search fails after replicates_base_dn has answered.

vharseko added a commit to vharseko/OpenDJ that referenced this pull request Oct 3, 2026
…gy's data, and clean every replication list

Review round 1 of OpenIdentityPlatform#1115:

- join.sh: a failed dsreplication initialize is no longer taken for a success (the
  exit status was that of an if without else), and it is bounded by its own
  REPLICATION_INITIALIZE_TIMEOUT, none by default; REPLICATION_ATTEMPT_TIMEOUT bounds
  the enable only, and the bound sits on the dsreplication call, as timeout cannot run
  a shell function.
- join.sh publishes pending/ready in the local entry cn=Docker Join,cn=config; a pending
  server enables and initializes only through a ready peer, the first peer seeds at
  once when every other one is pending and after the retries only while no peer that
  could hold the data answers. Fresh servers started together no longer take each
  other's bootstrap data for the topology's.
- The cleanup prunes departed servers from every replication list (replication server,
  BASE_DN, cn=schema, cn=admin data), and the lists get every listed, registered peer
  they lack; dsconfig inside the read loops reads /dev/null.
- Self-recognition: a name equal to hostname -f or cut from it at a dot, and the
  container's own addresses and their /etc/hosts names, compared whole; is_member
  checks every registered hostname; the generation ID comes from the domain's own
  monitor entry.
- run.sh stays unhealthy on a restart while $INITIALIZE_PENDING is on the volume.
- Dockerfile-alpine installs coreutils, whose timeout signals the whole process group.
- CI: both image jobs run .github/scripts/docker-test-replication.sh, which adds the
  initialize pin, held unhealthy checks, the full-list scale-down check, three fresh
  servers started at once, a master named by its own address, and the replicate.sh
  pins on the /dev/shm glob and the retry on exit 8.
- README documents all of the above.
vharseko added a commit to vharseko/OpenDJ that referenced this pull request Oct 3, 2026
…p, keep a seed ready across restarts, and take turns to enable

Review round 2 of OpenIdentityPlatform#1115:
- both image jobs check out .github/scripts, so the replication test runs at all
- run.sh marks the volume pending before the bootstrap, so a bootstrap that failed or was
  killed is initialized from the topology rather than published ready
- on a restart the volume is ready unless it is pending; a seed that no peer joined has no
  replication domain, and the join cannot bind once the root password was changed
- the first peer seeds after its retries on either road only while no other peer answers
  that may hold the data: ready, or publishing no state but replicating BASE_DN
- the state reaches the peers before the pending marker goes, or the join fails
- a server the survivors removed while it was away takes its replication configuration
  down and enables anew: dsreplication enable does not register a server whose replicated
  cn=admin data already matches the peer's
- servers registered by an address are not removed as departed
- the /tmp fallback of the password files has a name run.sh removes
- an empty list of replication servers changes nothing
- joins through one peer take turns, with a lock entry in the peer's cn=config: two enable
  runs at once leave the loser a member only in part, for good
- README: one ADMIN_PORT for the one-shot types, address-registered servers, the lock
- CI: a member ready while its peer is down, a seed ready after its root password changed,
  a failed bootstrap and a failed initialize, a scale-up on a retained volume, no lock left
  after the parallel start, and sdsr's exit 8 without a race: the replica starts on a network
  of its own rather than dj-0 leaving the topology's, which left dj-0's replication server a
  dead connection of its own directory server to route the replica's initialize into
@vharseko
vharseko force-pushed the feature/1086-replication-join branch from aca6ebe to 9ecb440 Compare October 3, 2026 05:55
@vharseko

vharseko commented Oct 3, 2026

Copy link
Copy Markdown
Member Author

Round 4 is in 9ecb440, rebased onto the current master (4870ff3). Every point held up. Two of the fixes go further than the snippet, and one pin is placed differently. Point by point:

Blocking

  1. The marker does not make the container unhealthy in time. Fixed with your approach. Before anything changes, reset_replication sets the backend of BASE_DN to writability-mode:internal-only. LocalBackendWorkflowElement.checkIfWritable lets internal and synchronization operations through. dsreplication writes to the data only through tasks, and a total update imports past that check. So the enable and the initialize still work, and a client's write gets 53. Three additions to the snippet:

    • seed() lets the writes in again as well. A first peer can seed after a reset, on the ready road once every other peer is pending, and on the pending road too. With only joined() restoring the mode, the backend would have stayed read-only for good.
    • The backend is looked up by ds-cfg-base-dn rather than named userRoot, because a custom BOOTSTRAP may call it something else.
    • The mode goes back to enabled only where the join changed it. hold_writes writes the backend's name into $REJOIN_PENDING before it changes the mode, and leaves alone a backend that is not enabled. That way an operator's read-only backend stays read-only, and a container killed between the two steps still gets its writes back once it rejoins.

    $REJOIN_PENDING goes only after the writes are let in again and the server published ready. If either step fails, the next start tries again and stays unready meanwhile. The comment at reset_replication now says what the marker does and what it does not do. The README says the same, and adds that a join that gives up leaves the server unhealthy and read-only. The pins check behaviour instead of the file: a client's add to dj-2 must fail with 53 once its replication is down, and again after its restart. After it rejoins, the existing add_ou dj-2 rejoined must go through.

Non-blocking

  1. A leading zero in the knobs. Fixed. I confirmed it on bash 3.2.57 and 5.2.15: test reads 08 as 8, $((08 + …)) fails, and bash drops the whole top-level if. whole_number() now reads all four knobs as decimal (10#). REPLICATION_RETRY_COUNT and REPLICATION_ATTEMPT_TIMEOUT must be at least 1, the interval and REPLICATION_INITIALIZE_TIMEOUT at least 0. Anything else falls back to the default with a log line, and the README table says so.
  2. A reset that could not take its lock is not tried again. Fixed, slightly differently. The check now starts every round on both roads (reset_if_unregistered). reset_replication makes one attempt at its own lock and returns 1 if the lock is taken. A round whose reset is due but did not run enables nothing, because an enable before the reset registers nobody. With the lock loop inside the reset kept, that loop would have nested inside the round loop: up to REPLICATION_RETRY_COUNT² rounds.
  3. ready while the domain is gone. Fixed with the rejoining state. The gap was wider than the reset itself: on a restart with $REJOIN_PENDING, the ready road published ready straight away. That start now publishes rejoining. peer_state knows the state; before this it would have turned into absent, which counts as trusted when the peers are not explicit. trusted() refuses it, any_other_may_hold_data counts it like ready, and all_others_pending does not treat it as pending. The server publishes ready again when it leaves the rejoin. On the pending road the reset publishes pending instead: that volume holds no data of the topology, and rejoining would keep a first peer from seeding.

Suggestions

  1. The server's own lock. Pinned, but not with the snippet's check, which does not kill the :557-564 mutant. With the reset's own lock removed, enable_through still stops at the same held lock and logs the same "…through this server, waiting for it". The negative check on "through dj-0" stays green as well. What the scale-up case checks now:

    • dj-2 leaves a lock held by hand in its own cn=config (hold_lock dj-2 before docker rm -f dj-2). After it starts, dj-2 must log the wait for it, publish rejoining, have no marker, and still hold its BASE_DN domain two rounds later. This kills the :557-564 mutant. Then the lock is dropped and the existing checks follow.
    • After the restart, dj-2's own lock is held by hand again, and then both survivors' locks are dropped. dj-2 must log the wait for its own lock, and two rounds later dj-0 must still not register it. This kills the :398 mutant: without its own lock, dj-2 would enable through the free dj-0.
    • The held lock is then relabelled under dj-2's own name (hostname -f, read from the container). dj-2 must log breaking the lock <name> took on localhost and turn healthy. This kills the :372 mutant: without the own-name break the lock stays below its TTL, and dj-2 waits.

    dj-2 now runs with REPLICATION_ATTEMPT_TIMEOUT=1800, because the first lock has been held since the scale-down and must not be broken as stale.

  2. The refusing peer. Pinned with your dj-f case beside dj-solo, with wait_until at 480 s, because dj-f bootstraps first. The failed-search road stays unpinned. As you say, it needs a peer whose cn=admin data search fails after replicates_base_dn has answered, and I am only recording that here.

One more fix in the same case: the check that the restarted dj-2 waits for dj-0 again counted waits from before docker restart. The join of the previous start could still log one while it was being stopped, so the wait could be satisfied before the new server was up. Now only lines logged after the restart count (logs_have_since). With a check that is only as strong as the marker, that race never showed.

Verification: .github/scripts/docker-test-replication.sh passed in full over the Debian image in 2411 s, including every new case. As before, the scripts of this head were laid over an image built from a server package of the branch. Alpine could not run locally, because the host ran out of memory. The mutants named in the pins were not run separately. The docker jobs of this head run the script on both images over the current master.

@vharseko
vharseko requested a review from maximthomas October 3, 2026 05:56

@maximthomas maximthomas left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

praise: 9ecb440 closes round 4's blocking finding where the bug is, and it fixes the knob parsing for every knob at once.

  • hold_writes (join.sh:620-634) fences only the backend of BASE_DN with writability-mode:internal-only. It records the backend in $REJOIN_PENDING before the change, so a kill in between still releases it. checkIfWritable lets the internal and synchronization writes of the rejoin through.
  • whole_number (join.sh:92-117) reads every knob with 10# before any $(( )), so 08 can no longer abandon a top-level command.
  • trusted() refuses rejoining right before the pending road's enable_through (join.sh:284-290, :860-861).

issue (blocking): On the ready road, enable_through runs against a peer that publishes rejoining. Two reset servers can enable with each other and form a lasting two-server topology beside the survivors, and both report healthy and writable.

opendj-packages/opendj-docker/bootstrap/join.sh:809-812, :341-343, :327-339, :704-712

dsreplication disable --disableAll empties the server's own registry: removeAdminData subtree-deletes cn=Servers (ReplicationCliMain.java:5504-5511, ADSContext.java:951-970). An enable between two servers whose registries hold one server or none registers just those two (ReplicationCliMain.java:4629-4648).

Take survivor S, which pruned R and Y. Both come back on their volumes and reset, so both publish rejoining. R's enable through S exits 75 because a third joiner holds S's lock, and the loop moves on to Y (:820). Y is pausing with its own lock free, so R and Y enable with each other and form registry {R,Y}. member_of_topology passes on Y's listing (:334-335), and joined() makes R writable and healthy. Y's next round finds itself registered through R and joins too.

repair_replication_servers adds only hosts from the local registry (:704-712), so S is never added, and every later start passes the same check. The pair serves stale data whose writes never reach S until S itself restarts. This breaks the guarantee README.md:160 and the comment at :572-574 give for rejoining.

Not run: it needs two pruned volumes, a survivor and a contender for the survivor's lock. The mechanism above is traced end to end.

      for peer in "${PEERS[@]}"; do
        is_self "$peer" && continue
        state=$(peer_state "$peer")
        if [ "$state" = rejoining ]; then
          echo "join: $peer is rejoining, not a peer to join through"
          continue
        fi
        echo "join: enabling replication with $peer"
        enable_through "$peer"

Or: the pending road's trusted "$state" check. It also keeps the ready road off pending peers, which the comment at :823-824 already expects.

Pin: restart two pruned volumes together while a third joiner holds the survivor's lock. Then assert that each of them lists the survivor in cn=admin data before it turns healthy.


issue (non-blocking): A volume left mid-rejoin and restarted without a join reports healthy, while its backend refuses every client write. Nothing in the log says why.

opendj-packages/opendj-docker/run.sh:154-155

"Without a join" means OPENDJ_REPLICATION_TYPE is no longer simple, or no peers are set. ! join_requested short-circuits the $REJOIN_PENDING test, so :155 touches the marker and :162 starts no join. internal-only is configuration, and only join.sh's release_writes sets it back. A join that gave up mid-rejoin (join.sh:762-763), or a container killed after hold_writes, then restarted with the replication environment removed, gets 53 on every write under BASE_DN indefinitely. That contradicts README.md:156-163. It is narrow: the environment has to change between starts while a rejoin is pending.

    if ! join_requested || { [ ! -f "$INITIALIZE_PENDING" ] && [ ! -f "$REJOIN_PENDING" ]; }; then
      if ! join_requested && [ -s "$REJOIN_PENDING" ]; then
        echo "The backend $(cat "$REJOIN_PENDING") still refuses client writes from an unfinished rejoin; set its writability-mode back to enabled with dsconfig set-backend-prop once its replication is settled"
      fi
      touch "$BOOTSTRAP_COMPLETE"

Or: release the hold in run.sh after start_server. Nothing will rejoin, so the hold no longer serves a purpose.


issue (non-blocking): joined() exits on a single failed leave_rejoin. A member that rejoined fine then stays read-only and unready until someone restarts the container.

opendj-packages/opendj-docker/bootstrap/join.sh:740, :728-735

leave_rejoin returns 1 when release_writes' dsconfig (:640) or publish_state ready (:734) fails once. exit 1 then ends join.sh, which run.sh starts once per container start (run.sh:163). That throws away the next round, which would reach is_member, joined() and the release again. Neither the HEALTHCHECK nor a readinessProbe restarts the container. The message at :731 describes a give-up, but nothing gave up here.

joined() {
  cleanup_departed
  leave_rejoin || return 1
  touch "$BOOTSTRAP_COMPLETE"

Change each caller from joined; exit to joined && exit. Inside the peer loop (:817-818) use joined && exit; break. A failed release then falls through to the round's pause, and give_up stays the end after the last round.


suggestion (non-blocking): Nothing pins seed()'s leave_rejoin || return 1. No CI seed follows a reset, so deleting it keeps CI green.

opendj-packages/opendj-docker/bootstrap/join.sh:751; .github/scripts/docker-test-replication.sh:232, :274, :503, :533, :402-446

The seed exits (join.sh:826, :835, :900, :912) never reach joined(), so this call is their only release. Every CI seed case starts on a fresh volume. The one reset case (dj-2) is not the first peer and leaves through joined(). If seed() stops releasing the writes, the server reports healthy and refuses every client write with 53.

Pin: restart the first peer dj-0, next to pending dj-1 and dj-2, on a volume left mid-reset after the disable: $REJOIN_PENDING naming userRoot, the backend internal-only, no domain. The ready road then seeds at :825-826. Once dj-0 is healthy, add_ou dj-0 seeded-after-reset must succeed; without the call it fails with 53.


suggestion (non-blocking): No CI case reaches whole_number's decimal reading or hold_writes' guard for a backend the operator set to something other than enabled.

opendj-packages/opendj-docker/bootstrap/join.sh:96, :627; .github/scripts/docker-test-replication.sh:402-446

Every knob the CI script passes is a plain decimal, and no case sets a writability mode before the dj-2 reset. Dropping 10#, or deleting [ "${mode:-enabled}" = enabled ] || return 0, keeps refuses_writes and add_ou dj-2 rejoined green. The second mutant turns an operator's internal-only backend into an enabled one at the rejoin, the opposite of README.md:158-159.

Pin:

  1. Run the dj-2 reset case with REPLICATION_RETRY_INTERVAL=08. Octal reading rejects 08, while 03 would pass either way. Assert that the case completes.
  2. Before that reset, run dsconfig set-backend-prop --backend-name userRoot --set writability-mode:internal-only on dj-2, and write the rejoined entry through dj-0. Then assert that dsconfig get-backend-prop --backend-name userRoot --property writability-mode on dj-2 still says internal-only.

issue (non-blocking): give_up reports "serves its data read-only" whenever $REJOIN_PENDING is non-empty. After a failed hold_writes dsconfig, the server is in fact writable and unreplicated.

opendj-packages/opendj-docker/bootstrap/join.sh:762-763, :632-633

hold_writes names the backend before its dsconfig, deliberately, for a kill in between. When the dsconfig fails, reset_replication keeps the server's replication (:589-592). A later give_up then tells the operator that the server refuses writes while it actually takes them. Decide the message from the backend's mode, and keep the write-before-dsconfig order.

  elif [ -s "$REJOIN_PENDING" ] && [ "$(data_backend | awk '{ print $2 }')" = internal-only ]; then

vharseko added a commit to vharseko/OpenDJ that referenced this pull request Oct 4, 2026
…gy's data, and clean every replication list

Review round 1 of OpenIdentityPlatform#1115:

- join.sh: a failed dsreplication initialize is no longer taken for a success (the
  exit status was that of an if without else), and it is bounded by its own
  REPLICATION_INITIALIZE_TIMEOUT, none by default; REPLICATION_ATTEMPT_TIMEOUT bounds
  the enable only, and the bound sits on the dsreplication call, as timeout cannot run
  a shell function.
- join.sh publishes pending/ready in the local entry cn=Docker Join,cn=config; a pending
  server enables and initializes only through a ready peer, the first peer seeds at
  once when every other one is pending and after the retries only while no peer that
  could hold the data answers. Fresh servers started together no longer take each
  other's bootstrap data for the topology's.
- The cleanup prunes departed servers from every replication list (replication server,
  BASE_DN, cn=schema, cn=admin data), and the lists get every listed, registered peer
  they lack; dsconfig inside the read loops reads /dev/null.
- Self-recognition: a name equal to hostname -f or cut from it at a dot, and the
  container's own addresses and their /etc/hosts names, compared whole; is_member
  checks every registered hostname; the generation ID comes from the domain's own
  monitor entry.
- run.sh stays unhealthy on a restart while $INITIALIZE_PENDING is on the volume.
- Dockerfile-alpine installs coreutils, whose timeout signals the whole process group.
- CI: both image jobs run .github/scripts/docker-test-replication.sh, which adds the
  initialize pin, held unhealthy checks, the full-list scale-down check, three fresh
  servers started at once, a master named by its own address, and the replicate.sh
  pins on the /dev/shm glob and the retry on exit 8.
- README documents all of the above.
vharseko added a commit to vharseko/OpenDJ that referenced this pull request Oct 4, 2026
…p, keep a seed ready across restarts, and take turns to enable

Review round 2 of OpenIdentityPlatform#1115:
- both image jobs check out .github/scripts, so the replication test runs at all
- run.sh marks the volume pending before the bootstrap, so a bootstrap that failed or was
  killed is initialized from the topology rather than published ready
- on a restart the volume is ready unless it is pending; a seed that no peer joined has no
  replication domain, and the join cannot bind once the root password was changed
- the first peer seeds after its retries on either road only while no other peer answers
  that may hold the data: ready, or publishing no state but replicating BASE_DN
- the state reaches the peers before the pending marker goes, or the join fails
- a server the survivors removed while it was away takes its replication configuration
  down and enables anew: dsreplication enable does not register a server whose replicated
  cn=admin data already matches the peer's
- servers registered by an address are not removed as departed
- the /tmp fallback of the password files has a name run.sh removes
- an empty list of replication servers changes nothing
- joins through one peer take turns, with a lock entry in the peer's cn=config: two enable
  runs at once leave the loser a member only in part, for good
- README: one ADMIN_PORT for the one-shot types, address-registered servers, the lock
- CI: a member ready while its peer is down, a seed ready after its root password changed,
  a failed bootstrap and a failed initialize, a scale-up on a retained volume, no lock left
  after the parallel start, and sdsr's exit 8 without a race: the replica starts on a network
  of its own rather than dj-0 leaving the topology's, which left dj-0's replication server a
  dead connection of its own directory server to route the replica's initialize into
@vharseko
vharseko force-pushed the feature/1086-replication-join branch from 9ecb440 to 442c5ff Compare October 4, 2026 12:16
@vharseko

vharseko commented Oct 4, 2026

Copy link
Copy Markdown
Member Author

Round 5 is in 442c5ff, rebased onto the current master (751a4d2). Every point held up. For the blocking one I took your alternative rather than the snippet, one fix differs from its snippet, and two pins are placed elsewhere. Point by point:

Blocking

  1. The ready road enables through a rejoining peer. Fixed with the trusted "$state" check, as on the pending road. Skipping only rejoining would not have been enough. Your trace holds for any two servers without replication: a reset server enabling with a pending one, or with a fresh one that still bootstraps and publishes nothing, forms the same two-server registry. member_of_topology accepts it because one peer of the pair lists the server, and the pending one then initializes from the reset one. A seed still works: a first peer whose other peers are all pending seeds through all_others_pending, and they enable through it on their own road. The cost is that when the survivor is gone for good and every other server is rejoining or pending, none of them turns healthy. That fits the rule the seed already follows: a lost topology shows up, it is not forked. The README and the header of join.sh now say that a server joins only through a ready peer.

    Pinned as you suggest, with one change: the survivor's lock is held by hand instead of by a third joiner. After the scale-up case, dj-1 and dj-2 are removed together and dj-0 restarts alone, which prunes both. With dj-0's lock held, dj-1 and then dj-2 start on their volumes. Each must log that it passes over the other as rejoining. Neither may list the other in its cn=admin data, neither may be ready, and dj-1 must refuse writes with 53. Once the lock is dropped, both turn healthy and register with dj-0, and a write on dj-1 reaches both other servers.

    While building the pin I found a gap that I am only recording here, not fixing. A server that comes back shows its peers the state it had before it left, because the state lives in its cn=config. That lasts until its first round finds that the peers no longer register it. Without care, dj-1 could have enabled through dj-2 in that window, so dj-2 keeps a lock held by hand on itself from before it left, and dj-1 is already rejoining before dj-2 starts. In that window the peer still has its old replication and its copy of cn=admin data, so it is not the empty pair of this finding. I have not traced what such an enable leaves.

Non-blocking

  1. A volume left mid-rejoin that starts without a join. Fixed with your log line. Releasing the hold in run.sh would need the server up: OpenDJ's dsconfig has no offline mode, and start_server ends in exec. So run.sh says which backend may still refuse writes and how to set it back. The README says the same.
  2. joined() exits on one failed leave_rejoin. Fixed, but not by changing the callers to joined && exit. On the pending road, joined follows a successful initialize with $INITIALIZE_PENDING already gone. A failed release there would fall through to the next round, where trusted_source runs a second full initialize. Instead, leave_rejoin now tries again itself, release_writes and publish_state ready together, for as long as the retries last, pausing between attempts. The join gives up only after that, with a message that says what failed: the writes, or the publish. seed() gets the same retry.
  3. "read-only" after a failed hold. Fixed as suggested. writes_held checks that $REJOIN_PENDING names a backend and that the backend is internal-only. give_up and leave_rejoin choose their message by it. The order in hold_writes (name the backend, then change the mode) stays.

Suggestions

  1. seed()'s release. Pinned with a cheaper fixture. Right after dj-0 seeds, and before dj-1 exists, dj-0 gets the state a reset cut off after its disable leaves: the backend internal-only, $REJOIN_PENDING naming userRoot, no domain. It is then restarted. It must leave its health to the join, seed again once its retries run out with no other peer answering, and then take add_ou dj-0 seeded-after-reset and publish ready. This is the seed exit at the end of the ready road. Both exits go through the same seed().
  2. 10# and the operator's mode. Both pinned, placed differently:
    • The leading zero is on the dj-f case next to dj-solo (REPLICATION_RETRY_INTERVAL=08), not on dj-2, whose case times its checks by interval 3. Just checking that the case completes does not kill the mutant there: on bash 3.2.57 and 5.2 the mutant abandons the rounds at their first pause and still ends with "could not join … after 2 attempts". So the case requires "(1 of 2), trying again in" and no fallback line.
    • A read-only backend set before dj-2's reset in the scale-up case would undo that case's pins: refuses_writes would pass whatever hold_writes does, and add_ou dj-2 rejoined could no longer show that the writes are let in again. So the operator's internal-only goes on dj-2 in the new two-server case, where dj-1 covers the release. After both rejoin, dj-2's writability-mode must still be internal-only.

Verification: .github/scripts/docker-test-replication.sh passed in full over the Debian image, every new case included. As before, the scripts of this head were laid over an image built from a server package of the branch. The local host was slow (its Docker VM had one CPU and 2 GB), so every wait_until got three times its budget, an enable 300 s instead of 90, and the health probe a 120 s timeout. These changes were made in the local copy only, and the run took 8070 s. An earlier run with the budgets as they are timed out on the dj-f give-up under a load average of about 800, after the case had logged the right first round. Alpine was not run locally. The mutants named in points 1, 5 and 6 were not run as separate builds; the 10# one was probed with the round loop on bash 3.2.57 and 5.2. The docker jobs of this head run the script on both images over the current master.

@vharseko
vharseko requested a review from maximthomas October 4, 2026 12:17
vharseko added a commit to vharseko/OpenDJ that referenced this pull request Oct 5, 2026
…gy's data, and clean every replication list

Review round 1 of OpenIdentityPlatform#1115:

- join.sh: a failed dsreplication initialize is no longer taken for a success (the
  exit status was that of an if without else), and it is bounded by its own
  REPLICATION_INITIALIZE_TIMEOUT, none by default; REPLICATION_ATTEMPT_TIMEOUT bounds
  the enable only, and the bound sits on the dsreplication call, as timeout cannot run
  a shell function.
- join.sh publishes pending/ready in the local entry cn=Docker Join,cn=config; a pending
  server enables and initializes only through a ready peer, the first peer seeds at
  once when every other one is pending and after the retries only while no peer that
  could hold the data answers. Fresh servers started together no longer take each
  other's bootstrap data for the topology's.
- The cleanup prunes departed servers from every replication list (replication server,
  BASE_DN, cn=schema, cn=admin data), and the lists get every listed, registered peer
  they lack; dsconfig inside the read loops reads /dev/null.
- Self-recognition: a name equal to hostname -f or cut from it at a dot, and the
  container's own addresses and their /etc/hosts names, compared whole; is_member
  checks every registered hostname; the generation ID comes from the domain's own
  monitor entry.
- run.sh stays unhealthy on a restart while $INITIALIZE_PENDING is on the volume.
- Dockerfile-alpine installs coreutils, whose timeout signals the whole process group.
- CI: both image jobs run .github/scripts/docker-test-replication.sh, which adds the
  initialize pin, held unhealthy checks, the full-list scale-down check, three fresh
  servers started at once, a master named by its own address, and the replicate.sh
  pins on the /dev/shm glob and the retry on exit 8.
- README documents all of the above.
vharseko added a commit to vharseko/OpenDJ that referenced this pull request Oct 5, 2026
…p, keep a seed ready across restarts, and take turns to enable

Review round 2 of OpenIdentityPlatform#1115:
- both image jobs check out .github/scripts, so the replication test runs at all
- run.sh marks the volume pending before the bootstrap, so a bootstrap that failed or was
  killed is initialized from the topology rather than published ready
- on a restart the volume is ready unless it is pending; a seed that no peer joined has no
  replication domain, and the join cannot bind once the root password was changed
- the first peer seeds after its retries on either road only while no other peer answers
  that may hold the data: ready, or publishing no state but replicating BASE_DN
- the state reaches the peers before the pending marker goes, or the join fails
- a server the survivors removed while it was away takes its replication configuration
  down and enables anew: dsreplication enable does not register a server whose replicated
  cn=admin data already matches the peer's
- servers registered by an address are not removed as departed
- the /tmp fallback of the password files has a name run.sh removes
- an empty list of replication servers changes nothing
- joins through one peer take turns, with a lock entry in the peer's cn=config: two enable
  runs at once leave the loser a member only in part, for good
- README: one ADMIN_PORT for the one-shot types, address-registered servers, the lock
- CI: a member ready while its peer is down, a seed ready after its root password changed,
  a failed bootstrap and a failed initialize, a scale-up on a retained volume, no lock left
  after the parallel start, and sdsr's exit 8 without a race: the replica starts on a network
  of its own rather than dj-0 leaving the topology's, which left dj-0's replication server a
  dead connection of its own directory server to route the replica's initialize into
@vharseko
vharseko force-pushed the feature/1086-replication-join branch from 442c5ff to 8038819 Compare October 5, 2026 08:53

@maximthomas maximthomas left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

praise: the round-5 island is closed on both roads now, and the new case checks it directly.

  • The ready road enables only through a peer that passes trusted "$state" (join.sh:838-842), so two reset servers no longer enable with each other.
  • admin_data_lacks dj-2 dj-1 / admin_data_lacks dj-1 dj-2 (docker-test-replication.sh:513-514) assert that the pair never registers itself, and both image cells passed them.
  • The dj-0 seed-after-reset case and REPLICATION_RETRY_INTERVAL=08 give seed()'s leave_rejoin and whole_number's decimal reading a CI observable.

issue (blocking): the case reads dj-1's stale healthy as the end of its rejoin and writes to it while its backend is still internal-only.

.github/scripts/docker-test-replication.sh:516-524

dj-1 starts on a volume without $REJOIN_PENDING, so run.sh:159 creates the marker and a probe passes. reset_replication removes the marker (join.sh:585), but HEALTHCHECK --interval=30s --retries=3 keeps the status healthy until three probes in a row fail, which is the delay the README describes. In build-docker job 111497688547, wait_healthy dj-1 (:517) returned while dj-1's enable was still running: its log ends at "Initializing registration information on server dj-1:4444", with no joined line. registered_in_dj0 passed because the enable registers dj-1 on dj-0 early. add_ou dj-1 then got 53 ("configured in read-only mode"), and set -e ended the step. Nothing else writes the marker during an enable: the other writers are run.sh:232, joined() and seed(), and the last two run only after leave_rejoin. leave_rejoin publishes ready only after release_writes (join.sh:743), so the case should wait for that instead:

drop_lock dj-0
published_ready() { [ "$(published_state "$1")" = ready ]; }
for n in dj-1 dj-2; do
  wait_until 480 "$n publishes itself ready again" published_ready $n
  wait_healthy $n
done

Both servers already publish rejoining at this point (:502 and :506 wait for it), so ready can only come from leave_rejoin. This fix does not turn build-docker-alpine green; the next comment covers that.


issue (blocking): a rejoin whose enable fails after it has configured the domains never recovers: every later enable exits 5, and no reset follows.

opendj-packages/opendj-docker/bootstrap/join.sh:823-852

In build-docker-alpine job 111497688528, dj-1's enable through dj-0 configured dc=example,dc=com and cn=schema and updated the registration on both servers. It then failed at "Initializing registration information ... with the contents of server dj-0:4444" with "Protocol error : a replication server is not expected to be the destination of a message of type ...InitializeRequestMsg" (STOPPED_BY_ERROR, exit 12). Each of the next 22 rounds, through dj-0 or dj-2, printed "There are no base DNs available to enable replication between the two servers ... already replicated" and "exited with 5 and membership is not there". reset_if_unregistered (:823) never fired, because the failed enable had already registered dj-1 on dj-0. The case's wait_healthy dj-1 (:517) timed out after 480 s. Outside CI, the same server would spend all of REPLICATION_RETRY_COUNT and end read-only and unhealthy. A rejoining server has nothing to keep, so take its replication down again after a failed enable instead of waiting for the peers to forget it:

        [ "$rc" -eq 75 ] || echo "join: dsreplication enable with $peer exited with $rc and membership is not there"
        # a failed enable may have configured the domains and registered this server with
        # its peer: no later round would reset it, and every later enable would exit 5
        if [ "$rc" -ne 0 ] && [ "$rc" -ne 75 ] && [ -f "$REJOIN_PENDING" ]; then
          reset_replication rejoining || true
        fi

Also add docker exec "$n" cat /opt/opendj/logs/errors to the ERR-trap dump. The exit 12 itself was not traced: the dump has container stdout only, not dj-0's errors log.


suggestion (non-blocking): the fixes in leave_rejoin, writes_held and run.sh's new message have no CI observable.

opendj-packages/opendj-docker/bootstrap/join.sh:731-758, opendj-packages/opendj-docker/run.sh:156-157

docker-test-replication.sh never asserts "not out of its rejoin yet", "serves its data read-only", "could not publish this server ready" or "may still refuse the writes". Reverting leave_rejoin to one attempt, reverting writes_held to [ -s "$REJOIN_PENDING" ], or deleting the run.sh line changes no observable in any case. The mutants were not run: CI is red at this head.

Pin: start the volume of the dj-0 seed-after-reset setup (backend internal-only, .replication-rejoin-pending naming userRoot) once without OPENDJ_REPLICATION_TYPE / REPLICATION_PEERS. Deleting run.sh:156-157 turns this red:

wait_until 300 "dj-0 says that no join lets the writes in" logs_have dj-0 "may still refuse the writes of clients"
wait_healthy dj-0
refuses_writes dj-0 held-without-join || fail "dj-0 took the writes of clients on a held backend"

The retry loop needs release_writes or publish_state to fail exactly once, and the docker CI has no hook for that.

…gy's data, and clean every replication list

Review round 1 of OpenIdentityPlatform#1115:

- join.sh: a failed dsreplication initialize is no longer taken for a success (the
  exit status was that of an if without else), and it is bounded by its own
  REPLICATION_INITIALIZE_TIMEOUT, none by default; REPLICATION_ATTEMPT_TIMEOUT bounds
  the enable only, and the bound sits on the dsreplication call, as timeout cannot run
  a shell function.
- join.sh publishes pending/ready in the local entry cn=Docker Join,cn=config; a pending
  server enables and initializes only through a ready peer, the first peer seeds at
  once when every other one is pending and after the retries only while no peer that
  could hold the data answers. Fresh servers started together no longer take each
  other's bootstrap data for the topology's.
- The cleanup prunes departed servers from every replication list (replication server,
  BASE_DN, cn=schema, cn=admin data), and the lists get every listed, registered peer
  they lack; dsconfig inside the read loops reads /dev/null.
- Self-recognition: a name equal to hostname -f or cut from it at a dot, and the
  container's own addresses and their /etc/hosts names, compared whole; is_member
  checks every registered hostname; the generation ID comes from the domain's own
  monitor entry.
- run.sh stays unhealthy on a restart while $INITIALIZE_PENDING is on the volume.
- Dockerfile-alpine installs coreutils, whose timeout signals the whole process group.
- CI: both image jobs run .github/scripts/docker-test-replication.sh, which adds the
  initialize pin, held unhealthy checks, the full-list scale-down check, three fresh
  servers started at once, a master named by its own address, and the replicate.sh
  pins on the /dev/shm glob and the retry on exit 8.
- README documents all of the above.
… out of command substitutions

set -E hands the ERR trap to every $(...): a command that failed there, as the
dsreplication enable the replicate.sh pin expects to fail, dumped the container logs
into the captured output and removed the containers under the running test. The trap
now acts in the main shell only.

join.sh waits for the peers to register with one search a round instead of one per
peer: every ldapsearch starts a JVM.
…p, keep a seed ready across restarts, and take turns to enable

Review round 2 of OpenIdentityPlatform#1115:
- both image jobs check out .github/scripts, so the replication test runs at all
- run.sh marks the volume pending before the bootstrap, so a bootstrap that failed or was
  killed is initialized from the topology rather than published ready
- on a restart the volume is ready unless it is pending; a seed that no peer joined has no
  replication domain, and the join cannot bind once the root password was changed
- the first peer seeds after its retries on either road only while no other peer answers
  that may hold the data: ready, or publishing no state but replicating BASE_DN
- the state reaches the peers before the pending marker goes, or the join fails
- a server the survivors removed while it was away takes its replication configuration
  down and enables anew: dsreplication enable does not register a server whose replicated
  cn=admin data already matches the peer's
- servers registered by an address are not removed as departed
- the /tmp fallback of the password files has a name run.sh removes
- an empty list of replication servers changes nothing
- joins through one peer take turns, with a lock entry in the peer's cn=config: two enable
  runs at once leave the loser a member only in part, for good
- README: one ADMIN_PORT for the one-shot types, address-registered servers, the lock
- CI: a member ready while its peer is down, a seed ready after its root password changed,
  a failed bootstrap and a failed initialize, a scale-up on a retained volume, no lock left
  after the parallel start, and sdsr's exit 8 without a race: the replica starts on a network
  of its own rather than dj-0 leaving the topology's, which left dj-0's replication server a
  dead connection of its own directory server to route the replica's initialize into
…ins, reset only on the peers' word, and lock both ends of an enable

- A server whose join takes its replication down to enable it anew stops reporting itself
  healthy first, and marks its volume ($REJOIN_PENDING) so that a restart does not report it
  healthy either until the enable registered it again.
- A failed search of a peer's cn=admin data no longer reads as "not registered": the reset
  needs at least one answering peer that replicates BASE_DN, and none of them listing it.
- A peer that refuses ROOT_PASSWORD (exit 49) may hold the data of the topology and holds the
  seed of the first peer off.
- The pending road decides membership with the peers as well, and resets a local domain they
  no longer register.
- An enable takes the join lock on this server as well as on the peer, the reset holds it,
  a lock left by an earlier start of this server is broken at once, and rounds wait a random
  part of the interval on top of it.
- REPLICATION_ATTEMPT_TIMEOUT that is not a positive number falls back to 120.
- CI pins: no reset on a plain restart, the reset and a restart during it stay unready while
  both peers' locks are held, a first peer next to a stateless member does not seed, and an
  address-registered master survives the move to REPLICATION_PEERS.
… rejoins, publish it as rejoining, retry the reset every round, and read the knobs as decimal

- Before a reset changes anything, the backend of BASE_DN is set to writability-mode
  internal-only, so clients cannot write while the server's replication is down. Replication
  and dsreplication still write. The health status lags by the probe's retries, so the marker
  alone could not stop these writes. joined() or seed() sets the mode back to enabled, and
  only where the join itself changed it ($REJOIN_PENDING names the backend).
- From the reset until it rejoins, a volume that holds the data publishes the state rejoining,
  on a restart too. No peer enables or initializes through it, and a first peer does not seed
  next to it.
- The reset makes one attempt at its own lock and runs at the start of every round on both
  roads. A reset that could not run no longer leaves the ready road enabling against the
  stale domain.
- REPLICATION_RETRY_COUNT, REPLICATION_RETRY_INTERVAL, REPLICATION_ATTEMPT_TIMEOUT and
  REPLICATION_INITIALIZE_TIMEOUT are read as decimal whole numbers. A leading zero no longer
  breaks the arithmetic and drops the ready road onto the pending one.
- CI pins: the reset waits for the server's own lock and keeps its domain meanwhile, an enable
  waits for the own lock while both peers are free, a lock left under the server's own name
  is broken at once, client writes are refused with 53 and the state is rejoining before and
  after a restart, and a first peer next to a peer that refuses the bind gives up instead of
  seeding. The wait after dj-2's restart counts only what the restarted join logged; the
  join of the start before could satisfy it while it was being stopped.
…ads, retry leaving a rejoin, and say when no join will let the writes in

- The ready road enables only through a peer that trusted() accepts, as the pending road
  already did. Two servers without replication - two that were reset, or one that was reset
  and one that is pending or still bootstraps - registered each other in a registry of their
  own, which passed for membership, and served a topology beside the survivors.
- leave_rejoin tries letting the writes in and publishing ready again for as long as the
  retries last, instead of ending the join on one failed dsconfig or publish. Ending it left a
  member that had rejoined read-only and unready until the next start, and a retry from the
  round loop would have initialized again on the pending road.
- give_up and leave_rejoin say "read-only" only when the backend is in fact internal-only: a
  hold whose dsconfig failed names the backend without changing its mode.
- run.sh logs that the backend may still refuse writes when a volume left mid-rejoin starts
  without a join, which is the only thing that lets them in again.
- CI pins: a seed after a reset lets the writes in again (dj-0 restarted with the backend
  held, $REJOIN_PENDING and no domain); REPLICATION_RETRY_INTERVAL=08 is read as 8 and the
  rounds go on; two servers reset together pass over each other while the survivor's lock is
  held, register with it once it is free, and a backend the operator made read-only stays so.
… the server's own cn=admin data does not register it, and wait for a rejoin by the published state in the CI test

- join.sh: a round also resets a server that replicates BASE_DN and is registered at its
  peers but not in its own cn=admin data (or has no cn=Servers there at all). An enable whose
  initialize of cn=admin data failed leaves that state, and every later enable exits 5 without
  registering it. A local search that fails otherwise decides nothing. The reset logs why.
- CI: the two servers scaled up together are waited for by their published ready state
  before their health. They started without a reset pending, so a probe may have found them
  healthy before their joins took the replication down, and the status turns unhealthy only
  after the probe's retries.
- CI: new case, a server whose own cn=admin data lacks its entry, then one that lacks
  cn=Servers, made by deletes with the replication repair control; each takes its
  replication down and rejoins.
- The plain-restart check looks for any reset, not only the one the peers cause.
- README: the second reason to reset.
… logs of a failed replication test, and pin the message of a held backend started without a join

- CI: the ERR trap also prints, for every running container, the last 1000 lines of the
  server's errors log and every detailed dsreplication log left in the instance's tmp
  directory (both under /opt/opendj/data): the container log says that an enable failed,
  these say why.
- CI: the dj-0 seed-after-reset volume (backend internal-only, $REJOIN_PENDING naming it) is
  started once without OPENDJ_REPLICATION_TYPE: it must say which backend may still refuse
  the writes, report itself healthy, and refuse a client's write. Then it is started with
  its join as before, in a new container, so the seed is looked for in that one's log.
@vharseko
vharseko force-pushed the feature/1086-replication-join branch from 8038819 to 667e9af Compare October 5, 2026 13:01
@vharseko

vharseko commented Oct 5, 2026

Copy link
Copy Markdown
Member Author

Both blocking items describe 442c5ff and its run 37201543789. The fix for that run was already on the branch: 8038819 was pushed at 08:54, before this review, but I had not posted its reply, so here it is together with this round. The lines you cite are those of 442c5ff: at 8038819, join.sh:823 is a comment, and the case waits for published_ready before wait_healthy.

The branch is rebased onto master 85b28b3, and the earlier commits are unchanged (range-diff all =); 8038819 is now 8938ae7. This round is 667e9af.

Round 6 (8938ae7, was 8038819): the two red docker jobs of 442c5ff

build-docker: add_ou dj-1 rejoined-together got 53 (read-only) right after wait_healthy dj-1 passed. This was the test, not the join. dj-1 and dj-2 start on volumes with no reset pending, so run.sh writes the health marker as it does for any member. A probe can find them healthy before their joins take the replication down, and the status turns unhealthy only after three failed probes. wait_healthy read that stale status. The case now waits for both servers to publish ready, which leave_rejoin does only after the writes are let in, and checks the health after that. This is the same as your snippet.

build-docker-alpine: dj-1 never turned healthy. Its enable through dj-0 failed with exit 12 while initializing cn=admin data. That left BASE_DN replicated and dj-1 registered at its peers, but not in its own cn=admin data. reset_if_unregistered asked only the peers, and they register it, so it never reset. is_member reads the local registry, so the server was never a member. Every later enable exited 5, "already replicated".

Round 6 also takes the replication down when the peers register the server but its own cn=admin data does not, or holds no cn=Servers at all. A local search that fails any other way decides nothing. A full member registers itself locally, and a server without a replication domain for BASE_DN is never checked, so a plain restart does not reset. The plain-restart check of the test now looks for any reset. The reset logs which reason applied.

New case: dj-1 deletes its own entry, and then dj-2 the whole cn=Servers, from their own cn=admin data only, with the replication repair control. Each must log the new reason, rejoin, publish ready, be registered locally and at dj-0, and turn healthy; a write then reaches the other two. Run on its own, over the Debian image with the scripts laid over it and every wait at three times its budget: it passes at the head (1644 s) and fails with join.sh of round 5 ("timed out waiting until dj-1 takes its half-enabled replication down"). CI of 8038819 (run 37286526709): its docker jobs were in the middle of "Docker test replication" on both images, with nothing failed so far, when the push of 667e9af cancelled them; the run of 667e9af covers the same cases.

This review

1. Stale healthy before add_ou dj-1 (blocking). Fixed in round 6, as above. One detail: dj-1's log in job 111497688547 does not end at "Initializing registration information on server dj-1:4444". It goes on to "Replication has been successfully enabled", with neither "joined" nor "exited with" after it. The conclusion is the same.

2. A failed enable that left the domains configured (blocking). Fixed in round 6, but not with the reset after every failed enable. The log does not show the registration updated on both servers. "Updating registration configuration on server dj-1:4444" is INFO_REPLICATION_ENABLE_CONFIGURING_ADS (ReplicationCliMain.java:6346): it configures the replication domain of cn=admin data on dj-1 and does not write to its registry. The enable registers both servers in the registry of the source only (:4629-4732: each branch that names a destination registers through the source's context alone). The destination gets its copy from the initialize that follows (:4887), and that initialize is what failed. So dj-1's own cn=admin data did not register it, which is the state round 6 resets.

I kept that check over the reset after a failed enable for two reasons:

  • It covers both roads. The same half-done enable can happen to a pending server's first join, where $REJOIN_PENDING does not exist.
  • It does not reset after failures that configured nothing: an unreachable peer (8), an attempt cut off by the timeout (124), or the exit 5 of a peer that has no BASE_DN yet.

The errors log dump is in 667e9af, with a different path. The image keeps the instance under /opt/opendj/data (the dsreplication log path in the job, /opt/opendj/data/tmp/opendj-replication-….log, shows it), so the trap prints the last 1000 lines of /opt/opendj/data/logs/errors, followed by every detailed dsreplication log still in /opt/opendj/data/tmp, for each running container. The cause of the exit 12 itself is still open. My guess is that the replication server the enable created on dj-1 got id 2861, the id of dj-0's directory server for cn=admin data, but the containers were gone. With the dump, the next occurrence will show it. That is a dsreplication matter, not one for join.sh.

3. No CI observable for leave_rejoin, writes_held and the run.sh message (suggestion). Your pin for run.sh is in 667e9af. The dj-0 seed-after-reset volume (backend internal-only, $REJOIN_PENDING naming userRoot) is first started without OPENDJ_REPLICATION_TYPE. It must log "The backend userRoot may still refuse the writes of clients", turn healthy, and refuse a write with 53. Then it is started with its join as before. That is a new container, so the case now looks for the seed in its log, not for a second seed in the old one. The retries of leave_rejoin, and the message writes_held picks, need a dsconfig or a publish that fails exactly once. As you say, the docker CI has no hook for that, so I record them here as unpinned and do not add one.

Verification of 667e9af: it was not run locally. Docker was not available on my host, and this round changes only the CI script. The docker jobs of run 37313618947 run the full script on both images, the new case included.

@vharseko
vharseko requested a review from maximthomas October 5, 2026 13:03

@maximthomas maximthomas left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

praise: The half-enabled reset closes the stuck rejoin from round 6, and the CI now proves it and keeps the evidence.

  • reset_if_unregistered (join.sh:836-839) catches the state behind round 6's red alpine cell. The failed enable's "Updating registration configuration" line is configureToReplicateBaseDN for cn=admin data (ReplicationCliMain.java:6343-6346), not a registry write. So the server's own cn=Servers lacks it while dj-0 registers it, which is exactly what registered_here checks.
  • half_enabled_rejoins builds both variants with the repair control. That control keeps the deletes local (MultimasterReplication.java:164-175 sets dontSynchronize), and the case asserts the precondition with registered_in_dj0.
  • The "two servers taken out together" case now waits on published_ready before the health status (docker-test-replication.sh:565-569). dump_logs adds the errors log and the dsreplication logs to the ERR-trap dump.

@vharseko
vharseko merged commit 947c0c9 into OpenIdentityPlatform:master Oct 6, 2026
24 checks passed
@vharseko
vharseko deleted the feature/1086-replication-join branch October 6, 2026 06:58
vharseko added a commit to vharseko/OpenDJ that referenced this pull request Oct 9, 2026
…did not complete unhealthy, the background join's too, and leave unanswering peers out of the replication server lists

Review round 1 of OpenIdentityPlatform#1184, and OpenIdentityPlatform#1185.

- run.sh marks a volume .upgrade-pending before it upgrades it to another
  version, and removes the mark once the upgrade, post-upgrade tasks
  included, succeeded (OpenIdentityPlatform#1185). The upgrade records the new version in
  buildinfo before those tasks - the index rebuilds and the verification of
  the DN equality indexes - so after one failed, or was cut off by a stop, a
  second run found nothing to do and the container reported itself healthy.
  A start over a marked volume whose version is already the one of the image
  stays unhealthy and says what to do: rebuild the indexes upgrade.log names
  and remove the mark, or restore the backup. An upgrade that failed before
  it recorded the version is run again, as before. The upgrade step of
  build.yml checks that a successful upgrade leaves no mark, that a mark
  over an upgraded volume keeps the container unhealthy until it is removed,
  and that an upgrade failed by a read-only buildinfo leaves the mark.
- docker-test-replication.sh: dj-bf, the background-join volume whose
  bootstrap fails (OpenIdentityPlatform#1115), followed the recovery that OpenIdentityPlatform#1182 now refuses, and
  failed both Docker jobs. It now stays unhealthy after a restart and runs no
  join, and initializes from dj-0 once the mark is removed by hand: the
  initialize would bring the data of the topology, but nothing else the
  bootstrap was to configure. The README says which road such a volume takes.
- run.sh runs the upgrade over a volume that carries .bootstrap-pending as
  well: start-ds refuses an instance of another version, and the setup could
  not be finished by hand under a newer image.
- join.sh: the repair of the replication server lists leaves every peer that
  does not answer for a later start. A combined server added a directory
  server that was down for good when nothing it asked showed another role,
  and the peers every list holds are never asked. A replication server left
  out adds this server to its own lists when it starts.
- join.sh: a replication server publishes rejoining on the ready road until a
  round finds it a member, and ready then, so that no peer enables through it
  before a reset takes it out. Not pending: that would count it among the
  peers that wait for data and let a fresh seed go ahead.
- build.yml: the upgrade step warns and is skipped when the latest release has
  no image on Docker Hub yet - release.yml pushes the images after it
  publishes the release.
- docker-test-replication.sh: dj-nm, which publishes no role, is given a
  replication server; a combined, a replication and a directory volume that
  lost .replication-role are recorded in their role again; setup.sh refuses
  an unknown REPLICATION_ROLE in a container and records nothing; each
  replication server of the OpenIdentityPlatform#534 topology is stopped in turn; the comment of
  roles_hold names the server that joined through dj-d0.
- Installation and Administration Guides (OpenIdentityPlatform#1179, now on master): a restart
  over a volume whose first start failed stays unhealthy, and a setup
  finished by hand goes on once .bootstrap-pending is removed; a restart
  after a failed upgrade stays unhealthy too, and .upgrade-pending is
  removed once the indexes are rebuilt, instead of "do not restart"; the
  replication note names REPLICATION_ROLE and links the standalone
  replication server and directory server sections.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Docker image: joining replication is one-shot, fixed-master and unchecked, which a StatefulSet cannot rely on

2 participants