Skip to content

[#1178] Give the background join of the Docker image a server role, keep a failed bootstrap or upgrade unhealthy, and test the upgrade from a released image - #1184

Open
vharseko wants to merge 2 commits into
OpenIdentityPlatform:masterfrom
vharseko:feature/1178-docker-replication-role
Open

vharseko wants to merge 2 commits into
OpenIdentityPlatform:masterfrom
vharseko:feature/1178-docker-replication-role

Conversation

@vharseko

@vharseko vharseko commented Oct 8, 2026 •

Copy link
Copy Markdown
Member

Summary

#1178: a server role for the background join

OPENDJ_REPLICATION_TYPE=simple made every server a directory server and a replication server at once. The one-shot srs and sdsr types, the only way to run a replication server without data or a directory server without one, are deprecated in its favour. REPLICATION_ROLE now sets the part a server of simple takes: combined (the default, as before), directory or replication.

  • Each side of dsreplication enable gets its own role. Without the role flags the tool gives both servers a replication server, so a join through a directory server handed it one (ReplicationCliMain.configureServer).
    • The peer's role is read from the peer, under its join lock, in one search of its cn=config. That search returns its replication server, its domain and its backend for BASE_DN, plus the role it publishes next to its state in cn=Docker Join.
    • A peer that publishes no state is combined, as every server was before roles. That covers a peer on an image before roles, or a master that runs no join.
  • The role is recorded on the volume and does not change after that. setup.sh checks REPLICATION_ROLE before it sets anything up and writes it to .replication-role in the data directory.
    • A later start with another value only logs that the server keeps its role. So does an unknown value.
    • A volume set up by an image before roles is read once from its replication configuration, then recorded. A replication server without a domain counts as the replication server of srs only when REPLICATION_ROLE=replication says so.
  • Membership, the reset of an unregistered server and the peers asked about registration follow the role.
    • A replication server is a member through its replication server, and holds no backend whose writes a reset would hold.
    • A peer whose only replication configuration is a replication server is asked about registration only when it publishes the role replication. On a combined server that state is an enable that stopped half-way.
  • A replication server holds no data and is never a seed.
    • It goes the pending road, and nothing initializes it. On a later start it publishes rejoining until a round finds it a member, so that no peer enables through it before a reset takes it out.
    • The first peer of REPLICATION_PEERS that holds data seeds the topology. Replication servers listed ahead of it are passed over, by the role they publish.
  • A directory server does not enable with a directory peer that no replication server joined yet. dsreplication enable refuses a topology without one, and the log says why.
  • The repair of the replication server lists adds only peers that run a replication server.
    • Each peer a list lacks is asked whether it runs one.
    • A peer that does not answer is left for a later start, whatever the roles: a directory server put into a list by mistake would stay there for good, while a replication server left out is added later, and adds this server to its own lists when it starts.
  • setup.sh creates no backend and loads no entries for a replication server of the background join.
  • BASE_DN goes into the search filters of the join escaped (RFC 4515).

#1182: a failed bootstrap stays unhealthy across restarts

run.sh marks a volume .bootstrap-pending before its bootstrap. It removes the mark only once the whole bootstrap has succeeded: the bootstrap script, and the replicate.sh of a one-shot type.

A start over a volume that still carries the mark starts the server so it can be looked at. It runs the upgrade (start-ds refuses an instance of another version) but not the join, never writes the health marker, and logs what to do: remove the volume, or finish the setup by hand and remove the mark. This holds for a volume of the background join as well: an initialize from the topology would bring its data, but nothing else the bootstrap was to configure. Once the mark is removed, such a volume joins and initializes from the topology. Volumes of earlier images carry no mark and start as before.

#1185: a failed upgrade stays unhealthy across restarts

The upgrade records the new version in buildinfo before its post-upgrade tasks (the index rebuilds and the verification of the DN equality indexes), so after one of them failed, or was cut off by a stop, the next run of upgrade found nothing to do and the container reported itself healthy.

run.sh marks a volume .upgrade-pending before it upgrades it to another version (major.minor.point of data/config/buildinfo against the image's template/config/buildinfo, as BuildVersion.equals compares them), and removes the mark once upgrade succeeded, post-upgrade tasks included. A mark over a volume whose version is already the image's can only come from an upgrade that recorded the version and did not complete: that start stays unhealthy and logs what to do (rebuild the indexes upgrade.log names and remove the mark, or restore the backup). An upgrade that failed before it recorded the version is run again, as before.

#1183: the upgrade from a released image in CI

Both Docker jobs bootstrap a volume with openidentityplatform/opendj:${release_version} (-alpine on the Alpine job). They then start the build image over it and check three things: the volume was upgraded (no "has already been upgraded"), the container reports healthy, and it serves the entries.

It also covers #1185: a successful upgrade leaves no .upgrade-pending; a mark over an upgraded volume keeps the container unhealthy until it is removed; an upgrade failed by a read-only buildinfo (it fails in changeBuildInfoVersion, after the upgrade tasks) leaves the mark. The step warns and is skipped when the latest release has no image on Docker Hub yet: release.yml pushes the images after it publishes the release.

A second new step covers #1182: a restart over a volume whose create-backend failed stays unhealthy, and turns healthy once the setup is finished by hand and the mark removed.

Documentation

README.md:

  • the new Server roles section;
  • how srs and sdsr map onto simple;
  • the Kubernetes layout with separate replication servers;
  • REPLICATION_ROLE in the environment table;
  • the .bootstrap-pending and .upgrade-pending marks.

Installation Guide (To Run OpenDJ in Docker, To Upgrade a Docker Container, from #1179): a restart after a failed first start or a failed upgrade stays unhealthy, and how to get past it by removing the mark, instead of "remove the volume" and "do not restart". Administration Guide, the replication NOTE: REPLICATION_ROLE, linked to the standalone directory server and replication server sections.

.github/scripts/docker-test-replication.sh gains:

  • the Missing Changes count is piling up during peak load testing with OpenDJ 4.9.4 #534 topology (two directory and two replication servers): roles held across restarts and when started again in another role, the roles published and recorded, each replication server stopped in turn;
  • the role of a volume set up before roles (no .replication-role) inferred again, for combined, directory and replication volumes;
  • setup.sh refusing an unknown role in a container;
  • the failed background-join bootstrap (dj-bf) staying unhealthy until the mark is removed, then initializing from the topology;
  • a replication server and a directory server joining through a replication server;
  • directory servers with no replication server;
  • a replication server listed first;
  • an unknown role;
  • a replica joining through a master that runs no join, which is given a replication server.

Testing

Run locally on a 1-CPU Docker VM, with images built from this branch over a 5.2.0-SNAPSHOT build of master, and the timeouts of the script tripled:

  • The whole of docker-test-replication.sh on Debian passed, in 20403 s.
  • On Alpine it failed at [#1086] Join replication in the background on every start of the Docker image #1115's dj-2, the third server's first join: its first dsreplication enable ran past the 270 s attempt bound and was cut off, and every later round's reset failed in dsreplication disable (exit 8, ERROR_CONNECTING). That code is master's and the case passed on Debian; it reads as the recovery after an enable cut off half-way, exposed by an overloaded host, and is not proven either way.
  • Round 1 (before review): the REPLICATION_ROLE block passed, and failed on the scripts of master as expected (the directory server dj-d0 runs a replication server).
  • The CI steps failed bootstrap and upgrade from a released image (with the Docker image: a restart after a failed post-upgrade task reports healthy, with the upgrade never completed #1185 checks) were run as GitHub runs a bash step (bash -eo pipefail), with the same text as in the workflow, on Debian and Alpine: they pass, the upgrade from 5.1.2 and 5.1.2-alpine. failed bootstrap on the run.sh of master fails: after the restart the container reports healthy with no userRoot.
  • The functions that tell a role from a replication configuration were checked against 30 combinations of input.
  • The changed AsciiDoc chapters render with AsciidoctorJ 2.5.3 (-v) without warnings.

Not run locally: the rest of the workflow, which CI runs.

Notes

  • dsreplication enable between two replication servers finds BASE_DN only on a directory server of the topology it can reach. A replication server that joins through another one while every directory server is down waits for one to come back. This is documented in the README.

Fixes #1178
Fixes #1182
Fixes #1183
Fixes #1185

@vharseko
vharseko requested a review from maximthomas October 8, 2026 18:48
@vharseko vharseko added enhancement bug docker replication kubernetes Kubernetes / Helm / OpenShift deployment CI upgrade Upgrading between versions and migrating from other directory servers tests Test suites: fixing, enabling, un-disabling docs setup setup / upgrade / uninstall tools (quicksetup) and the launcher scripts labels Oct 8, 2026

@maximthomas maximthomas left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

praise: The role is settled up front and kept with the volume, so a later environment change cannot quietly turn a server into something else.

  • setup.sh:35-46 refuses an unknown REPLICATION_ROLE before anything is set up, and writes the role to .replication-role.
  • join.sh:1037-1047 refuses a .replication-role that holds anything else. It infers the role of a pre-roles volume once, with legacy_role, and records it.

issue (blocking): The unchanged dj-bf case contradicts the new .bootstrap-pending branch, and both build-docker and build-docker-alpine are red at 23a3a58.

.github/scripts/docker-test-replication.sh:403-419, opendj-packages/opendj-docker/run.sh:149-153

dj-bf is a background-join node whose BOOTSTRAP fails. The case restarts it and expects wait_healthy dj-bf and "initializing from dj-0" (:414-417). With #1182, that restart takes the pending branch: no upgrade, no join.sh, no health marker. dj-bf logs "The bootstrap of this volume did not complete…" and the script fails with timed out waiting until dj-bf is healthy (run 37827037318, step "Docker test replication", both jobs). Every step after it was skipped. That covers the new REPLICATION_ROLE block of the same script, "Docker test failed bootstrap" and "Docker test upgrade from a released image", none of which has run in CI. Which contract do you want? Should a background-join volume whose bootstrap failed still recover on restart by initializing from the topology, as #1115's dj-bf pins? Or should it stay unhealthy until someone fixes it by hand, as #1182 now does? Your answer picks the fix below, and either fix needs a green run of both Docker jobs.

# dj-bf under the #1182 contract
docker restart dj-bf >/dev/null
wait_until 300 "dj-bf reports its pending bootstrap" logs_have dj-bf "The bootstrap of this volume did not complete"
[ "$(docker inspect -f '{{.State.Health.Status}}' dj-bf)" != healthy ] || fail "dj-bf turned healthy over a failed bootstrap"
docker exec dj-bf rm /opt/opendj/data/.bootstrap-pending
docker restart dj-bf >/dev/null
wait_healthy dj-bf
logs_have dj-bf "initializing from dj-0" || fail "dj-bf joined with the data of its failed bootstrap"

Or keep #1115's recovery: let a volume that still carries INITIALIZE_PENDING bypass the pending branch, so the initialize from the topology replaces whatever the failed bootstrap left (if [ -f "$BOOTSTRAP_PENDING" ] && ! { join_requested && [ -f "$INITIALIZE_PENDING" ]; }; then). Then make the README's .bootstrap-pending paragraph say which road a background-join volume takes.


issue (non-blocking): A start over a .bootstrap-pending volume also skips ./upgrade, so under a newer image start-ds refuses to start and the container crash-loops.

opendj-packages/opendj-docker/run.sh:149-153

The upgrade sits in the elif. A marked volume started under an image of another version therefore reaches BuildVersion.checkVersionMismatch() (DirectoryServer.java:1354, BuildVersion.java:100-104), which throws when the binary and instance versions differ. start-ds exits, PID 1 exits, and the "finish the setup by hand" road that the log line offers cannot be taken under that image. Same-image restarts, which is what the CI step does, are not affected.

if [ -f "$BOOTSTRAP_PENDING" ]; then
  sh ./upgrade -n --force || true
  echo "The bootstrap of this volume did not complete, this container will not report itself healthy: ..."
elif sh ./upgrade -n --force; then

issue (non-blocking): repair_replication_servers adds a directory server that does not answer to every replication server list of a combined server, when that directory server is the only peer the lists lack.

opendj-packages/opendj-docker/bootstrap/join.sh:879-911

mixed turns yes only when a lacking peer answers without server. Peers that every list already holds are never asked. Take combined A and B and directory C, with C down while A restarts. B is skipped, C fails replication_config, mixed stays no, and C:REPLICATION_PORT is added to A's server list and to its BASE_DN, cn=schema and cn=admin data domain lists. On later starts every list holds C, so C is never asked again, and cleanup_departed keeps it. That is the outcome the comment at :891-896 says this arm prevents. Replication keeps working, but A dials a port that nothing listens on for good. A peer replication server left out is added on a later start, as the comment already accepts, so the unknown hosts can be deferred on every role:

for host in $unknown; do
  echo "join: could not ask $host whether it runs a replication server, its place in the replication server lists is checked on a later start"
done

issue (non-blocking): The upgrade step pulls openidentityplatform/opendj:${release_version} from Docker Hub with no fallback, so Docker CI goes red on every PR between a GitHub release and its Docker push.

.github/workflows/build.yml:829-833

release_version is the name of releases/latest (:499-501). release.yml publishes that release in release-maven and pushes the Docker tags only in release-docker, which needs release-maven. In that window, and for as long as a failed release-docker stays unfixed, the pull fails in both jobs on every PR, through nothing the PR did.

docker pull "$RELEASED" || { echo "::warning::$RELEASED is not on Docker Hub yet, the upgrade test is skipped"; exit 0; }

issue (non-blocking): On the ready road, a replication server publishes ready before its first membership check. On the pending road, ready means it is already a member.

opendj-packages/opendj-docker/bootstrap/join.sh:1104-1112

The comment's premise, "this volume holds the data of the topology", does not hold for a replication server. Two servers reach this road: an RS that the survivors no longer register, and a legacy srs volume restarted as simple with REPLICATION_ROLE=replication. Either one advertises ready until reset_if_unregistered publishes rejoining in its first round. A pending peer that enables through it in that window may be counted a member through the RS's own registry. Not run: this needs a multi-container run, and I have not traced whether that leaves a lasting island or heals after the RS resets. Publishing pending for ROLE=replication until a round finds it a member would match the pending road.


suggestion (non-blocking): The dj-nm/dj-nmr case does not pin that a peer which publishes no role is taken for combined.

.github/scripts/docker-test-replication.sh:742-747

dj-nm runs no join, so published_role falls through to *) echo combined (join.sh:400-406). Consider a mutant that returns directory for a peer that publishes nothing. It passes --noReplicationServer1 for dj-nm, dj-nmr's own RS carries the initialize, and both assertions stay green while the legacy master is left without a replication server.

runs_replication_server dj-nm || fail "dj-nm, which publishes no role, was not given a replication server"

Pin: after :747, kills the mutant above.


suggestion (non-blocking): legacy_role, which infers the role of a volume set up before roles, has no CI observable.

opendj-packages/opendj-docker/bootstrap/join.sh:1018-1048

It runs only when .replication-role is missing. Every replication volume in the script is set up by this image, which writes the file, and the released-image upgrade step sets no OPENDJ_REPLICATION_TYPE. A mutant such as server) always echoing combined stays green. Under that mutant, an srs replication server upgraded with REPLICATION_ROLE=replication is recorded as combined for good.

docker exec dj-r0 rm /opt/opendj/data/.replication-role
docker restart dj-r0 >/dev/null
wait_healthy dj-r0
[ "$(recorded_role_of dj-r0)" = replication ] || fail "dj-r0 was recorded $(recorded_role_of dj-r0) from its replication configuration"

Pin: the same on a combined volume, expecting combined.


suggestion (non-blocking): No case exercises setup.sh refusing an unknown REPLICATION_ROLE, which is the road a real container takes.

opendj-packages/opendj-docker/bootstrap/setup.sh:35-46, .github/scripts/docker-test-replication.sh:722-725

The only unknown-role case runs join.sh through --entrypoint, with no volume, no run.sh and no setup.sh. A mutant that drops setup.sh's *) … exit 1 writes bogus to .replication-role, sets the volume up as a data holder, and stays green.

start_node dj-bogus dj-0,dj-bogus -e REPLICATION_ROLE=bogus
wait_until 120 "setup refuses the role" logs_have dj-bogus "REPLICATION_ROLE=bogus is not one of combined, directory or replication"
if docker exec dj-bogus test -e /opt/opendj/data/.replication-role; then fail "setup recorded an unknown role"; fi

suggestion (non-blocking): "one replication server is enough" stops dj-r1 without checking that any directory server was connected to it.

.github/scripts/docker-test-replication.sh:868-871

If both directory servers sit on dj-r0, the step passes with no failover at all. Not traced: whether the broker always spreads two equal-weight replication servers one per directory server. Stopping each RS in turn makes the step independent of that choice.

docker start dj-r1 >/dev/null; wait_healthy dj-r1
docker stop dj-r0 >/dev/null
add_ou dj-d0 roles-other-rs
wait_has dj-d1 "ou=roles-other-rs,dc=example,dc=com"

nitpick (non-blocking): roles_hold's comment "every server joined through dj-d0 at some point" is false at its third call.

.github/scripts/docker-test-replication.sh:791-792

Before :866, dj-r1 and dj-d1 come back on empty volumes with dj-d0 stopped, and the script asserts that they joined through dj-r0 and dj-r. The check still holds because dj-r0 joined through dj-d0. Suggested wording: "dj-r0 joined through dj-d0: an enable that took dj-d0 for a combined server would have given it a replication server".

…age a server role, keep a failed bootstrap unhealthy, and test the upgrade from a released image

OPENDJ_REPLICATION_TYPE=simple made every server a directory server and a
replication server at once, and the one-shot srs and sdsr types - the only way
to run a replication server without data, or a directory server without a
replication server - are deprecated in its favour. REPLICATION_ROLE now sets
the part a server of simple takes: combined (the default, as before),
directory or replication.

- join.sh passes each side of dsreplication enable its own role. The peer's is
  read from the peer in one search of its cn=config - its replication server,
  its domain and its backend for BASE_DN - or, for a seed nobody joined yet,
  from the role it publishes next to its state in cn=Docker Join; a peer that
  publishes none, an image before roles or a master that runs no join, is
  combined. The peer is read under its join lock. Without the role flags the tool gives both servers a
  replication server, so a join through a directory server handed it one
  (ReplicationCliMain.configureServer).
- setup.sh records the role on the volume, and that is the role of the server
  from then on: a start in another role is logged and changes nothing. A
  volume set up by an image before roles is read once from its replication
  configuration - a replication server without a domain is the replication
  server of srs where REPLICATION_ROLE=replication says so - and recorded.
- Membership, the reset of an unregistered server and the peers asked about
  registration follow the role: a replication server is a member through its
  replication server, and holds no backend whose writes a reset would hold.
- A replication server goes the pending road and is never initialized or a
  seed: the first peer of REPLICATION_PEERS that holds data seeds, replication
  servers listed ahead of it passed over by the role they publish.
- A directory server does not enable with a directory seed no replication
  server joined yet, which dsreplication enable refuses, and says why.
- The repair of the replication server lists adds only peers that run a
  replication server, asking each peer a list lacks; one that does not answer
  goes in as before on a combined server whose answering peers all run one.
- setup.sh creates no backend and loads no entries for a replication server
  of the background join, and refuses an unknown REPLICATION_ROLE before it
  sets anything up.
- An unknown REPLICATION_ROLE ends the join.
- BASE_DN goes into the search filters of the join escaped (RFC 4515).

README: Server roles, the mapping of srs and sdsr onto simple, and the
Kubernetes layout. docker-test-replication.sh: the OpenIdentityPlatform#534 topology - two
directory and two replication servers - with roles held across restarts and
one replication server stopped, the roles each server publishes and records, a
replication server and a directory server joining through a replication
server, directory servers with no replication server,
an unknown role, a replication server listed first, servers started again in
another role, and a replica joining through a master that runs no join.

OpenIdentityPlatform#1182: run.sh marks a volume .bootstrap-pending before its bootstrap and removes
the mark only once the whole bootstrap - the bootstrap script, and the
replicate.sh of a one-shot type - succeeded. A start over a volume that still
carries it starts the server to be looked at, runs neither the upgrade nor the
join, never writes the health marker, and logs what to do: remove the volume,
or finish the setup by hand and remove the mark. Volumes of earlier images
carry no mark and start as before.

OpenIdentityPlatform#1183: both Docker jobs bootstrap a volume with the latest released image of
their flavour, start the build image over it, and check that it upgraded the
volume, reports healthy and serves the entries. They also check that a
restart over a volume whose create-backend failed stays unhealthy, and that
it turns healthy once the setup is finished by hand and the mark removed.
…did not complete unhealthy, the background join's too, and leave unanswering peers out of the replication server lists

Review round 1 of OpenIdentityPlatform#1184, and OpenIdentityPlatform#1185.

- run.sh marks a volume .upgrade-pending before it upgrades it to another
  version, and removes the mark once the upgrade, post-upgrade tasks
  included, succeeded (OpenIdentityPlatform#1185). The upgrade records the new version in
  buildinfo before those tasks - the index rebuilds and the verification of
  the DN equality indexes - so after one failed, or was cut off by a stop, a
  second run found nothing to do and the container reported itself healthy.
  A start over a marked volume whose version is already the one of the image
  stays unhealthy and says what to do: rebuild the indexes upgrade.log names
  and remove the mark, or restore the backup. An upgrade that failed before
  it recorded the version is run again, as before. The upgrade step of
  build.yml checks that a successful upgrade leaves no mark, that a mark
  over an upgraded volume keeps the container unhealthy until it is removed,
  and that an upgrade failed by a read-only buildinfo leaves the mark.
- docker-test-replication.sh: dj-bf, the background-join volume whose
  bootstrap fails (OpenIdentityPlatform#1115), followed the recovery that OpenIdentityPlatform#1182 now refuses, and
  failed both Docker jobs. It now stays unhealthy after a restart and runs no
  join, and initializes from dj-0 once the mark is removed by hand: the
  initialize would bring the data of the topology, but nothing else the
  bootstrap was to configure. The README says which road such a volume takes.
- run.sh runs the upgrade over a volume that carries .bootstrap-pending as
  well: start-ds refuses an instance of another version, and the setup could
  not be finished by hand under a newer image.
- join.sh: the repair of the replication server lists leaves every peer that
  does not answer for a later start. A combined server added a directory
  server that was down for good when nothing it asked showed another role,
  and the peers every list holds are never asked. A replication server left
  out adds this server to its own lists when it starts.
- join.sh: a replication server publishes rejoining on the ready road until a
  round finds it a member, and ready then, so that no peer enables through it
  before a reset takes it out. Not pending: that would count it among the
  peers that wait for data and let a fresh seed go ahead.
- build.yml: the upgrade step warns and is skipped when the latest release has
  no image on Docker Hub yet - release.yml pushes the images after it
  publishes the release.
- docker-test-replication.sh: dj-nm, which publishes no role, is given a
  replication server; a combined, a replication and a directory volume that
  lost .replication-role are recorded in their role again; setup.sh refuses
  an unknown REPLICATION_ROLE in a container and records nothing; each
  replication server of the OpenIdentityPlatform#534 topology is stopped in turn; the comment of
  roles_hold names the server that joined through dj-d0.
- Installation and Administration Guides (OpenIdentityPlatform#1179, now on master): a restart
  over a volume whose first start failed stays unhealthy, and a setup
  finished by hand goes on once .bootstrap-pending is removed; a restart
  after a failed upgrade stays unhealthy too, and .upgrade-pending is
  removed once the indexes are rebuilt, instead of "do not restart"; the
  replication note names REPLICATION_ROLE and links the standalone
  replication server and directory server sections.
@vharseko
vharseko force-pushed the feature/1178-docker-replication-role branch from 23a3a58 to dd6113e Compare October 9, 2026 15:55
@vharseko

vharseko commented Oct 9, 2026

Copy link
Copy Markdown
Member Author

Thank you for the review. All nine points hold against the code; they are in dd6113e, on master f3c7a3c (which brought #1179).

dj-bf against .bootstrap-pending. Confirmed, and it is my miss: locally I ran the REPLICATION_ROLE block and the new steps, never the whole script, so the #1115 case never met the #1182 branch. I took the #1182 contract, as the issue's Expected states it ("whatever the replication settings") and as the README already described it. The initialize brings the data of the topology, but nothing else a BOOTSTRAP was to configure. dj-bf fails exactly that way: its custom part fails after setup.sh. A replication server never carries INITIALIZE_PENDING, so the bypass would not have covered it anyway. dj-bf now follows your first sketch. It checks the pending log line after the restart, uses stays_unhealthy rather than one look at the status (a single look right after a restart reads starting), and checks that no join ran. It then removes the mark and expects the initialize from dj-0. The README paragraph now says that a background-join volume takes the same road and joins once the mark is gone. The Installation Guide from #1179 says so as well: a restart over a failed first start stays unhealthy, and removing the mark after finishing the setup by hand is the alternative to removing the volume.

Upgrade over a pending volume. Taken as suggested. A failed upgrade is logged and the server is started anyway.

#1185 in this PR. The upgrade counterpart of #1182 is in the same commit. run.sh marks a volume .upgrade-pending before it upgrades it to another version: the major.minor.point of data/config/buildinfo differs from the image's template/config/buildinfo, as BuildVersion.equals compares them. It removes the mark once upgrade succeeded, post-upgrade tasks included. On a later start, a mark over a volume whose version is already the image's can only come from an upgrade that recorded the version and did not complete. That start stays unhealthy and says what to do: rebuild the indexes upgrade.log names and remove the mark, or restore the backup. An upgrade that failed before it recorded the version is run again, as before. The upgrade step of build.yml now also checks three things: a successful upgrade leaves no mark; a mark over an upgraded volume keeps the container unhealthy until it is removed; an upgrade failed by a read-only buildinfo (it fails in changeBuildInfoVersion, after the upgrade tasks) leaves the mark. The "do not restart" advice in To Upgrade a Docker Container is replaced accordingly.

Repair adds an unanswering directory server. Confirmed: the peers that every list holds are never asked, so mixed could not see C's role. I took the deferral for every role and dropped mixed. The cost is what the comment already accepted for a replication server left out. A peer that was down also adds this server to its own lists when it starts, which connects the two already. The README paragraph is updated as well.

Released image not on Docker Hub yet. Confirmed against release.yml (release-docker needs release-maven). Taken as suggested, with a comment on why.

ready before the first membership check on a replication server. I agree with the premise and the fix, with one change: the replication server publishes rejoining, not pending, until a round finds it a member, and ready from joined() then. Both keep peers from enabling through it, since only ready is trusted. But all_others_pending counts pending as "waits for data". A first peer that starts from scratch at the same time could then seed next to the topology that this replication server serves. rejoining keeps any_other_may_hold_data true for it, as ready did. The state comment at the top of join.sh and the README name this. Like you, I have not built a multi-container case for the window itself. The restarts of dj-r0 and dj-r1 in the role block now go through it, and dj-r1 still joins through dj-r0.

dj-nm pin. Taken. runs_replication_server and recorded_role_of moved up with the other helpers, because the pin runs before the role block where they were defined.

legacy_role pin. Taken, for all three roles: dj-1 (combined) loses .replication-role while it is down in the restart case of the early block. dj-r0 (replication) and dj-d1 (directory) lose it in a restart of their own in the role block, and roles_hold reads what they recorded.

setup.sh and an unknown role. Taken, with one change: the container exits once setup.sh fails, because nothing creates data/config. docker exec would then fail and pass the check vacuously. The volume is therefore read through a throwaway container of the image.

One replication server stopped. Taken: dj-r1 and then dj-r0 are stopped in turn, with a change from the other directory server each time.

roles_hold comment. Taken as worded.

Docs. Besides the two Installation Guide procedures above, the replication NOTE of the Administration Guide now names REPLICATION_ROLE and links the standalone directory server and replication server sections. All four Docker chapters render with AsciidoctorJ 2.5.3 -v without warnings.

Testing (local, 1-CPU Docker VM, images built from this branch, with the local timeouts ×3):

  • the whole docker-test-replication.sh on Debian: passed, in 20403 s;
  • the upgrade step, with the Docker image: a restart after a failed post-upgrade task reports healthy, with the upgrade never completed #1185 checks, and the failed-bootstrap step, on Debian and Alpine: passed;
  • the whole script on Alpine: failed at [#1086] Join replication in the background on every start of the Docker image #1115's dj-2, the third server's first join. Its first dsreplication enable with dj-0 ran past the 270 s attempt bound and was cut off (124). Every later round found "the peers register this server but its own cn=admin data does not", and the reset's dsreplication disable exited 8 (ERROR_CONNECTING) each time. reset_replication is master's here, except that it skips hold_writes for a replication server, and the same case passed on Debian. So I read it as the recovery after an enable cut off half-way, which an overloaded host exposed, rather than as something this round changed. I have not proven that, and the output of that disable does not reach the log. CI runs both jobs on faster runners; if dj-2 fails there too, I will take it up in a separate issue.

@vharseko vharseko changed the title [#1178] Give the background join of the Docker image a server role, keep a failed bootstrap unhealthy, and test the upgrade from a released image [#1178] Give the background join of the Docker image a server role, keep a failed bootstrap or upgrade unhealthy, and test the upgrade from a released image Oct 9, 2026
@vharseko
vharseko requested a review from maximthomas October 9, 2026 15:56

@maximthomas maximthomas left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

praise: Round 2 closes every issue of round 1 where it was, and pins the hardest road of the role topology end to end.

  • .github/scripts/docker-test-replication.sh:881-891 stops dj-d0, brings dj-r1 back on an empty volume and requires "joined the replication topology through dj-r0", so a replication server enabling through another replication server (--onlyReplicationServer on both sides) is shown to work, green in both Docker jobs.
  • .github/workflows/build.yml:849-866 pins the #1185 mark from both sides: a clean upgrade leaves no .upgrade-pending, and a mark over an upgraded volume keeps the container unhealthy until it is removed (run.sh:83-97).
  • The round-1 roads are fixed where they were: dj-bf removes .bootstrap-pending (docker-test-replication.sh:435), the .bootstrap-pending start still runs the upgrade (run.sh:181-188), and unanswering peers stay out of the replication server lists (join.sh:891).

issue (non-blocking): The new pre-roles volume cases restart every volume under the REPLICATION_ROLE it was set up with, so they cannot tell legacy_role from an echo of $REPLICATION_ROLE.

.github/scripts/docker-test-replication.sh:860-867, :337-346

dj-1 loses .replication-role at :337 and restarts under the default combined; dj-r0 and dj-d1 lose it at :861, and docker restart keeps replication and directory from :789-791. Each case expects exactly the env's value, and the swap at :869-870 comes after the role is recorded again, so it never reaches legacy_role. The mutant legacy_role() { echo "$REPLICATION_ROLE"; } passes the assert at :345 and both roles_hold calls. The description lists "the role of a volume set up before roles inferred again, for combined, directory and replication volumes" among what the script gains. The property the inference exists for has no observable: a pre-roles directory or combined volume started under another role keeps its configured role. A second mutant survives too: the server) branch always echoing replication.

# after the loop at :866: a pre-roles directory volume under another REPLICATION_ROLE
docker exec dj-d1 rm /opt/opendj/data/.replication-role
docker rm -f dj-d1 >/dev/null
start_node dj-d1 $ROLE_PEERS -e REPLICATION_ROLE=combined -e REPLICATION_RETRY_COUNT=30
wait_until 300 "the pre-roles dj-d1 is a member again" logs_have dj-d1 "already a member of the replication topology"
[ "$(recorded_role_of dj-d1)" = directory ] || fail "dj-d1 was recorded $(recorded_role_of dj-d1) from its directory replication configuration"

Pin: the echo mutant records combined and fails the last line.


issue (non-blocking): dj-bogus and vol-dj-bogus are missing from NODES and VOLUMES.

.github/scripts/docker-test-replication.sh:26-27, :749

Only the success path removes them by hand (:754-755). When the wait at :750 times out or the check at :751 fails, the ERR trap's dump_logs loops over $NODES and never prints the one container log that explains the failure. cleanup also leaves dj-bogus and its volume behind, so on a reused host the next run's docker run --name dj-bogus fails on the name. On ephemeral runners only the missing log matters.

NODES="... dj-nm dj-nmr dj-bogus"
VOLUMES="... vol-dj-bogus"

suggestion (non-blocking): The released-image pull turns any failure into a green skip, not only a tag that is not on Docker Hub yet.

.github/workflows/build.yml:836, :1317

docker pull "$RELEASED" || { ...; exit 0; } ends the step with success before the first upgrade assertion on a rate limit, a registry or network outage, or a renamed repository, and the warning names the release window whatever the cause. The jobs API then reports "Docker test upgrade from a released image" as passed either way. Run 37955316819 did pull: it carries no such warning.

if ! out=$(docker pull "$RELEASED" 2>&1); then
  case "$out" in
    *"manifest unknown"*|*"not found"*) echo "::warning::$RELEASED is not on Docker Hub yet, the upgrade test is skipped"; exit 0 ;;
    *) echo "$out"; exit 1 ;;
  esac
fi

suggestion (non-blocking): Neither join.sh change of this round has a case that fails without it.

opendj-packages/opendj-docker/bootstrap/join.sh:1108, :891

The script checks published_state only on the combined dj-0/dj-1/dj-2. The replication-server restarts (:850-856, :860-865, :869-872, :903-913) check the "already a member" line, health and roles_hold, which reads the role, not the state. No case leaves a registered peer down while another server repairs its lists. Two reverts leave the script green: dropping || [ "$ROLE" = replication ] at :1108, which brings back ready at a replication server's start, and restoring the mixed=no fallback that adds an unanswering peer.

# with dj-r0's first membership round held back, as the dj-held block does for dj-2 (:498-503)
[ "$(published_state dj-r0)" = rejoining ] || fail "dj-r0 publishes itself ready before its membership check"

Pin: the first revert publishes ready and fails this line. For the second, stop a registered peer missing from one list, restart another server, and assert that the list still lacks it.


suggestion (non-blocking): The setup.sh refusal case checks that no role was recorded, not that the refusal comes "before it sets anything up".

.github/scripts/docker-test-replication.sh:748-753

Move the role check in setup.sh after the /opt/opendj/setup call (:92). The mutant then sets the instance up, logs the same refusal and exits before it records a role, so the wait at :750 and the test -e at :751 both pass. That leaves the hazard setup.sh:30-32 names unpinned: a volume set up for a role nobody asked for.

if docker run --rm --entrypoint test -v vol-dj-bogus:/opt/opendj/data "$IMAGE" -e /opt/opendj/data/config/buildinfo; then
  fail "setup.sh set the instance up before it refused REPLICATION_ROLE=bogus"
fi

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug CI docker docs enhancement kubernetes Kubernetes / Helm / OpenShift deployment replication setup setup / upgrade / uninstall tools (quicksetup) and the launcher scripts tests Test suites: fixing, enabling, un-disabling upgrade Upgrading between versions and migrating from other directory servers

Projects

None yet

2 participants