Repository navigation
[#1178] Give the background join of the Docker image a server role, keep a failed bootstrap or upgrade unhealthy, and test the upgrade from a released image - #1184
Conversation
maximthomas
left a comment
There was a problem hiding this comment.
praise: The role is settled up front and kept with the volume, so a later environment change cannot quietly turn a server into something else.
setup.sh:35-46refuses an unknownREPLICATION_ROLEbefore anything is set up, and writes the role to.replication-role.join.sh:1037-1047refuses a.replication-rolethat holds anything else. It infers the role of a pre-roles volume once, withlegacy_role, and records it.
issue (blocking): The unchanged dj-bf case contradicts the new .bootstrap-pending branch, and both build-docker and build-docker-alpine are red at 23a3a58.
.github/scripts/docker-test-replication.sh:403-419, opendj-packages/opendj-docker/run.sh:149-153
dj-bf is a background-join node whose BOOTSTRAP fails. The case restarts it and expects wait_healthy dj-bf and "initializing from dj-0" (:414-417). With #1182, that restart takes the pending branch: no upgrade, no join.sh, no health marker. dj-bf logs "The bootstrap of this volume did not complete…" and the script fails with timed out waiting until dj-bf is healthy (run 37827037318, step "Docker test replication", both jobs). Every step after it was skipped. That covers the new REPLICATION_ROLE block of the same script, "Docker test failed bootstrap" and "Docker test upgrade from a released image", none of which has run in CI. Which contract do you want? Should a background-join volume whose bootstrap failed still recover on restart by initializing from the topology, as #1115's dj-bf pins? Or should it stay unhealthy until someone fixes it by hand, as #1182 now does? Your answer picks the fix below, and either fix needs a green run of both Docker jobs.
# dj-bf under the #1182 contract
docker restart dj-bf >/dev/null
wait_until 300 "dj-bf reports its pending bootstrap" logs_have dj-bf "The bootstrap of this volume did not complete"
[ "$(docker inspect -f '{{.State.Health.Status}}' dj-bf)" != healthy ] || fail "dj-bf turned healthy over a failed bootstrap"
docker exec dj-bf rm /opt/opendj/data/.bootstrap-pending
docker restart dj-bf >/dev/null
wait_healthy dj-bf
logs_have dj-bf "initializing from dj-0" || fail "dj-bf joined with the data of its failed bootstrap"Or keep #1115's recovery: let a volume that still carries INITIALIZE_PENDING bypass the pending branch, so the initialize from the topology replaces whatever the failed bootstrap left (if [ -f "$BOOTSTRAP_PENDING" ] && ! { join_requested && [ -f "$INITIALIZE_PENDING" ]; }; then). Then make the README's .bootstrap-pending paragraph say which road a background-join volume takes.
issue (non-blocking): A start over a .bootstrap-pending volume also skips ./upgrade, so under a newer image start-ds refuses to start and the container crash-loops.
opendj-packages/opendj-docker/run.sh:149-153
The upgrade sits in the elif. A marked volume started under an image of another version therefore reaches BuildVersion.checkVersionMismatch() (DirectoryServer.java:1354, BuildVersion.java:100-104), which throws when the binary and instance versions differ. start-ds exits, PID 1 exits, and the "finish the setup by hand" road that the log line offers cannot be taken under that image. Same-image restarts, which is what the CI step does, are not affected.
if [ -f "$BOOTSTRAP_PENDING" ]; then
sh ./upgrade -n --force || true
echo "The bootstrap of this volume did not complete, this container will not report itself healthy: ..."
elif sh ./upgrade -n --force; thenissue (non-blocking): repair_replication_servers adds a directory server that does not answer to every replication server list of a combined server, when that directory server is the only peer the lists lack.
opendj-packages/opendj-docker/bootstrap/join.sh:879-911
mixed turns yes only when a lacking peer answers without server. Peers that every list already holds are never asked. Take combined A and B and directory C, with C down while A restarts. B is skipped, C fails replication_config, mixed stays no, and C:REPLICATION_PORT is added to A's server list and to its BASE_DN, cn=schema and cn=admin data domain lists. On later starts every list holds C, so C is never asked again, and cleanup_departed keeps it. That is the outcome the comment at :891-896 says this arm prevents. Replication keeps working, but A dials a port that nothing listens on for good. A peer replication server left out is added on a later start, as the comment already accepts, so the unknown hosts can be deferred on every role:
for host in $unknown; do
echo "join: could not ask $host whether it runs a replication server, its place in the replication server lists is checked on a later start"
doneissue (non-blocking): The upgrade step pulls openidentityplatform/opendj:${release_version} from Docker Hub with no fallback, so Docker CI goes red on every PR between a GitHub release and its Docker push.
.github/workflows/build.yml:829-833
release_version is the name of releases/latest (:499-501). release.yml publishes that release in release-maven and pushes the Docker tags only in release-docker, which needs release-maven. In that window, and for as long as a failed release-docker stays unfixed, the pull fails in both jobs on every PR, through nothing the PR did.
docker pull "$RELEASED" || { echo "::warning::$RELEASED is not on Docker Hub yet, the upgrade test is skipped"; exit 0; }issue (non-blocking): On the ready road, a replication server publishes ready before its first membership check. On the pending road, ready means it is already a member.
opendj-packages/opendj-docker/bootstrap/join.sh:1104-1112
The comment's premise, "this volume holds the data of the topology", does not hold for a replication server. Two servers reach this road: an RS that the survivors no longer register, and a legacy srs volume restarted as simple with REPLICATION_ROLE=replication. Either one advertises ready until reset_if_unregistered publishes rejoining in its first round. A pending peer that enables through it in that window may be counted a member through the RS's own registry. Not run: this needs a multi-container run, and I have not traced whether that leaves a lasting island or heals after the RS resets. Publishing pending for ROLE=replication until a round finds it a member would match the pending road.
suggestion (non-blocking): The dj-nm/dj-nmr case does not pin that a peer which publishes no role is taken for combined.
.github/scripts/docker-test-replication.sh:742-747
dj-nm runs no join, so published_role falls through to *) echo combined (join.sh:400-406). Consider a mutant that returns directory for a peer that publishes nothing. It passes --noReplicationServer1 for dj-nm, dj-nmr's own RS carries the initialize, and both assertions stay green while the legacy master is left without a replication server.
runs_replication_server dj-nm || fail "dj-nm, which publishes no role, was not given a replication server"Pin: after :747, kills the mutant above.
suggestion (non-blocking): legacy_role, which infers the role of a volume set up before roles, has no CI observable.
opendj-packages/opendj-docker/bootstrap/join.sh:1018-1048
It runs only when .replication-role is missing. Every replication volume in the script is set up by this image, which writes the file, and the released-image upgrade step sets no OPENDJ_REPLICATION_TYPE. A mutant such as server) always echoing combined stays green. Under that mutant, an srs replication server upgraded with REPLICATION_ROLE=replication is recorded as combined for good.
docker exec dj-r0 rm /opt/opendj/data/.replication-role
docker restart dj-r0 >/dev/null
wait_healthy dj-r0
[ "$(recorded_role_of dj-r0)" = replication ] || fail "dj-r0 was recorded $(recorded_role_of dj-r0) from its replication configuration"Pin: the same on a combined volume, expecting combined.
suggestion (non-blocking): No case exercises setup.sh refusing an unknown REPLICATION_ROLE, which is the road a real container takes.
opendj-packages/opendj-docker/bootstrap/setup.sh:35-46, .github/scripts/docker-test-replication.sh:722-725
The only unknown-role case runs join.sh through --entrypoint, with no volume, no run.sh and no setup.sh. A mutant that drops setup.sh's *) … exit 1 writes bogus to .replication-role, sets the volume up as a data holder, and stays green.
start_node dj-bogus dj-0,dj-bogus -e REPLICATION_ROLE=bogus
wait_until 120 "setup refuses the role" logs_have dj-bogus "REPLICATION_ROLE=bogus is not one of combined, directory or replication"
if docker exec dj-bogus test -e /opt/opendj/data/.replication-role; then fail "setup recorded an unknown role"; fisuggestion (non-blocking): "one replication server is enough" stops dj-r1 without checking that any directory server was connected to it.
.github/scripts/docker-test-replication.sh:868-871
If both directory servers sit on dj-r0, the step passes with no failover at all. Not traced: whether the broker always spreads two equal-weight replication servers one per directory server. Stopping each RS in turn makes the step independent of that choice.
docker start dj-r1 >/dev/null; wait_healthy dj-r1
docker stop dj-r0 >/dev/null
add_ou dj-d0 roles-other-rs
wait_has dj-d1 "ou=roles-other-rs,dc=example,dc=com"nitpick (non-blocking): roles_hold's comment "every server joined through dj-d0 at some point" is false at its third call.
.github/scripts/docker-test-replication.sh:791-792
Before :866, dj-r1 and dj-d1 come back on empty volumes with dj-d0 stopped, and the script asserts that they joined through dj-r0 and dj-r. The check still holds because dj-r0 joined through dj-d0. Suggested wording: "dj-r0 joined through dj-d0: an enable that took dj-d0 for a combined server would have given it a replication server".
…age a server role, keep a failed bootstrap unhealthy, and test the upgrade from a released image OPENDJ_REPLICATION_TYPE=simple made every server a directory server and a replication server at once, and the one-shot srs and sdsr types - the only way to run a replication server without data, or a directory server without a replication server - are deprecated in its favour. REPLICATION_ROLE now sets the part a server of simple takes: combined (the default, as before), directory or replication. - join.sh passes each side of dsreplication enable its own role. The peer's is read from the peer in one search of its cn=config - its replication server, its domain and its backend for BASE_DN - or, for a seed nobody joined yet, from the role it publishes next to its state in cn=Docker Join; a peer that publishes none, an image before roles or a master that runs no join, is combined. The peer is read under its join lock. Without the role flags the tool gives both servers a replication server, so a join through a directory server handed it one (ReplicationCliMain.configureServer). - setup.sh records the role on the volume, and that is the role of the server from then on: a start in another role is logged and changes nothing. A volume set up by an image before roles is read once from its replication configuration - a replication server without a domain is the replication server of srs where REPLICATION_ROLE=replication says so - and recorded. - Membership, the reset of an unregistered server and the peers asked about registration follow the role: a replication server is a member through its replication server, and holds no backend whose writes a reset would hold. - A replication server goes the pending road and is never initialized or a seed: the first peer of REPLICATION_PEERS that holds data seeds, replication servers listed ahead of it passed over by the role they publish. - A directory server does not enable with a directory seed no replication server joined yet, which dsreplication enable refuses, and says why. - The repair of the replication server lists adds only peers that run a replication server, asking each peer a list lacks; one that does not answer goes in as before on a combined server whose answering peers all run one. - setup.sh creates no backend and loads no entries for a replication server of the background join, and refuses an unknown REPLICATION_ROLE before it sets anything up. - An unknown REPLICATION_ROLE ends the join. - BASE_DN goes into the search filters of the join escaped (RFC 4515). README: Server roles, the mapping of srs and sdsr onto simple, and the Kubernetes layout. docker-test-replication.sh: the OpenIdentityPlatform#534 topology - two directory and two replication servers - with roles held across restarts and one replication server stopped, the roles each server publishes and records, a replication server and a directory server joining through a replication server, directory servers with no replication server, an unknown role, a replication server listed first, servers started again in another role, and a replica joining through a master that runs no join. OpenIdentityPlatform#1182: run.sh marks a volume .bootstrap-pending before its bootstrap and removes the mark only once the whole bootstrap - the bootstrap script, and the replicate.sh of a one-shot type - succeeded. A start over a volume that still carries it starts the server to be looked at, runs neither the upgrade nor the join, never writes the health marker, and logs what to do: remove the volume, or finish the setup by hand and remove the mark. Volumes of earlier images carry no mark and start as before. OpenIdentityPlatform#1183: both Docker jobs bootstrap a volume with the latest released image of their flavour, start the build image over it, and check that it upgraded the volume, reports healthy and serves the entries. They also check that a restart over a volume whose create-backend failed stays unhealthy, and that it turns healthy once the setup is finished by hand and the mark removed.
…did not complete unhealthy, the background join's too, and leave unanswering peers out of the replication server lists Review round 1 of OpenIdentityPlatform#1184, and OpenIdentityPlatform#1185. - run.sh marks a volume .upgrade-pending before it upgrades it to another version, and removes the mark once the upgrade, post-upgrade tasks included, succeeded (OpenIdentityPlatform#1185). The upgrade records the new version in buildinfo before those tasks - the index rebuilds and the verification of the DN equality indexes - so after one failed, or was cut off by a stop, a second run found nothing to do and the container reported itself healthy. A start over a marked volume whose version is already the one of the image stays unhealthy and says what to do: rebuild the indexes upgrade.log names and remove the mark, or restore the backup. An upgrade that failed before it recorded the version is run again, as before. The upgrade step of build.yml checks that a successful upgrade leaves no mark, that a mark over an upgraded volume keeps the container unhealthy until it is removed, and that an upgrade failed by a read-only buildinfo leaves the mark. - docker-test-replication.sh: dj-bf, the background-join volume whose bootstrap fails (OpenIdentityPlatform#1115), followed the recovery that OpenIdentityPlatform#1182 now refuses, and failed both Docker jobs. It now stays unhealthy after a restart and runs no join, and initializes from dj-0 once the mark is removed by hand: the initialize would bring the data of the topology, but nothing else the bootstrap was to configure. The README says which road such a volume takes. - run.sh runs the upgrade over a volume that carries .bootstrap-pending as well: start-ds refuses an instance of another version, and the setup could not be finished by hand under a newer image. - join.sh: the repair of the replication server lists leaves every peer that does not answer for a later start. A combined server added a directory server that was down for good when nothing it asked showed another role, and the peers every list holds are never asked. A replication server left out adds this server to its own lists when it starts. - join.sh: a replication server publishes rejoining on the ready road until a round finds it a member, and ready then, so that no peer enables through it before a reset takes it out. Not pending: that would count it among the peers that wait for data and let a fresh seed go ahead. - build.yml: the upgrade step warns and is skipped when the latest release has no image on Docker Hub yet - release.yml pushes the images after it publishes the release. - docker-test-replication.sh: dj-nm, which publishes no role, is given a replication server; a combined, a replication and a directory volume that lost .replication-role are recorded in their role again; setup.sh refuses an unknown REPLICATION_ROLE in a container and records nothing; each replication server of the OpenIdentityPlatform#534 topology is stopped in turn; the comment of roles_hold names the server that joined through dj-d0. - Installation and Administration Guides (OpenIdentityPlatform#1179, now on master): a restart over a volume whose first start failed stays unhealthy, and a setup finished by hand goes on once .bootstrap-pending is removed; a restart after a failed upgrade stays unhealthy too, and .upgrade-pending is removed once the indexes are rebuilt, instead of "do not restart"; the replication note names REPLICATION_ROLE and links the standalone replication server and directory server sections.
23a3a58 to
dd6113e
Compare
|
Thank you for the review. All nine points hold against the code; they are in dd6113e, on master f3c7a3c (which brought #1179). dj-bf against Upgrade over a pending volume. Taken as suggested. A failed upgrade is logged and the server is started anyway. #1185 in this PR. The upgrade counterpart of #1182 is in the same commit. Repair adds an unanswering directory server. Confirmed: the peers that every list holds are never asked, so Released image not on Docker Hub yet. Confirmed against release.yml (
dj-nm pin. Taken.
One replication server stopped. Taken: dj-r1 and then dj-r0 are stopped in turn, with a change from the other directory server each time.
Docs. Besides the two Installation Guide procedures above, the replication NOTE of the Administration Guide now names Testing (local, 1-CPU Docker VM, images built from this branch, with the local timeouts ×3):
|
maximthomas
left a comment
There was a problem hiding this comment.
praise: Round 2 closes every issue of round 1 where it was, and pins the hardest road of the role topology end to end.
.github/scripts/docker-test-replication.sh:881-891stops dj-d0, brings dj-r1 back on an empty volume and requires "joined the replication topology through dj-r0", so a replication server enabling through another replication server (--onlyReplicationServeron both sides) is shown to work, green in both Docker jobs..github/workflows/build.yml:849-866pins the #1185 mark from both sides: a clean upgrade leaves no.upgrade-pending, and a mark over an upgraded volume keeps the container unhealthy until it is removed (run.sh:83-97).- The round-1 roads are fixed where they were: dj-bf removes
.bootstrap-pending(docker-test-replication.sh:435), the.bootstrap-pendingstart still runs the upgrade (run.sh:181-188), and unanswering peers stay out of the replication server lists (join.sh:891).
issue (non-blocking): The new pre-roles volume cases restart every volume under the REPLICATION_ROLE it was set up with, so they cannot tell legacy_role from an echo of $REPLICATION_ROLE.
.github/scripts/docker-test-replication.sh:860-867, :337-346
dj-1 loses .replication-role at :337 and restarts under the default combined; dj-r0 and dj-d1 lose it at :861, and docker restart keeps replication and directory from :789-791. Each case expects exactly the env's value, and the swap at :869-870 comes after the role is recorded again, so it never reaches legacy_role. The mutant legacy_role() { echo "$REPLICATION_ROLE"; } passes the assert at :345 and both roles_hold calls. The description lists "the role of a volume set up before roles inferred again, for combined, directory and replication volumes" among what the script gains. The property the inference exists for has no observable: a pre-roles directory or combined volume started under another role keeps its configured role. A second mutant survives too: the server) branch always echoing replication.
# after the loop at :866: a pre-roles directory volume under another REPLICATION_ROLE
docker exec dj-d1 rm /opt/opendj/data/.replication-role
docker rm -f dj-d1 >/dev/null
start_node dj-d1 $ROLE_PEERS -e REPLICATION_ROLE=combined -e REPLICATION_RETRY_COUNT=30
wait_until 300 "the pre-roles dj-d1 is a member again" logs_have dj-d1 "already a member of the replication topology"
[ "$(recorded_role_of dj-d1)" = directory ] || fail "dj-d1 was recorded $(recorded_role_of dj-d1) from its directory replication configuration"Pin: the echo mutant records combined and fails the last line.
issue (non-blocking): dj-bogus and vol-dj-bogus are missing from NODES and VOLUMES.
.github/scripts/docker-test-replication.sh:26-27, :749
Only the success path removes them by hand (:754-755). When the wait at :750 times out or the check at :751 fails, the ERR trap's dump_logs loops over $NODES and never prints the one container log that explains the failure. cleanup also leaves dj-bogus and its volume behind, so on a reused host the next run's docker run --name dj-bogus fails on the name. On ephemeral runners only the missing log matters.
NODES="... dj-nm dj-nmr dj-bogus"
VOLUMES="... vol-dj-bogus"suggestion (non-blocking): The released-image pull turns any failure into a green skip, not only a tag that is not on Docker Hub yet.
.github/workflows/build.yml:836, :1317
docker pull "$RELEASED" || { ...; exit 0; } ends the step with success before the first upgrade assertion on a rate limit, a registry or network outage, or a renamed repository, and the warning names the release window whatever the cause. The jobs API then reports "Docker test upgrade from a released image" as passed either way. Run 37955316819 did pull: it carries no such warning.
if ! out=$(docker pull "$RELEASED" 2>&1); then
case "$out" in
*"manifest unknown"*|*"not found"*) echo "::warning::$RELEASED is not on Docker Hub yet, the upgrade test is skipped"; exit 0 ;;
*) echo "$out"; exit 1 ;;
esac
fisuggestion (non-blocking): Neither join.sh change of this round has a case that fails without it.
opendj-packages/opendj-docker/bootstrap/join.sh:1108, :891
The script checks published_state only on the combined dj-0/dj-1/dj-2. The replication-server restarts (:850-856, :860-865, :869-872, :903-913) check the "already a member" line, health and roles_hold, which reads the role, not the state. No case leaves a registered peer down while another server repairs its lists. Two reverts leave the script green: dropping || [ "$ROLE" = replication ] at :1108, which brings back ready at a replication server's start, and restoring the mixed=no fallback that adds an unanswering peer.
# with dj-r0's first membership round held back, as the dj-held block does for dj-2 (:498-503)
[ "$(published_state dj-r0)" = rejoining ] || fail "dj-r0 publishes itself ready before its membership check"Pin: the first revert publishes ready and fails this line. For the second, stop a registered peer missing from one list, restart another server, and assert that the list still lacks it.
suggestion (non-blocking): The setup.sh refusal case checks that no role was recorded, not that the refusal comes "before it sets anything up".
.github/scripts/docker-test-replication.sh:748-753
Move the role check in setup.sh after the /opt/opendj/setup call (:92). The mutant then sets the instance up, logs the same refusal and exits before it records a role, so the wait at :750 and the test -e at :751 both pass. That leaves the hazard setup.sh:30-32 names unpinned: a volume set up for a role nobody asked for.
if docker run --rm --entrypoint test -v vol-dj-bogus:/opt/opendj/data "$IMAGE" -e /opt/opendj/data/config/buildinfo; then
fail "setup.sh set the instance up before it refused REPLICATION_ROLE=bogus"
fi
Summary
#1178: a server role for the background join
OPENDJ_REPLICATION_TYPE=simplemade every server a directory server and a replication server at once. The one-shotsrsandsdsrtypes, the only way to run a replication server without data or a directory server without one, are deprecated in its favour.REPLICATION_ROLEnow sets the part a server ofsimpletakes:combined(the default, as before),directoryorreplication.dsreplication enablegets its own role. Without the role flags the tool gives both servers a replication server, so a join through a directory server handed it one (ReplicationCliMain.configureServer).cn=config. That search returns its replication server, its domain and its backend forBASE_DN, plus the role it publishes next to its state incn=Docker Join.combined, as every server was before roles. That covers a peer on an image before roles, or a master that runs no join.setup.shchecksREPLICATION_ROLEbefore it sets anything up and writes it to.replication-rolein the data directory.srsonly whenREPLICATION_ROLE=replicationsays so.replication. On acombinedserver that state is an enable that stopped half-way.rejoininguntil a round finds it a member, so that no peer enables through it before a reset takes it out.REPLICATION_PEERSthat holds data seeds the topology. Replication servers listed ahead of it are passed over, by the role they publish.dsreplication enablerefuses a topology without one, and the log says why.setup.shcreates no backend and loads no entries for a replication server of the background join.BASE_DNgoes into the search filters of the join escaped (RFC 4515).#1182: a failed bootstrap stays unhealthy across restarts
run.shmarks a volume.bootstrap-pendingbefore its bootstrap. It removes the mark only once the whole bootstrap has succeeded: the bootstrap script, and thereplicate.shof a one-shot type.A start over a volume that still carries the mark starts the server so it can be looked at. It runs the upgrade (start-ds refuses an instance of another version) but not the join, never writes the health marker, and logs what to do: remove the volume, or finish the setup by hand and remove the mark. This holds for a volume of the background join as well: an initialize from the topology would bring its data, but nothing else the bootstrap was to configure. Once the mark is removed, such a volume joins and initializes from the topology. Volumes of earlier images carry no mark and start as before.
#1185: a failed upgrade stays unhealthy across restarts
The upgrade records the new version in
buildinfobefore its post-upgrade tasks (the index rebuilds and the verification of the DN equality indexes), so after one of them failed, or was cut off by a stop, the next run ofupgradefound nothing to do and the container reported itself healthy.run.shmarks a volume.upgrade-pendingbefore it upgrades it to another version (major.minor.point ofdata/config/buildinfoagainst the image'stemplate/config/buildinfo, asBuildVersion.equalscompares them), and removes the mark onceupgradesucceeded, post-upgrade tasks included. A mark over a volume whose version is already the image's can only come from an upgrade that recorded the version and did not complete: that start stays unhealthy and logs what to do (rebuild the indexesupgrade.lognames and remove the mark, or restore the backup). An upgrade that failed before it recorded the version is run again, as before.#1183: the upgrade from a released image in CI
Both Docker jobs bootstrap a volume with
openidentityplatform/opendj:${release_version}(-alpineon the Alpine job). They then start the build image over it and check three things: the volume was upgraded (no "has already been upgraded"), the container reports healthy, and it serves the entries.It also covers #1185: a successful upgrade leaves no
.upgrade-pending; a mark over an upgraded volume keeps the container unhealthy until it is removed; an upgrade failed by a read-onlybuildinfo(it fails inchangeBuildInfoVersion, after the upgrade tasks) leaves the mark. The step warns and is skipped when the latest release has no image on Docker Hub yet: release.yml pushes the images after it publishes the release.A second new step covers #1182: a restart over a volume whose
create-backendfailed stays unhealthy, and turns healthy once the setup is finished by hand and the mark removed.Documentation
README.md:srsandsdsrmap ontosimple;REPLICATION_ROLEin the environment table;.bootstrap-pendingand.upgrade-pendingmarks.Installation Guide (To Run OpenDJ in Docker, To Upgrade a Docker Container, from #1179): a restart after a failed first start or a failed upgrade stays unhealthy, and how to get past it by removing the mark, instead of "remove the volume" and "do not restart". Administration Guide, the replication NOTE:
REPLICATION_ROLE, linked to the standalone directory server and replication server sections..github/scripts/docker-test-replication.shgains:.replication-role) inferred again, for combined, directory and replication volumes;setup.shrefusing an unknown role in a container;Testing
Run locally on a 1-CPU Docker VM, with images built from this branch over a
5.2.0-SNAPSHOTbuild of master, and the timeouts of the script tripled:docker-test-replication.shon Debian passed, in 20403 s.dsreplication enableran past the 270 s attempt bound and was cut off, and every later round's reset failed indsreplication disable(exit 8,ERROR_CONNECTING). That code is master's and the case passed on Debian; it reads as the recovery after an enable cut off half-way, exposed by an overloaded host, and is not proven either way.REPLICATION_ROLEblock passed, and failed on the scripts ofmasteras expected (the directory server dj-d0 runs a replication server).bash -eo pipefail), with the same text as in the workflow, on Debian and Alpine: they pass, the upgrade from5.1.2and5.1.2-alpine. failed bootstrap on therun.shofmasterfails: after the restart the container reportshealthywith nouserRoot.-v) without warnings.Not run locally: the rest of the workflow, which CI runs.
Notes
dsreplication enablebetween two replication servers findsBASE_DNonly on a directory server of the topology it can reach. A replication server that joins through another one while every directory server is down waits for one to come back. This is documented in the README.Fixes #1178
Fixes #1182
Fixes #1183
Fixes #1185