Repository navigation
Conversation
init.sh and fix-split-brain.sh reached Sentinel through the headless service, which lists only ready pods. A replica that loses its master fails readiness, so once the master was gone long enough no pod could find Sentinel. init.sh then fell back to following announce-0 even when announce-0 was itself a replica, and the set ended with no master. Query each sentinel through its announce service, which publishes not-ready addresses, and use the most common answer. Send forced failovers and resets the same way. Signed-off-by: Robert Moucha <robert.moucha@gooddata.com> Co-authored-by: Claude <noreply@anthropic.com>
When init.sh could not get a usable master from sentinel, a replica followed announce-0 without checking it. If announce-0 was itself a replica, the set formed a replication loop with no master, which sentinel cannot break: it rejects every candidate with NOGOODSLAVE. - Follow a member only when it reports role:master. - Never accept this pod's own announce address as the master; the address may still reach the predecessor that is shutting down. - Look for a master again right before forcing a failover, so a healthy new master is not moved a second time. - With no master anywhere, wait for one. Pod 0 still makes itself master when no other member answers, or after MASTERLESS_WAIT_ROUNDS rounds of 10s (default 12). Signed-off-by: Robert Moucha <robert.moucha@gooddata.com> Co-authored-by: Claude <noreply@anthropic.com>
…inel When the master pod is replaced behind the same announce address, sentinel can keep naming that address while it runs as a replica. Every member agrees with sentinel, so split-brain-fix never acts, and sentinel rejects every candidate with NOGOODSLAVE. Poll the announce services for the member that really holds the master role. Repoint the sentinels at it when one exists. When none does over MASTERLESS_CONFIRMATIONS checks, promote the reachable member with the highest replication offset. Squashed from DandyDeveloper#413 (commits 0a3bcb0 and ac939b0), without its chart version bump. Signed-off-by: ShmuelOps <shmuel@tennasys.com> Signed-off-by: Robert Moucha <robert.moucha@gooddata.com> Co-authored-by: Claude <noreply@anthropic.com>
Harden the masterless healer: - Also count a named master that does not answer, once a quorum of sentinels flags it as down. A replaced master that cannot start, while sentinel finds no good replica, otherwise leaves the set without a master and the healer never acts. - After repointing a sentinel, restore what sentinel.conf sets for the master: auth-pass and the sentinel.config values. A bare monitor entry falls back to sentinel defaults and cannot authenticate to the master. - Resolve the promoted member through getent_hosts, so announceHostnames works. - Force failovers through the announce services, and reuse the shared redis and sentinel client helpers. Signed-off-by: Robert Moucha <robert.moucha@gooddata.com> Co-authored-by: Claude <noreply@anthropic.com>
Kubernetes stops all containers of a pod at once. The sentinel container had no preStop hook and exited on SIGTERM, so the redis container's hook could not ask it to fail over: it always failed with "Connection refused" and the master was replaced only after down-after-milliseconds. Add a default sentinel preStop hook that keeps sentinel running while the local redis is master. With sentinel.auth, the redis container also lacked SENTINELAUTH, so the hook could not authenticate; pass it in. During a forced failover, the old master kept accepting writes until sentinel demoted it on its next check, about 10s after the promotion. Those writes were lost. The redis preStop hook now pauses writes (CLIENT PAUSE WRITE) before it asks for the failover, follows the promoted node as soon as sentinel names it, and then unpauses. A held write fails with READONLY, and the client retries it on the new master. When sentinel aborts the failover, the hook unpauses at once. The hook is now POSIX sh, like the other scripts. Signed-off-by: Robert Moucha <robert.moucha@gooddata.com> Co-authored-by: Claude <noreply@anthropic.com>
When the redis container of the master restarts inside its pod, for example after an OOM kill or a crash, it is back within seconds. Sentinel has not noticed the outage, so the node returns as master with the data it last saved. Every replica then resyncs from that older data, and the whole set loses the writes made since the last save. Start redis through redis-start.sh. config-init clears a marker on every pod start, so the script can tell a container restart from a pod start. On a container restart, if sentinel still names this node as master, it asks sentinel to promote a replica and starts as that replica's replica. Signed-off-by: Robert Moucha <robert.moucha@gooddata.com> Co-authored-by: Claude <noreply@anthropic.com>
When a replica pod and the master pod stop at the same time, sentinel can promote the replica that is shutting down. The set then waits for sentinel to notice the second loss and fail over again. A terminating replica now sets replica-priority 0, then watches its role for 10s. If a failover that started before sentinel saw the new priority promotes it anyway, it hands the master role on with the same handoff the master uses. The old master sets replica-priority 0 after its own handoff. The sentinel preStop hook stays up for the replica's watch. Signed-off-by: Robert Moucha <robert.moucha@gooddata.com> Co-authored-by: Claude <noreply@anthropic.com>
Minor bump: the sentinel container gets a default preStop hook, and the redis container starts through redis-start.sh. Signed-off-by: Robert Moucha <robert.moucha@gooddata.com> Co-authored-by: Claude <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this PR does / why we need it:
Our production redis-ha 4.39.0 (3 replicas, quorum 2, on EKS with Karpenter and Istio) lost its master for 43 minutes and could not recover on its own. Karpenter consolidation evicted the master pod. The preStop handoff failed with
FailedPreStopHook. The replacement pod could not reach Sentinel and followedannounce-0, which was itself a replica, so all three nodes ended up replicating each other with no master. Every failover from then on was aborted with-failover-abort-no-good-slave, andSENTINEL FAILOVERreturnedNOGOODSLAVE. Deleting pod 0 by hand was the only way out. This happened several times.We traced it to several defects that compound, and while load-testing the fixes we found more. This PR fixes them, one commit each:
0abcfa1).init.shandfix-split-brain.shreached Sentinel through the headless service, which lists only ready pods. A replica that has lost its master fails readiness, so once the master was down for about 75 s, no pod could find Sentinel. The scripts now query each Sentinel through itsannounce-Nservice, which publishes not-ready addresses, and use the majority answer. Forced failover andsentinel resetgo the same way.e2296f5).init.shno longer followsannounce-0blindly. It never accepts its own announce address as the master, and it checks Sentinel again right before forcing a failover, so a healthy new master is not moved a second time. When no member is master, it waits instead of guessing. Pod 0 still promotes itself when no other member answers (a fresh install) or afterMASTERLESS_WAIT_ROUNDSrounds of 10 s (default 12), so deleting pod 0 still works as a manual recovery.1064207). This is [redis-ha] Self-heal a masterless set instead of waiting for sentinel #413 by @ShmuelOps, squashed as one commit without its version bump. When the member Sentinel names is not actually master, it repoints the Sentinels at the member that is, or promotes the reachable member with the highest replication offset.4e1739a):SENTINEL MONITORit restoresauth-passand thesentinel.configvalues, the gap reported in [redis-ha] Self-heal a masterless set instead of waiting for sentinel #413;announceHostnames, and sends forced failovers through the announce services.627071f):sentinel.lifecyclepreStop hook keeps Sentinel running while the local redis is master.SENTINELAUTH, so the hook also works withsentinel.auth.CLIENT PAUSE WRITE) during the handoff, follows the promoted node at once, and unpauses. Writes held during the pause fail withREADONLYand are retried on the new master instead of being lost. If Sentinel aborts the failover, writes are released at once.sh.b427c22). When the master's redis container restarts inside its pod (an OOM kill, a crash, a failed liveness probe), it is back before Sentinel notices. It used to come back as master with its last RDB, and every replica resynced from that older data. Redis now starts throughredis-start.sh. On a container restart, as opposed to a pod start, while Sentinel still names this node as master, the script has a replica promoted first and starts as its replica.4de1538). When a replica and the master stop together, Sentinel could promote the replica that was shutting down. A terminating replica now setsreplica-priority 0. If it is promoted anyway, it hands the role on.377de48). The defaults change: the sentinel container gets a preStop hook, and the redis container starts throughredis-start.sh.Which issue this PR fixes
It also addresses #207 and #283, which were closed without a fix. It includes #413 and supersedes #377.
Special notes for your reviewer:
Tests. We ran an automated regression suite on EKS: 10 value sets, 81 scenario runs, plus reruns of the fixed cases.
minAvailable: 1PDBauth+sentinel.authresolveHostnames/announceHostnamespodManagementPolicy: ParallelCLIENT PAUSE ALL)replica-priority 0on the replicas, then deleting the masterA probe sampled every member's role and every Sentinel's answer once a second. Two clients wrote unique keys during each scenario, one following Sentinel and one pinned to the old master's address, and we counted acknowledged writes missing afterwards.
Results with this PR:
FailedPreStopHook; master replaced afterdown-after-millisecondsBehavior changes:
sentinel.lifecycle. It now has a preStop hook. A user who overridessentinel.lifecycleloses it.sh /readonly-config/redis-start.sh, whichexecsredis-serverwith the same arguments. Settingredis.customCommandbypasses it.splitBrainDetection.interval, so about 3 minutes with the default 60 s.MAX_QUORUM_FAILURES, so there are no new chart values:MASTERLESS_WAIT_ROUNDS,MAX_STALE_MASTER_FAILURES,MASTERLESS_CONFIRMATIONS,MASTERLESS_CONFIRM_INTERVALandPROMOTE_SCAN_ATTEMPTS.Known limitations, not addressed:
lookup, which does not work under Argo CD. Workaround: fail over before scaling down.min-replicas-to-write: 1. The master rejects writes until a replica rejoins.Checklist
[stable/mychartname])