Skip to content

[stable/redis-ha] Make master failover reliable when pods are disrupted - #425

Open
mouchar wants to merge 8 commits into
DandyDeveloper:masterfrom
mouchar:redis-ha-failover-fixes
Open

mouchar wants to merge 8 commits into
DandyDeveloper:masterfrom
mouchar:redis-ha-failover-fixes

Conversation

@mouchar

@mouchar mouchar commented Oct 9, 2026

Copy link
Copy Markdown

What this PR does / why we need it:

Our production redis-ha 4.39.0 (3 replicas, quorum 2, on EKS with Karpenter and Istio) lost its master for 43 minutes and could not recover on its own. Karpenter consolidation evicted the master pod. The preStop handoff failed with FailedPreStopHook. The replacement pod could not reach Sentinel and followed announce-0, which was itself a replica, so all three nodes ended up replicating each other with no master. Every failover from then on was aborted with -failover-abort-no-good-slave, and SENTINEL FAILOVER returned NOGOODSLAVE. Deleting pod 0 by hand was the only way out. This happened several times.

We traced it to several defects that compound, and while load-testing the fixes we found more. This PR fixes them, one commit each:

  1. Ask every sentinel through its announce service (0abcfa1). init.sh and fix-split-brain.sh reached Sentinel through the headless service, which lists only ready pods. A replica that has lost its master fails readiness, so once the master was down for about 75 s, no pod could find Sentinel. The scripts now query each Sentinel through its announce-N service, which publishes not-ready addresses, and use the majority answer. Forced failover and sentinel reset go the same way.
  2. Follow only a member that holds the master role (e2296f5). init.sh no longer follows announce-0 blindly. It never accepts its own announce address as the master, and it checks Sentinel again right before forcing a failover, so a healthy new master is not moved a second time. When no member is master, it waits instead of guessing. Pod 0 still promotes itself when no other member answers (a fresh install) or after MASTERLESS_WAIT_ROUNDS rounds of 10 s (default 12), so deleting pod 0 still works as a manual recovery.
  3. Self-heal a masterless set (1064207). This is [redis-ha] Self-heal a masterless set instead of waiting for sentinel #413 by @ShmuelOps, squashed as one commit without its version bump. When the member Sentinel names is not actually master, it repoints the Sentinels at the member that is, or promotes the reachable member with the highest replication offset.
  4. Harden that healer (4e1739a):
    • it also acts when the named master does not answer at all and a quorum of Sentinels flags it down, which is exactly the state from our outage;
    • after SENTINEL MONITOR it restores auth-pass and the sentinel.config values, the gap reported in [redis-ha] Self-heal a masterless set instead of waiting for sentinel #413;
    • it supports announceHostnames, and sends forced failovers through the announce services.
  5. Hand over the master without losing writes (627071f):
    • Shutdown race: Kubernetes stops all containers at once, so the sentinel container exited before the redis preStop hook could ask it to fail over ([stable/redis-ha][BUG] Pre-stop hook unsuccessful #207). A new default sentinel.lifecycle preStop hook keeps Sentinel running while the local redis is master.
    • Missing password: the redis container now receives SENTINELAUTH, so the hook also works with sentinel.auth.
    • Lost writes: the hook pauses writes (CLIENT PAUSE WRITE) during the handoff, follows the promoted node at once, and unpauses. Writes held during the pause fail with READONLY and are retried on the new master instead of being lost. If Sentinel aborts the failover, writes are released at once.
    • Portability: the hook is POSIX sh.
  6. Do not restart a crashed master with stale data (b427c22). When the master's redis container restarts inside its pod (an OOM kill, a crash, a failed liveness probe), it is back before Sentinel notices. It used to come back as master with its last RDB, and every replica resynced from that older data. Redis now starts through redis-start.sh. On a container restart, as opposed to a pod start, while Sentinel still names this node as master, the script has a replica promoted first and starts as its replica.
  7. Keep the master role off terminating replicas (4de1538). When a replica and the master stop together, Sentinel could promote the replica that was shutting down. A terminating replica now sets replica-priority 0. If it is promoted anyway, it hands the role on.
  8. Bump the chart version to 4.40.0 (377de48). The defaults change: the sentinel container gets a preStop hook, and the redis container starts through redis-start.sh.

Which issue this PR fixes

It also addresses #207 and #283, which were closed without a fix. It includes #413 and supersedes #377.

Special notes for your reviewer:

Tests. We ran an automated regression suite on EKS: 10 value sets, 81 scenario runs, plus reruns of the fixed cases.

  • Value sets:
    • chart defaults
    • Redis 7.2.4 with the exporter and a minAvailable: 1 PDB
    • auth + sentinel.auth
    • 5 replicas with quorum 3
    • resolveHostnames/announceHostnames
    • TLS for Redis, Sentinel and replication
    • no persistence
    • podManagementPolicy: Parallel
    • with and without the Istio sidecar
    • HAProxy
  • Scenarios:
    • delete the master; delete a replica; force-delete the master
    • kill redis in the master; hang the master for 30 s (CLIENT PAUSE ALL)
    • kill every Sentinel
    • delete the master and a replica together; delete all pods
    • rolling restart; scale 3→5→3
    • our masterless outage, reproduced with replica-priority 0 on the replicas, then deleting the master

A probe sampled every member's role and every Sentinel's answer once a second. Two clients wrote unique keys during each scenario, one following Sentinel and one pinned to the old master's address, and we counted acknowledged writes missing afterwards.

Results with this PR:

Scenario Before After
Delete master FailedPreStopHook; master replaced after down-after-milliseconds handoff in ~2–5 s, 0 writes lost (65 lost without the write pause)
Masterless outage no recovery (43 min until manual action) self-heals in 175–216 s in every value set
Delete master + replica up to 65 s without a master; could form the loop at most 9 s without a master
Redis crash in the master 371–625 acknowledged writes lost 0 lost
Rolling restart 10 s outage per master restart 2–3 s write gap, 0 lost

Behavior changes:

  • Default sentinel.lifecycle. It now has a preStop hook. A user who overrides sentinel.lifecycle loses it.
  • Redis start command. The redis container runs sh /readonly-config/redis-start.sh, which execs redis-server with the same arguments. Setting redis.customCommand bypasses it.
  • Slower replica shutdown. A replica pod takes up to 10 s longer to stop, while it watches for an unwanted promotion.
  • Healer timing. The healer acts after 3 checks of splitBrainDetection.interval, so about 3 minutes with the default 60 s.
  • New tuning knobs. They are environment variables only, like MAX_QUORUM_FAILURES, so there are no new chart values: MASTERLESS_WAIT_ROUNDS, MAX_STALE_MASTER_FAILURES, MASTERLESS_CONFIRMATIONS, MASTERLESS_CONFIRM_INTERVAL and PROMOTE_SCAN_ATTEMPTS.

Known limitations, not addressed:

  • Scaling down while the master sits on a removed ordinal gives about 25 s without a master. Helm deletes that pod's announce Service before the pod stops. A fix would need a pre-upgrade hook that compares replica counts with lookup, which does not work under Argo CD. Workaround: fail over before scaling down.
  • Deleting all pods at once can lose the last few writes, or everything without persistence.
  • The write gap after losing both replicas comes from the default min-replicas-to-write: 1. The master rejects writes until a replica rejoins.

Checklist

  • DCO signed
  • Chart Version bumped
  • Variables are documented in the README.md
  • Title of the PR starts with chart name (e.g. [stable/mychartname])

mouchar and others added 8 commits October 9, 2026 22:25
init.sh and fix-split-brain.sh reached Sentinel through the headless
service, which lists only ready pods. A replica that loses its master
fails readiness, so once the master was gone long enough no pod could
find Sentinel. init.sh then fell back to following announce-0 even when
announce-0 was itself a replica, and the set ended with no master.

Query each sentinel through its announce service, which publishes
not-ready addresses, and use the most common answer. Send forced
failovers and resets the same way.

Signed-off-by: Robert Moucha <robert.moucha@gooddata.com>
Co-authored-by: Claude <noreply@anthropic.com>
When init.sh could not get a usable master from sentinel, a replica
followed announce-0 without checking it. If announce-0 was itself a
replica, the set formed a replication loop with no master, which
sentinel cannot break: it rejects every candidate with NOGOODSLAVE.

- Follow a member only when it reports role:master.
- Never accept this pod's own announce address as the master; the
  address may still reach the predecessor that is shutting down.
- Look for a master again right before forcing a failover, so a
  healthy new master is not moved a second time.
- With no master anywhere, wait for one. Pod 0 still makes itself
  master when no other member answers, or after MASTERLESS_WAIT_ROUNDS
  rounds of 10s (default 12).

Signed-off-by: Robert Moucha <robert.moucha@gooddata.com>
Co-authored-by: Claude <noreply@anthropic.com>
…inel

When the master pod is replaced behind the same announce address,
sentinel can keep naming that address while it runs as a replica. Every
member agrees with sentinel, so split-brain-fix never acts, and sentinel
rejects every candidate with NOGOODSLAVE.

Poll the announce services for the member that really holds the master
role. Repoint the sentinels at it when one exists. When none does over
MASTERLESS_CONFIRMATIONS checks, promote the reachable member with the
highest replication offset.

Squashed from DandyDeveloper#413
(commits 0a3bcb0 and ac939b0), without its chart version bump.

Signed-off-by: ShmuelOps <shmuel@tennasys.com>
Signed-off-by: Robert Moucha <robert.moucha@gooddata.com>
Co-authored-by: Claude <noreply@anthropic.com>
Harden the masterless healer:

- Also count a named master that does not answer, once a quorum of
  sentinels flags it as down. A replaced master that cannot start, while
  sentinel finds no good replica, otherwise leaves the set without a
  master and the healer never acts.
- After repointing a sentinel, restore what sentinel.conf sets for the
  master: auth-pass and the sentinel.config values. A bare monitor entry
  falls back to sentinel defaults and cannot authenticate to the master.
- Resolve the promoted member through getent_hosts, so announceHostnames
  works.
- Force failovers through the announce services, and reuse the shared
  redis and sentinel client helpers.

Signed-off-by: Robert Moucha <robert.moucha@gooddata.com>
Co-authored-by: Claude <noreply@anthropic.com>
Kubernetes stops all containers of a pod at once. The sentinel container
had no preStop hook and exited on SIGTERM, so the redis container's hook
could not ask it to fail over: it always failed with "Connection
refused" and the master was replaced only after down-after-milliseconds.
Add a default sentinel preStop hook that keeps sentinel running while
the local redis is master. With sentinel.auth, the redis container also
lacked SENTINELAUTH, so the hook could not authenticate; pass it in.

During a forced failover, the old master kept accepting writes until
sentinel demoted it on its next check, about 10s after the promotion.
Those writes were lost. The redis preStop hook now pauses writes
(CLIENT PAUSE WRITE) before it asks for the failover, follows the
promoted node as soon as sentinel names it, and then unpauses. A held
write fails with READONLY, and the client retries it on the new master.
When sentinel aborts the failover, the hook unpauses at once.

The hook is now POSIX sh, like the other scripts.

Signed-off-by: Robert Moucha <robert.moucha@gooddata.com>
Co-authored-by: Claude <noreply@anthropic.com>
When the redis container of the master restarts inside its pod, for
example after an OOM kill or a crash, it is back within seconds. Sentinel
has not noticed the outage, so the node returns as master with the data
it last saved. Every replica then resyncs from that older data, and the
whole set loses the writes made since the last save.

Start redis through redis-start.sh. config-init clears a marker on every
pod start, so the script can tell a container restart from a pod start.
On a container restart, if sentinel still names this node as master, it
asks sentinel to promote a replica and starts as that replica's replica.

Signed-off-by: Robert Moucha <robert.moucha@gooddata.com>
Co-authored-by: Claude <noreply@anthropic.com>
When a replica pod and the master pod stop at the same time, sentinel
can promote the replica that is shutting down. The set then waits for
sentinel to notice the second loss and fail over again.

A terminating replica now sets replica-priority 0, then watches its role
for 10s. If a failover that started before sentinel saw the new priority
promotes it anyway, it hands the master role on with the same handoff
the master uses. The old master sets replica-priority 0 after its own
handoff. The sentinel preStop hook stays up for the replica's watch.

Signed-off-by: Robert Moucha <robert.moucha@gooddata.com>
Co-authored-by: Claude <noreply@anthropic.com>
Minor bump: the sentinel container gets a default preStop hook, and the redis container starts through redis-start.sh.

Signed-off-by: Robert Moucha <robert.moucha@gooddata.com>
Co-authored-by: Claude <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

2 participants