| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
Sorry, something went wrong.
When Redis (or a Redis-compatible store like Dragonfly/KeyDB) restarts,
MutexOwner would immediately crash with {:shutdown, :err_keeping_mutex}.
Since MutexedSupervisor uses :one_for_all strategy, this cascades and
takes down the entire Runtime.Supervisor including all consumers.
The fix:
- MutexOwner now retries up to 5 times with backoff on Redis errors
while in :has_mutex state, giving Redis time to come back
- Resets the error counter on successful reconnection
- Also fixes the GenStateMachine return value (was {:shutdown, reason}
which is invalid - now {:stop, {:shutdown, reason}})
- LiveView pages (index.ex, show.ex) now handle Redis errors gracefully
in metrics loading instead of crashing with MatchError
Fixes sequinstream#2072
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…s errors Instead of giving up after N errors, MutexOwner now retries indefinitely with exponential backoff capped at 1 hour. Redis going down should never crash Sequin — it should degrade gracefully and self-heal when Redis returns. Integration tests use iptables REJECT to simulate a real Redis/Dragonfly redeploy and verify the process survives the outage and recovers. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…ation tests Unit tests (6, <0.1s): verify backoff math, error counter behavior, and state struct defaults without needing Redis. Integration tests (2, ~35s each, tagged :integration): use iptables REJECT to simulate Redis going down, verifying the process survives and recovers. Run with: mix test --include integration Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
Since upstream appears unmaintained (no maintainer activity since Feb 2026), this fix is now merged and released in the maintained fork at https://github.com/triptechtravel/sequin — release v0.14.6-tt2, image ghcr.io/triptechtravel/sequin:v0.14.6-tt2. The fork continues from upstream's final v0.14.6 with production-driven maintenance (security/crash/bug fixes). |
Sorry, something went wrong.
| Back | FazBrowse Home | New Git URL |
Summary
Fixes #2072 — When Redis (or a Redis-compatible store like Dragonfly/KeyDB) restarts, Sequin enters an unrecoverable failure state requiring a full restart. This has bitten us (and others per #2072) multiple times in production.
Root cause: MutexOwner in :has_mutex state calls acquire_mutex which returns :error when Redis is unreachable. The handler immediately returns {:shutdown, :err_keeping_mutex} — an invalid GenStateMachine return that crashes the process with {:bad_return_from_state_function, ...}. Due to MutexedSupervisor's :one_for_all strategy, this cascades and takes down the entire Runtime.Supervisor including all consumers.
Changes:
Context
We self-host Sequin on Railway with Dragonfly (Redis-compatible) as the backing store. Railway periodically auto-updates Dragonfly, which causes a brief restart. Every time this happens, Sequin fails to self-heal and requires a manual restart — we've hit this 3 times now. The :await_mutex state already retries correctly on Redis errors; this PR brings the same resilience to the :has_mutex state, but with exponential backoff so it doesn't spam during extended outages.
Test plan
Note: Integration tests require NET_ADMIN capability (for iptables) and are tagged :integration so they can be excluded from normal test runs: mix test --exclude integration
🤖 Generated with Claude Code