Closes e2b-dev#3605.
Two new histograms:
api.redis_storage.expiration_index.size ({entry})
ZCARD of sandbox:storage:global:expiration sampled once per heal
pass (every 5 min). Use max/p99 to detect unbounded growth caused
by stale ZSET entries that are not being swept.
api.redis_storage.expiration_index.sweep_duration (ms)
Wall-clock duration of each ExpiredItems call (ZRANGEBYSCORE +
MGET pipeline). Rising p99 alongside rising index size confirms
the O(log N + K) cost growth described in e2b-dev#3605.
Collection points:
- heal.go: ZCARD after each healExpirationIndex pass — one extra
Redis round-trip per 5 min, negligible next to forEachSandboxBatch.
- items.go: defer-based timer wrapping the full ExpiredItems body,
recorded on every evictor tick.
Summary
Closes #3605.
Adds two histograms to make the global expiration ZSET observable in production:
Why these two signals
sandbox:storage:global:expiration is a singleton ZSET shared by every sandbox.
With 22,458 members (~3 MB) observed in production on 2026-08-14 and the evictor
calling ExpiredItems at pollInterval = 50 ms, this key is read 20 × N times
per second (N = API allocation count) — yet neither its size trend nor the
per-sweep cost was visible in any metric.
it causes latency spikes; use as a before/after signal when fix(api): prune stale team index entries when ZSET orphans are swept #3567 (orphan SREM
fix) or future sharding work lands.
rising p99 alongside rising size confirms the cost is growing and justifies
remediation.
Implementation
heal.go — one ZCARD call appended after healExpirationIndex completes.
Cost: one Redis round-trip per 5 min, negligible next to the forEachSandboxBatch
SSCAN scan that precedes it. Sampling here avoids adding any load to the 50 ms
evictor hot path.
items.go — defer-based timer wrapping the full ExpiredItems body.
Records on every evictor tick, capturing the end-to-end cost of
ZRANGEBYSCORE + MGET pipeline + orphan cleanup.
Related
/cc @jakubno @dobrac @ValentaTomas