…ed eid key
The executionID check in startTransitionScript previously decoded the full
sandbox JSON blob on every pinned removal call. Redis executes all Lua on its
single-threaded event loop, so cjson.decode (O(N) in JSON size) blocked all
other commands for the duration.
Introduce a dedicated execution-ID key (sandbox:storage:{team}:sandboxes:{id}:eid)
that stores only the executionID string. All three Lua scripts are updated to
write and delete it atomically alongside the main sandbox key. The hot-path
check in startTransitionScript now does a single GET + string compare (O(1))
instead of the full JSON decode.
A migration fallback is retained: when the eid key is absent (sandboxes created
before this deploy) the script falls back to cjson.decode so existing sandboxes
are still handled correctly. The fallback becomes unreachable once all pre-deploy
sandboxes have expired and can be removed in a follow-up.
Fixes: e2b-dev#3602
Problem
Closes #3602.
startTransitionScript checks the expected executionID by decoding the full
sandbox JSON blob with cjson.decode. Because Redis executes all Lua on a
single-threaded event loop, this O(N) decode blocks every other Redis command
for its duration. Under production load (thousands of concurrent sandboxes) this
creates measurable latency spikes on all storage operations.
Solution
Introduce a dedicated execution-ID key per sandbox:
sandbox:storage:{teamID}:sandboxes:{sandboxID}:eidThe key uses the same {teamID} hash tag as the main sandbox key, so it lands
in the same cluster slot and can be accessed atomically within Lua scripts.
Changes
Hot path before vs after
Before — every StartRemoving call with ExpectExecutionID set:
After — O(1) string compare:
Migration safety
The eid key does not exist for sandboxes created before this deploy. The script
detects this case and falls back to cjson.decode so pre-deploy sandboxes
continue to work correctly. The fallback becomes unreachable once all pre-deploy
sandboxes have expired and can be removed in a follow-up cleanup.
Test plan