If you DELETE a paused sandbox while a resume for it is still running, the
delete returns 204 and drops the snapshot, but the resume then publishes
the sandbox back as running. The client is told it is gone while it keeps
running until timeout, and the snapshot it would resume from is gone too.
The reason is that a paused sandbox only has a snapshot, no running-store
record. StartRemoving has nothing to lock or pin, so the kill handler just
deletes the snapshot without recording any intent. Resume publishes through
storage.Add, which is a plain SET+SADD with no check. Nothing serializes
the two.
Rather than add a separate tombstone key, I reused the reservation a resume
already holds the whole time it runs. Before touching the snapshot the kill
handler calls ClaimKill, which looks at the pending set and the storage
index in one script: if a resume is in flight (or already finished and back
in the index) it bails out and the handler returns 409, so the caller
retries the kill against the running sandbox. Otherwise it writes a
short-lived claim that reserveScript rejects, so any resume that starts
after we commit to the delete loses. The claim only has to outlive the
snapshot soft-delete becoming durable - after that a resume fails when it
fetches the snapshot - and it is released early if the delete fails.
This is the same idea as the ExpectExecutionID pin we already use for
running sandboxes: make the write that could bring a removed sandbox back
check, atomically, that no kill was accepted first. Returning 409 also
matches what resume already does when a sandbox is snapshotting.
Fixes e2b-dev#3636
Problem
If a client DELETEs a paused sandbox while a resume for the same sandbox is still in flight, the delete returns 204 and soft-deletes the snapshot, but the resume then publishes the sandbox back as running. The client is told the sandbox is gone while it keeps running until its timeout (and is billable), and the snapshot it would have resumed from has been deleted, so a later pause/resume of that sandbox ends up unrecoverable.
Reproduced deterministically on a dev cluster: create + pause, fire an async resume, then DELETE ~20-60 ms later. DELETE returns 204, resume returns 201, and GET then reports state=running.
Fixes #3636.
Root cause
A paused sandbox has no running-store record — only a snapshot row. So on DELETE:
The two paths share no lock, so they interleave: the delete can remove the snapshot and return 204 in the same window the resume is restoring the node and about to Add.
This is the same class of stale-write race that RemoveOpts.ExpectExecutionID already fences for running sandboxes (startTransitionScript refuses to overwrite a newer incarnation). The paused case is unprotected precisely because there is no record and no execution ID to pin against.
Fix
Rather than introduce a separate tombstone key with its own TTL and GC, this reuses the reservation a resume already holds for its entire lifecycle (Reserve .. finishStart) as the rendezvous point.
Before touching the snapshot, the kill handler calls ClaimKill, a Lua script that checks the pending set and the storage index atomically:
The claim only has to outlive the snapshot soft-delete becoming durable (after that, a resume fails when it fetches the snapshot), so its TTL is short, and it is released eagerly if the delete fails. Net effect: an accepted kill is irreversible — no concurrent resume can publish after it.
Why this shape
Testing