| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
|
Standardize it, and I'd argue for going slightly beyond a bare timing hint, because timing turns out to be only half of what clients need. I've spent the past weeks building exactly this layer as a transparent proxy (mcp-fuse, MIT) and can share what the payload converged to after contact with real servers. The fragmentation you list matches what I ran into. My classifier literally has to regex retry_after out of free text because every server spells it differently. That alone makes the case for a standard shape. What I'd add to the retry-after field, based on implementation experience: A category enum plus a retryable boolean as the minimal contract (transient, rate_limit, timeout, auth, invalid_input, and a few more). A client that understands nothing else still behaves correctly; timing refines it. A defined carrier at each layer where failures surface: error.data for JSON-RPC errors, result._meta for tool results with isError, and a response header for the HTTP transport case, so the signal survives an opaque 5xx body. The timing hint should mean "not before", not "sleep this long". When a server says wait 12s and the host times the call out at 10s, sleeping is the wrong move. What worked in practice: fail fast, remember the earliest retry time, and reject repeat calls cheaply until it passes. Auto-retry needs to be gated on tool annotations. readOnlyHint and idempotentHint already exist in the spec. A client that silently replays an ambiguous failure against a tool with side effects risks double-executing it. Making that gating normative would also give server authors a concrete reason to annotate accurately. Draft schema, JSON Schema file and conformance examples are here: https://github.com/YoadElkayam/mcp-fuse/tree/main/spec. The proxy synthesizes the payload for servers that don't emit it, so I also have data on what a classifier can and can't recover from free text; happy to share if useful. There's a related thread in #2930 about the classification half of the same problem. On the title question: SDK-only conventions are how the ecosystem got three incompatible retry_after fields. Even a small normative payload, or a reserved _meta key with a defined shape, would let opencode and vercel/ai trigger the retry logic they already have. |
|
Strong +1 on standardizing, and @YoadElkayam's point about gating auto-retry on idempotentHint/readOnlyHint is the one I'd build on — because we went and measured how well that gate actually holds today, and the data argues for making it normative rather than advisory. We scanned 671 published MCP servers (27,153 declared tools) looking specifically at retry safety. Two findings relevant here: 1. The annotation that would make a retry-timing hint safe is usually absent. Of the servers doing real writes (80% of the sample), 32% had no visible idempotency guard of any kind. A retryAfter tells a client when it may try again; it says nothing about whether the effect is safe to repeat. If the side-effect annotation isn't populated, a well-behaved client that honors the timing hint still double-executes — it just waits politely first. 2. When the annotation is present, it can be decorative. In at least one server we read closely, a create tool was marked non-idempotent in its metadata, but nothing at runtime consulted that flag before the retry loop fired. So "the annotation exists" and "the annotation gates retry" are separate claims, and a normative spec is what closes the gap between them. The piece I'd add to the payload, beyond category + timing: a reconciliation pointer for the ambiguous case. The hardest failure isn't a clean rate-limit — it's the timeout where the request may have already landed and only the response was lost. retryable + timing don't resolve that; the only safe move is to ask whether it happened before retrying. A field naming the read that answers "did this land?" (a get-by-id, an idempotency-key lookup) turns "I don't know" from a guess into a query. Absent that, the correct client behavior on an ambiguous timeout against a side-effecting tool is to escalate, not retry — and the spec could say so. I sketched this as a per-tool "retry contract" — effect class, caller-key support (+ scope + retention), duplicate-key/different-payload behavior, ambiguous-timeout reconciliation, and whether downstream dedup is relied upon — deliberately as a superset of the existing hints rather than a competitor to them: https://github.com/aurumflux20/fencescan/blob/main/docs/RETRY-CONTRACT.md. It's a draft; I'd rather have holes poked in it than adopt it, and the honesty boundary (marking what a scanner can't verify from code alone) is the part I'm least sure about. Happy to share the scan dataset/method if the "how common is this actually" question is useful to the SEP — the aggregate is public, no server names. |
|
Relevant data point for the normative-vs-advisory question, from trying to answer it a different way. Rather than asking servers, I tried to infer retry-safety from source: does a retry construct wrap a payment call with no idempotency identity in scope? Built it as a public tool: agent-money-test (MIT, free, no signup). It turned out to be a good argument for a normative field, not against one. Even with careful static analysis — matching only real retry constructs, stripping comments and string literals, scoping identity checks per-file — hand-verifying the results still found the tool wrong about a third of what it flagged: a frontend component displaying payment status read as a payment call, a status-poll loop read as a payment retry, a tutorial's "solution" file counted as a live service. If inferring retry-safety from source is this unreliable even when done carefully, that is the case for the server declaring it directly rather than leaving callers to guess. Happy to be wrong — the code and the false-positive writeup are both public, and the harness self-tests against a known-broken and a known-safe target before it trusts its own verdict on anything else. |
|
That one-in-three false positive rate from careful static analysis is the number the motivation section needs. It also matches the view from the runtime side: sitting between host and server, the proxy can observe failures but cannot infer whether an effect is safe to repeat, so the only reliable input it has is what the tool declares. Inference and observation both point at declaration. One update on process. devmaha's idempotency SEP (#3182) was closed last week under the repo's AI contribution policy, so I don't think waiting to fold into that draft is the right plan anymore. I'd rather we start the companion as a human-written draft, keep it small, put the implementation and the scan data behind it, and open it with full disclosure. I set up a shared working draft with the section split we agreed on and TODO markers per owner: https://github.com/YoadElkayam/mcp-fuse/tree/main/sep. PRs welcome from anyone here, it moves to the spec repo's seps/ directory once it has a sponsor. @aurumflux20 @johnyzaguirre-glean your sections are marked. @devmaha your retryAfter work is welcome in it if you want to bring it over, you're listed as an invited co-author. The SEP guidelines say to bring a proposal to a relevant interest or working group before a cold submission. The Interceptors WG looks like the natural home for this one, so I'll raise it there next. For the record and per AI_POLICY.md: I use Claude Code to help draft my comments and code, including this one. The design decisions, the proxy implementation and its testing are mine, and I stand behind all of it. |
|
Works for me on all three counts. A small human-written companion with the implementations and data behind it is the honest shape anyway — the argument for this SEP was always the evidence, not the prose. I've read the draft and my sections are the right cut: 3.3 and 4.3 are the two halves of one argument (inference cannot replace declaration; therefore the tool must declare, and the client must consult). I'll PR them this week in that order:
On the _meta-dropping open question in 4.2: I'd vote normative preservation requirement rather than a dedicated field — a new field just moves the problem, and SDKs that drop _meta today are already out of spec in spirit. Interceptors WG as the venue sounds right. I'll have my sections in the working draft before that conversation so it lands with the evidence attached. Per AI_POLICY.md: I use Claude Code to help draft comments and code, including this one. The scan, the datasets, the RETRY-CONTRACT design and the implementations behind my sections are mine, and I stand behind all of it. |
|
Verdict space in 4.3, and I will keep 4.1's reconcile as a bare pointer with a cross-reference to it. Your framing is the right one: the pointer says what to run, 4.3 says how to read the answer, and "could not determine is terminal" is the sentence that closes the gate. Agree on _meta preservation over a new field; I wrote it up that way in the draft but marked it as a question for the sponsor, since a preservation rule lands on SDK maintainers and they should see it coming. Looking forward to the PRs. (drafted with Claude Code, same disclosure as above) |
|
Small follow-up on the conformance angle, since the SEP will eventually need one: we packaged the failure-mode battery we'd been running by hand as a runnable tool — a facilitator that deliberately produces each ambiguous outcome (settle-then-timeout, 5xx-after-settle, double-402, slow-answer) and counts how many distinct payments a client actually made for one purchase. One command against any client that reads a facilitator URL from an env var; a correct client settles once, a client that mints a fresh authorization on the ambiguous retry gets caught paying twice. github.com/aurumflux20/hostile-facilitator (MIT, stdlib-only, no keys/chain) Two ways it might be useful here, no strings: as an executable illustration of the retry-timing hint's motivation (each battery mode is one of the cases the hint would let a client handle correctly), and — if the SEP lands with the declaration + verdict-space shape — as a starting point for its conformance suite, since "does the client honor the declaration under each ambiguous outcome" is exactly what it measures. Happy to align its modes/verdicts to whatever the draft finalizes, or for it to be superseded by something official. |
|
Ported it. The MCP-shaped version is a hostile server rather than a facilitator: a charge tool that records its effect before the failure mode fires, verify_charge as the reconciliation read (ground truth), and a runner that drives a client through settle-then-timeout, 5xx-after-settle, duplicate, slow-answer, plus an honest control. It lives in the working draft under sep/conformance with the fixture in examples/hostile-server, your battery credited as the origin: https://github.com/YoadElkayam/mcp-fuse/tree/main/sep/conformance It earned its place on the first run. mcp-fuse double-executed in 2 of 5 modes: the gate treated a 503 and a 429 as "guaranteed not processed" and replayed them after the effect had landed. I tightened it to connection-level phrases only, and it failed again, because the 503 body quoted Envoy's "reset before headers". Both versions were inference from response text. The fix is structural: a tool not declared safe to replay is retried only when the request provably never left the client. All five modes pass now, shipped as mcp-fuse 0.2.0, and the whole sequence is written up in the conformance README because it is the SEP's argument demonstrated on our own code, twice. Two open items noted there: a mode where the reconciliation read itself fails (to assert "could not determine" is terminal, per your 4.3), and a declared-safe mode to make sure the gate is not over-cautious either. (drafted with Claude Code, same disclosure as above) |
|
The port catching mcp-fuse's own gate twice is the whole argument in miniature: a 503 or 429 body is not a settlement outcome, and any inference from response text eventually reads "landed" as "never left." The structural rule — retry only when the request provably never left the client — is the right closure, and that it took two passes on code written by someone who understands the problem is exactly why this wants to be normative, not advisory. Both open items are load-bearing; I'd argue neither is optional:
I'll take my marked sections — the effect declaration and the reconciliation pointer, scan as motivation — and PR against sep/ this week. Agreed it's one proposal; the declaration ("may I ever replay this") and the runtime payload ("what do I do right now") are two reads of one seam, and neither client is safe with only one. |
|
Field data for §4.3, since the room is back: https://aurumflux.co/state-of-retry-safety/ — ten agent-payment money paths read this week, seven treat an ambiguous settlement as "did not happen" and pay twice; the three that don't all implement §4.3's rule literally (unknown is terminal; re-present, never re-mint). Four now carry red/green tests on their own issues and one is reproduced on chain. Posting it here because it's the strongest evidence I have that the four-verdict space isn't theoretical — "applied more than once" is the state you discover when the retry policy has been double-charging. |
|
Reading §4.3's verdict space and the "could not determine is terminal" rule — that's the right The reconciliation read can fail in a way that is indistinguishable from a negative answer, and
Neither is exotic; both are the default idiom in the language they're written in. The pattern So the structural rule I'd propose alongside "never infer settlement from response text": For the failure-mode battery: alongside the hostile server that fires each ambiguous outcome, I'd The positive-control discipline that eventually caught mine, if useful as a phrasing for the |
|
That pipefail failure is a beautiful catch, the checker succeeded and the plumbing reported it as absent. Folded in: 4.3 now carries the rule (read-failure stays distinct from authoritative absence through the client's own code, and success-shaped non-answers like soft-404s map to could not determine, never absent), and your positive-control line went in nearly verbatim, a checker validated only against the negative is indistinguishable from a constant. The three endpoint-misbehavior cases (oversized body, soft-404, truncation) are in the battery's open items next to verify-unavailable, which we now note only covers the clean-failure half. Credited to you in both places: https://github.com/YoadElkayam/mcp-fuse/tree/main/sep (drafted with Claude Code, same disclosure as above) |
|
Correction carried, verb and all: the draft does not quote the number today, and if it ever does it will say "seven of the ten read can charge twice", capability not incidence. Both battery principles are folded in as normative conformance text in 4.3, credited: the fixture serves the reconciliation read raw (a fixture handing over clean verdicts is testing itself), and a run that never attempts the read scores not_exercised, reported and never counted as a pass. And we took our own medicine: our runner now shows mcp-fuse as NOT-EXERCISED on verify-unavailable, because it holds without ever attempting reconciliation. That row was PASS last week. It was a false green of exactly the kind you describe, and implementing the reconcile read in the proxy is now the roadmap item that turns it green honestly. Your slow-read and stale-replica cases are recorded as open modes too; the stale one probably does need read-your-writes language before it is testable, agreed on leaving it rather than guessing. On the disclosure: welcomed, and it changes nothing here. This thread has judged work by whether it can be verified, and yours has been, every time. (drafted with Claude Code, same disclosure as above) |
|
Closing the loop on the not_exercised row: mcp-fuse 0.3.0 now implements the reconciliation read, which makes it, as far as I know, the first client-side implementation of 4.3's verdict space. A tool declares which read answers "did this land" (proxy config or the tool's own _meta), the proxy runs it only if that read is declared read-only, and acts on the four verdicts: found once returns a verified-complete result to the model instead of a dead-end warning, authoritatively absent replays exactly once, found multiple surfaces the divergence for manual review, and could not determine holds, never collapsing into absent. A failed or unparseable read is could not determine, per the plumbing rule. Battery after: 9 pass, 0 fail, 0 not exercised. The four settled-but-response-lost modes now resolve as "Verified: the operation already completed exactly once", which is the first time the battery has shown an agent receiving the truth about a lost response rather than advice about uncertainty. not_exercised stays in the runner as a regression guard. Draft and battery: https://github.com/YoadElkayam/mcp-fuse/tree/main/sep (drafted with Claude Code, same disclosure as above) |
| Back | FazBrowse Home | New Git URL |
Uh oh!
There was an error while loading. Please reload this page.
{{title}}
Uh oh!
There was an error while loading. Please reload this page.
MCP has no standard way for a server to tell a client "this failure is retryable, and here's roughly how long to wait" — error.data is intentionally application-defined, and JSON-RPC gives it no shape for this. Three independently-built MCP servers already fill that gap today, each incompatibly:
Docdex: custom error code -32029, fields retry_after_ms/retry_at
slack-mcp: a dedicated exception class, field retry_after
mcp-time-server-node: code -32000, field retryAfter (seconds)
That's not a hypothetical concern — it has a real, documented cost. opencode (a widely-used open-source MCP client) has an issue where it aborts an entire in-progress task the moment a tool returns a rate-limit error, despite already having working retry logic for LLM-provider rate limits elsewhere in the same client — the logic exists, there's just nothing standard to trigger it on. vercel/ai and awslabs/mcp both have independent issues describing the same underlying gap with different symptoms (a hallucinated retry loop; a server wrongly marked "unavailable").
For what it's worth as outside precedent: Cloudflare's production edge network standardizes exactly this shape of signal — a retry-timing value kept in sync between an HTTP header and a JSON body field — across seven different error causes, only one of which is a rate limit. Same design choice, at real scale, for the same reason.
I've worked this into a full SEP draft with a minimal design (one field, error.data.retryAfter, an integer/null/absent) and a working reference implementation, but before opening that as a PR I want to check the more basic question: does this belong in the protocol, or is it better left as something each SDK exposes on its own? Happy to share the fuller writeup, FAQ, or reference implementation if useful context.