[Provider API] Extreme TTFT variance on long-context requests (identical payloads differ 10.6x within one minute) — deepseek/deepseek-v4.1-flash
Date: 2026-10-07, 09:16–09:18 UTC (inside DeepSeek off-peak window)
Endpoint: POST https://api.commandcode.ai/provider/v1/chat/completions (streaming)
Model: deepseek/deepseek-v4.1-flash, temperature=0, GOAT plan
Client network: direct public internet via Cloudflare
Scale of testing: ~120 requests total. Short-context (≤2K tokens) success rate 100%. The problem is specific to long context.
Summary
Time-to-first-token on ~136K-token prompts varies randomly by more than 10x between requests of identical size sent one minute apart. Agent workloads with large context die to client/gateway timeouts when a request lands on a slow upstream channel. This matches and extends #978.
Decisive evidence (trace_ids provided for server-side lookup)
Exp 1 — identical payload resent 13 seconds apart
| # |
UTC |
TTFT |
prompt_tokens |
cached_tokens |
x-trace-id |
| E1 |
09:16:27 |
76.71s |
135,913 |
0 |
cc30adf37c224ce0f2471ddc7beded15 |
| E2 |
09:16:40 |
13.05s |
135,913 |
135,808 |
00acc98c75d1f6a70ed1af13aa93f1b2 |
Exp 2 — same-size cold-cache requests ~1 minute apart (decisive)
| # |
UTC |
TTFT |
derived fresh-prefill throughput |
x-trace-id |
| E3 |
09:18:12 |
91.83s |
≈1,770 tok/s |
5bcbd6781535818ee2613f2463fbba61 |
| E4 |
09:18:20 |
8.63s |
≈15,750 tok/s |
824cd8ce8e50c236ea9a5fd536209e29 |
Both are ~135.9K tokens, cold cache (cached=0), identical structure, differing in one word near the head. 8.9x throughput difference between upstream channels within the same minute.
Also observed (same test session)
- Qwen/Qwen3.8-Max: TTFT 63.42s on a ~2K-token request — c775fb6ec03fbfd0c0c1d490226e2e86
- moonshotai/Kimi-K3: TTFT 34.92s on ~2K tokens — 579ae80dbf513da01600a9bb4745462f
- Cache-hit read speed itself varies 2.5K–5.7K tok/s across fully-cached runs
- A 400K-token request containing a 240K prefix sent seconds earlier came back with cached_tokens=0 — prompt cache does not reliably propagate across requests/channels
- Earlier ladder: 240K tokens TTFT 77.36s vs 400K tokens TTFT 9.95s (smaller request slower) — 217e1626... / 1379cb99...
Ruled out
- Client concurrency: 16 parallel requests → 16/16 HTTP 200, no 429
- Malformed requests: all short requests succeed; all long requests return correct output (problem is latency only)
- Peak-hour throttling: tests ran inside the DeepSeek off-peak discount window
Related existing issues
#978 (TTFT instability — same mechanism), #991 (502 on large sessions), #976 (tool-heavy long requests 400), #989 (non-streaming 524), #955 (422 above ~256K)
Secondary defects found along the way
- meta/muse-spark-1.2 and muse-spark-1.2-contributor (listed in GOAT) persistently return 403 "Authentication failed. Please check your credentials." — traces c8b4b62bc99b02f5b032e58de5a86a3a, 31badfb8cb23444665e9b9319fbc724c
- google/gemini-3.8-flash self-identifies as "Gemini 3.7 Flash" in identity probes
- deepseek/deepseek-v4-pro self-identified as "ChatGPT (GPT-4o) by OpenAI" in an identity probe
Requested
- Investigate upstream channel pool prefill throttling and prompt-cache account affinity (why cache doesn't hit across channels)
- Publish or improve P95/P99 TTFT for long-context requests
- A public status page would prevent most of these reports
Repro: any ~136K-token plain-text prompt, streaming, temperature 0. Full JSON results and generation script available on request.
[Provider API] Extreme TTFT variance on long-context requests (identical payloads differ 10.6x within one minute) — deepseek/deepseek-v4.1-flash
Date: 2026-10-07, 09:16–09:18 UTC (inside DeepSeek off-peak window)
Endpoint: POST https://api.commandcode.ai/provider/v1/chat/completions (streaming)
Model: deepseek/deepseek-v4.1-flash, temperature=0, GOAT plan
Client network: direct public internet via Cloudflare
Scale of testing: ~120 requests total. Short-context (≤2K tokens) success rate 100%. The problem is specific to long context.
Summary
Time-to-first-token on ~136K-token prompts varies randomly by more than 10x between requests of identical size sent one minute apart. Agent workloads with large context die to client/gateway timeouts when a request lands on a slow upstream channel. This matches and extends #978.
Decisive evidence (trace_ids provided for server-side lookup)
Exp 1 — identical payload resent 13 seconds apart
Exp 2 — same-size cold-cache requests ~1 minute apart (decisive)
Both are ~135.9K tokens, cold cache (cached=0), identical structure, differing in one word near the head. 8.9x throughput difference between upstream channels within the same minute.
Also observed (same test session)
Ruled out
Related existing issues
#978 (TTFT instability — same mechanism), #991 (502 on large sessions), #976 (tool-heavy long requests 400), #989 (non-streaming 524), #955 (422 above ~256K)
Secondary defects found along the way
Requested
Repro: any ~136K-token plain-text prompt, streaming, temperature 0. Full JSON results and generation script available on request.