Summary
I consume the Command Code provider API from a third-party OpenAI-compatible client. With model=deepseek/deepseek-v4.1-flash and stream=true, time-to-first-token is the clear weak point: it is not a stable cost — it varies widely between requests, but in my day-to-day use it is long far more often than not. 20 seconds or more is common, and I have measured it as high as 146.8 seconds.
Once the first token finally arrives, generation is fast. Here is one instrumented request from my client (screenshot attached):
| start |
total |
first-token latency |
generation |
throughput |
| 2026-10-03 08:18:27 |
167.9s |
146.8s |
21.1s |
175.9 tok/s |
Those numbers are the whole report: of 167.9 seconds total, 146.8 seconds (87%) elapsed before the first token arrived, while the generation phase afterwards ran at 175.9 tok/s — a perfectly healthy rate. Nothing is wrong with how fast the model generates; the problem is entirely in front of the first token.
Expected Behavior
- On a fast model such as V4.1 Flash, a streaming request should normally begin returning content within a few seconds, and should not routinely sit silent for 20s+.
- Even when the upstream model needs a long time to think (V4.1 Flash is a reasoning model with thinking on by default), the gateway should flush SSE response headers and/or emit periodic SSE comment heartbeats so clients and intermediaries know the connection is alive rather than stalled.
- A client should not have to distinguish "model is thinking" from "edge is buffering" by guesswork; a server-timing/TTFB header would make this trivial.
Actual Behavior
- Time-to-first-token is long far more often than not, commonly 20s+, with a measured worst case of 146.8s. It is not every single request — the number swings quite a bit between requests, and some requests do come back considerably faster.
- While waiting, the client receives nothing to render: the prompt sits there with no content. (I have not captured a raw packet trace, so I can't state whether the HTTP response headers themselves arrive early; what I can confirm is that no token content reaches the client during that window.)
- After the first token, throughput is normal — 175.9 tok/s in the measurement above.
Steps to reproduce the issue
- Configure an OpenAI-compatible client (I use DSH) with the Command Code provider base URL and API key.
- Set model=deepseek/deepseek-v4.1-flash, stream=true.
- Send a series of ordinary prompts and record the client's time-to-first-token for each.
- Observe that TTFT is frequently 20s or more, varies widely from request to request, and in at least one case reached 146.8s.
- Note that once generation begins, throughput is normal (~175.9 tok/s) — total wall-clock time is dominated by the pre-token wait.
Command Code Version
n/a — consumed via the hosted provider API endpoint, not the CLI.
Operating System
Windows
Terminal/IDE
DSH (third-party OpenAI-compatible API client)
Shell
n/a — HTTP API client
Session file (optional)
No response
Fix prompt (optional)
No response
Additional context
Client-side request timing for one streaming call (screenshot attached):
请求计时 (request timing)
开始时间 (start) 2026-10-03 08:18:27.224
总时长 (total) 167.9 s
首 token 延迟 (TTFT) 146.8 s
生成 (generation) 21.1 s
吞吐量 (throughput) 175.9 tok/s
- Network: mainland China, direct connection, no proxy/VPN. Dynamic requests fail from Celerity/PuntoNet Ecuador (Cloudflare UIO) but work via Movistar #790 documents a region-specific relay/edge-node failure (Cloudflare UIO), so the edge node serving mainland China is a plausible contributor and worth checking.
- Model: deepseek/deepseek-v4.1-flash (reasoning model, thinking on by default), not the default deepseek/deepseek-v4-flash. Think-token looping / raw (unrepaired) tool-call failures on DeepSeek V4.1 Flash & GLM 5.3 #967 reports think-token looping on this same model, which would also present as high TTFT. I can't fully rule that out — a pure API client doesn't expose reasoning tokens — but the healthy 175.9 tok/s afterwards and the large request-to-request variance both argue against a model that is simply reasoning slowly every time.
- Related issues: API 524 on non-streaming requests that take too long (~180s); streaming works fine #696 (relay buffers non-streaming responses → 180s 524 — buffering on the same relay), Bug: First prompt after suspend/resume hangs for seconds to minutes #819 (first prompt stalls — different trigger), Think-token looping / raw (unrepaired) tool-call failures on DeepSeek V4.1 Flash & GLM 5.3 #967 (think-token looping on V4.1 Flash / GLM 5.3), Dynamic requests fail from Celerity/PuntoNet Ecuador (Cloudflare UIO) but work via Movistar #790 (region-specific relay failure), deepseek V4 flash connot close thinking mode #776 (thinking mode cannot be disabled).
- Why I suspect the gateway rather than the model: generation after the first token runs at 175.9 tok/s, so the model is fast. And the TTFT is not a stable, prompt-dependent cost — it swings widely between otherwise ordinary requests, which is characteristic of queueing, cold connections or edge-node scheduling. Model inference time tends to be comparatively steady and to show up as reduced throughput, not as a silent wait followed by a fast stream.
Summary
I consume the Command Code provider API from a third-party OpenAI-compatible client. With model=deepseek/deepseek-v4.1-flash and stream=true, time-to-first-token is the clear weak point: it is not a stable cost — it varies widely between requests, but in my day-to-day use it is long far more often than not. 20 seconds or more is common, and I have measured it as high as 146.8 seconds.
Once the first token finally arrives, generation is fast. Here is one instrumented request from my client (screenshot attached):
Those numbers are the whole report: of 167.9 seconds total, 146.8 seconds (87%) elapsed before the first token arrived, while the generation phase afterwards ran at 175.9 tok/s — a perfectly healthy rate. Nothing is wrong with how fast the model generates; the problem is entirely in front of the first token.
Expected Behavior
Actual Behavior
Steps to reproduce the issue
Command Code Version
n/a — consumed via the hosted provider API endpoint, not the CLI.
Operating System
Windows
Terminal/IDE
DSH (third-party OpenAI-compatible API client)
Shell
n/a — HTTP API client
Session file (optional)
No response
Fix prompt (optional)
No response
Additional context
Client-side request timing for one streaming call (screenshot attached):