FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

Provider API: unstable and frequently very long time-to-first-byte on streaming requests (commonly 20s+, up to 146.8s measured) — deepseek/deepseek-v4.1-flash from mainland China · Issue #978 · CommandCodeAI/command-code · GitHub

Repository navigation

Provider API: unstable and frequently very long time-to-first-byte on streaming requests (commonly 20s+, up to 146.8s measured) — deepseek/deepseek-v4.1-flash from mainland China #978

Description

Summary

I consume the Command Code provider API from a third-party OpenAI-compatible client. With model=deepseek/deepseek-v4.1-flash and stream=true, time-to-first-token is the clear weak point: it is not a stable cost — it varies widely between requests, but in my day-to-day use it is long far more often than not. 20 seconds or more is common, and I have measured it as high as 146.8 seconds.

Once the first token finally arrives, generation is fast. Here is one instrumented request from my client (screenshot attached):

start total first-token latency generation throughput
2026-10-03 08:18:27 167.9s 146.8s 21.1s 175.9 tok/s

Those numbers are the whole report: of 167.9 seconds total, 146.8 seconds (87%) elapsed before the first token arrived, while the generation phase afterwards ran at 175.9 tok/s — a perfectly healthy rate. Nothing is wrong with how fast the model generates; the problem is entirely in front of the first token.

Expected Behavior

  • On a fast model such as V4.1 Flash, a streaming request should normally begin returning content within a few seconds, and should not routinely sit silent for 20s+.
  • Even when the upstream model needs a long time to think (V4.1 Flash is a reasoning model with thinking on by default), the gateway should flush SSE response headers and/or emit periodic SSE comment heartbeats so clients and intermediaries know the connection is alive rather than stalled.
  • A client should not have to distinguish "model is thinking" from "edge is buffering" by guesswork; a server-timing/TTFB header would make this trivial.

Actual Behavior

  • Time-to-first-token is long far more often than not, commonly 20s+, with a measured worst case of 146.8s. It is not every single request — the number swings quite a bit between requests, and some requests do come back considerably faster.
  • While waiting, the client receives nothing to render: the prompt sits there with no content. (I have not captured a raw packet trace, so I can't state whether the HTTP response headers themselves arrive early; what I can confirm is that no token content reaches the client during that window.)
  • After the first token, throughput is normal — 175.9 tok/s in the measurement above.

Steps to reproduce the issue

  1. Configure an OpenAI-compatible client (I use DSH) with the Command Code provider base URL and API key.
  2. Set model=deepseek/deepseek-v4.1-flash, stream=true.
  3. Send a series of ordinary prompts and record the client's time-to-first-token for each.
  4. Observe that TTFT is frequently 20s or more, varies widely from request to request, and in at least one case reached 146.8s.
  5. Note that once generation begins, throughput is normal (~175.9 tok/s) — total wall-clock time is dominated by the pre-token wait.

Command Code Version

n/a — consumed via the hosted provider API endpoint, not the CLI.

Operating System

Windows

Terminal/IDE

DSH (third-party OpenAI-compatible API client)

Shell

n/a — HTTP API client

Session file (optional)

No response

Fix prompt (optional)

No response

Additional context

Client-side request timing for one streaming call (screenshot attached):

请求计时 (request timing)
  开始时间 (start)      2026-10-03 08:18:27.224
  总时长 (total)        167.9 s
  首 token 延迟 (TTFT)  146.8 s
  生成 (generation)      21.1 s
  吞吐量 (throughput)    175.9 tok/s

Activity

  1. Preacher7306 commented on Oct 4, 2026

    I've been using the API for two months now, and it has been unstable the whole time. If you continue to be this irresponsible, I will make my friends never use your product again!!!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions


      Back | FazBrowse Home | New Git URL