FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

Recurring 502 "Stream ended unexpectedly before completion" on large sessions — turns lost after retry exhaustion (macOS, v0.1.47) · Issue #991 · CommandCodeAI/command-code · GitHub

Repository navigation

Recurring 502 "Stream ended unexpectedly before completion" on large sessions — turns lost after retry exhaustion (macOS, v0.1.47) #991

Description

Summary

Turns on large sessions repeatedly fail with:

502 Stream ended unexpectedly before completion (no finish event) — response was truncated

The failure hits mid-stream. The client retries internally (stream_restart up to 3×, api_retry up to 8× with ~10s backoff), but when retries exhaust (~10 min) the entire turn is lost and the app ends it as run_error. Retrying later sometimes succeeds. Small/medium sessions were fine during the same window; the failures concentrate on large sessions (transcript ≈ 27 MB, requests ≈ 350–400K tokens including many base64 image blocks). Even a 2-character prompt ("hi") failed after ~10 min of retries.

Expected Behavior

The turn completes normally. If the upstream stream truncates, automatic retries recover (ideally by resuming the partial response) instead of losing the whole turn.

Actual Behavior

Turns end with "502 Stream ended unexpectedly before completion (no finish event) — response was truncated"; after internal retries exhaust, the app ends the turn with run_error and the response is lost. Failed attempts consume 0 tokens.

Traces (times UTC+8; session A unless noted):

  • Oct 5 22:03 — b85bc5fd3a4e49d382dab12462f34ec9
  • Oct 5 22:35 — 7c64a91e28e9eee75ba94ecda838fc56
  • Oct 5 22:50 — ecb6851860fe19c292187e3c158d7dce
  • Oct 5 23:02 — 328b5cd1d47675ea88a464dc4618c035
  • Oct 6 00:04 — 5a6eb72090d1b091c6ca6cf36ee3a6a3
  • Oct 6 00:17 — 5c8fe322d8b2c029063623048a7da8a9
  • Oct 6 02:26 — 85f2080cb801b9fe1be786c664bdac90
  • Oct 5 17:48 — 1ae311d885cfc033964bc7c929e2fdc1 (session B)
  • Oct 5 17:51 — 0dd705cddacc282fb3c9763be47ae635 (session B)
  • Oct 5 18:36 — 6cdcd37f8bd0af0a5eb5a617565f53ef (session B)

Steps to reproduce the issue

  1. Open a large session (transcript ≈ 27 MB; recent context includes many image tool results).
  2. Send any message (short ones included) with model deepseek/deepseek-v4.1-flash at reasoning effort max.
  3. The response stream truncates mid-way; after internal retries exhaust the turn ends with run_error.
  4. Same behavior via headless: cmd --session -p "" — retries for minutes, then exits with "The API server encountered an error. Please try again later."

Command Code Version

0.1.47

Operating System

macOS

Terminal/IDE

N/A (desktop app)

Shell

zsh

Session file (optional)

No response

Fix prompt (optional)

Check the gateway logs for the traces above. Investigate stream truncation (no finish event) on very large requests (large context + image blocks), and make failures more recoverable — e.g. resume/continue the truncated stream instead of discarding the whole response, and clearer error semantics when retries exhaust.

Additional context

Environment: macOS (Apple Silicon); Command Code desktop 0.1.47; model deepseek/deepseek-v4.1-flash, reasoning effort max. Reproduced in the desktop app and via headless CLI on the same night.

Notes:

  • Workaround that helps: instructing the model to keep responses small / step-by-step — turns then usually complete.
  • Failed attempts show 0 tokens consumed; sessions stayed intact and resumable (no data loss).
  • Healthy small sessions in the same window suggests this is tied to very large request payloads / long streams.

Activity

  1. added theissue type on Oct 6, 2026
  2. Kernlix commented on Oct 7, 2026

    Cross-referencing with some quantified evidence — I hit the same error on Windows (v1.74.1, Electron app), so it doesn't look platform-specific. The measurements point strongly at request payload size, with base64 image data making up ~90% of it.

    Timeline (UTC+8), same session:

    Time Context size Result
    2026-10-07 22:19:52 ~23 MB ❌ 502
    2026-10-07 22:26:42 ~23 MB (unchanged — plain retry) ❌ 502
    2026-10-07 22:31:08 — auto-compaction ran (tokensSaved: 400540)
    2026-10-07 22:32:31 much smaller ✅ success

    Payload composition at the failing request (measured from the on-disk session transcript — all 779 lines parse as valid JSON, so nothing was corrupted locally):

    Component Size Share
    Base64 image data 20.9 MB 90.3%
    Text / tool output / thinking 2.2 MB 9.7%
    Total 23.1 MB 100%

    63 image tool-results had accumulated. They stay in context and are re-sent on every subsequent turn:

    after image  1:   0.8 MB
    after image 20:   5.5 MB
    after image 40:  12.0 MB
    after image 60:  22.5 MB
    

    This matches the retry behaviour you describe: an identical payload fails again, because the thing that breaks it hasn't changed. Once compaction shrank the payload, the very next attempt went through.

    Full details — trace IDs, exact repro steps, and suggested fixes (e.g. return a 413 with a /compact hint instead of a mid-stream 502) — are in #997.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions


      Back | FazBrowse Home | New Git URL