FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

Provider API: non-streaming MiniMax-M3 chat completions die at ~126s with HTTP 524 and an empty body · Issue #989 · CommandCodeAI/command-code · GitHub

Repository navigation

Provider API: non-streaming MiniMax-M3 chat completions die at ~126s with HTTP 524 and an empty body #989

Description

Summary

I am using the Provider API to write a book with MiniMaxAI/MiniMax-M3, one section per request. Each request is a single non-streaming chat completion: a book spec, a short review of that spec, then a chapter of about 4000 words. reasoning_effort is low and max_tokens is 50000.

The spec and the review both return in under 50 seconds. The chapter sits silent until about 126 seconds, then the gateway returns HTTP 524 with an empty body. The same chapter sent again does the same thing. A third try a minute later returns HTTP 503 at 60 seconds, also with no JSON body.

Expected Behavior

The chapter call returns the finished chapter once the model is done, including when that takes several minutes.

When a call fails, the body is the JSON error described in the Provider API docs, with the upstream message and a request id.

Actual Behavior

Endpoint: POST https://api.commandcode.ai/provider/v1/chat/completions

Same key, same model, same minute. Times are 2026-10-04, America/Denver (MDT, UTC-6). stream was omitted, so it is false. One user message. No tools, no images.

Two short calls succeeded:

Call Wall time Prompt tokens Completion tokens Reasoning tokens finish_reason
Book spec 41.13s 2903 1979 513 stop
Review of that spec 47.86s 4159 3556 3229 stop

The review already took 48 seconds, and 3229 of its 3556 completion tokens were reasoning, for about 220 words of visible text.

The chapter used the same body shape. The prompt was about 3100 words / 19 KB and asked for about 4000 words. Three attempts:

Attempt Started (MDT) Duration Result
1 21:40:27 125.80s HTTP 524, empty body
2 21:42:33 126.84s HTTP 524, empty body. Same prompt, sent again immediately.
3 21:44:40 60.05s HTTP 503 Service Unavailable. No JSON body.

Client error text:

Command Code Provider API request to 'https://api.commandcode.ai/provider/v1/chat/completions' failed (HTTP 524): Response status code does not indicate success: 524 (<none>).
Command Code Provider API request to 'https://api.commandcode.ai/provider/v1/chat/completions' failed (HTTP 503): Response status code does not indicate success: 503 (Service Unavailable).

No partial text was returned. An earlier run the same evening did the same thing: a spec of about 50 seconds succeeded, and the chapter returned 524.

Steps to reproduce the issue

  1. Call POST https://api.commandcode.ai/provider/v1/chat/completions with model MiniMaxAI/MiniMax-M3, reasoning_effort set to low, and max_tokens set to 50000. Leave stream unset.
  2. Send a short user message and confirm a normal chat.completion comes back in a few seconds. This checks the key and the model id.
  3. Send one user message of a few thousand words that asks for a chapter of about 4000 words. The failing prompt here was about 3100 words and 19 KB.
  4. Wait. At about 126 seconds the client receives HTTP 524 and an empty body.
  5. Send that same request again. The second attempt also returns HTTP 524 at about 126 seconds. A third attempt may return HTTP 503 after about 60 seconds.

Request body:

{
  "model": "MiniMaxAI/MiniMax-M3",
  "messages": [{ "role": "user", "content": "<about 3100 words of story context and a chapter brief>" }],
  "max_tokens": 50000,
  "reasoning_effort": "low"
}

The client is PowerShell Invoke-WebRequest with no client-side timeout, so the status codes came from the gateway.

Command Code Version

1.74.1

Operating System

Windows

Terminal/IDE

Windows Terminal

Shell

pwsh

Session file (optional)

No response

Fix prompt (optional)

No response

Additional context

This matches Cloudflare error 524. The proxy connects, and the origin produces no response bytes before the roughly 100 second limit. With stream false, the socket stays silent until the full completion is ready. A chapter, on top of reasoning tokens, does not finish inside that window. The 48 second review is already halfway there because most of its completion tokens were reasoning.

The two 524s were likely still generating upstream when the proxy closed the client connection. Those tokens are wasted. The follow-up 503 looks like the gateway or the upstream was then overloaded. A request id on the error would make that checkable.

Activity

  1. added theissue type on Oct 5, 2026
  2. naymurdev commented on Oct 5, 2026

    are you still having that issue? And have you tried using another model?

  3. FireInWinter commented on Oct 6, 2026

    Author

    Yes. It is still happening on a non-streaming chat completion with max_tokens set to 50000.

    MiniMaxAI/MiniMax-M3 is faster than it was when this was filed, but reasoning_effort: max still dies at 125 seconds with failed (HTTP 524). Qwen/Qwen3.7-Flash does the same thing: it runs until 125 seconds, then returns 524.

    deepseek/deepseek-v4.1-flash ran for 240 seconds on the same kind of call and finished with finish_reason: stop. That one call did not hit the error. We do not know the cause, so this does not show that DeepSeek is clear of it.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions


      Back | FazBrowse Home | New Git URL