Summary
I am using the Provider API to write a book with MiniMaxAI/MiniMax-M3, one section per request. Each request is a single non-streaming chat completion: a book spec, a short review of that spec, then a chapter of about 4000 words. reasoning_effort is low and max_tokens is 50000.
The spec and the review both return in under 50 seconds. The chapter sits silent until about 126 seconds, then the gateway returns HTTP 524 with an empty body. The same chapter sent again does the same thing. A third try a minute later returns HTTP 503 at 60 seconds, also with no JSON body.
Expected Behavior
The chapter call returns the finished chapter once the model is done, including when that takes several minutes.
When a call fails, the body is the JSON error described in the Provider API docs, with the upstream message and a request id.
Actual Behavior
Endpoint: POST https://api.commandcode.ai/provider/v1/chat/completions
Same key, same model, same minute. Times are 2026-10-04, America/Denver (MDT, UTC-6). stream was omitted, so it is false. One user message. No tools, no images.
Two short calls succeeded:
| Call |
Wall time |
Prompt tokens |
Completion tokens |
Reasoning tokens |
finish_reason |
| Book spec |
41.13s |
2903 |
1979 |
513 |
stop |
| Review of that spec |
47.86s |
4159 |
3556 |
3229 |
stop |
The review already took 48 seconds, and 3229 of its 3556 completion tokens were reasoning, for about 220 words of visible text.
The chapter used the same body shape. The prompt was about 3100 words / 19 KB and asked for about 4000 words. Three attempts:
| Attempt |
Started (MDT) |
Duration |
Result |
| 1 |
21:40:27 |
125.80s |
HTTP 524, empty body |
| 2 |
21:42:33 |
126.84s |
HTTP 524, empty body. Same prompt, sent again immediately. |
| 3 |
21:44:40 |
60.05s |
HTTP 503 Service Unavailable. No JSON body. |
Client error text:
Command Code Provider API request to 'https://api.commandcode.ai/provider/v1/chat/completions' failed (HTTP 524): Response status code does not indicate success: 524 (<none>).
Command Code Provider API request to 'https://api.commandcode.ai/provider/v1/chat/completions' failed (HTTP 503): Response status code does not indicate success: 503 (Service Unavailable).
No partial text was returned. An earlier run the same evening did the same thing: a spec of about 50 seconds succeeded, and the chapter returned 524.
Steps to reproduce the issue
- Call POST https://api.commandcode.ai/provider/v1/chat/completions with model MiniMaxAI/MiniMax-M3, reasoning_effort set to low, and max_tokens set to 50000. Leave stream unset.
- Send a short user message and confirm a normal chat.completion comes back in a few seconds. This checks the key and the model id.
- Send one user message of a few thousand words that asks for a chapter of about 4000 words. The failing prompt here was about 3100 words and 19 KB.
- Wait. At about 126 seconds the client receives HTTP 524 and an empty body.
- Send that same request again. The second attempt also returns HTTP 524 at about 126 seconds. A third attempt may return HTTP 503 after about 60 seconds.
Request body:
{
"model": "MiniMaxAI/MiniMax-M3",
"messages": [{ "role": "user", "content": "<about 3100 words of story context and a chapter brief>" }],
"max_tokens": 50000,
"reasoning_effort": "low"
}
The client is PowerShell Invoke-WebRequest with no client-side timeout, so the status codes came from the gateway.
Command Code Version
1.74.1
Operating System
Windows
Terminal/IDE
Windows Terminal
Shell
pwsh
Session file (optional)
No response
Fix prompt (optional)
No response
Additional context
This matches Cloudflare error 524. The proxy connects, and the origin produces no response bytes before the roughly 100 second limit. With stream false, the socket stays silent until the full completion is ready. A chapter, on top of reasoning tokens, does not finish inside that window. The 48 second review is already halfway there because most of its completion tokens were reasoning.
The two 524s were likely still generating upstream when the proxy closed the client connection. Those tokens are wasted. The follow-up 503 looks like the gateway or the upstream was then overloaded. A request id on the error would make that checkable.
Summary
I am using the Provider API to write a book with MiniMaxAI/MiniMax-M3, one section per request. Each request is a single non-streaming chat completion: a book spec, a short review of that spec, then a chapter of about 4000 words. reasoning_effort is low and max_tokens is 50000.
The spec and the review both return in under 50 seconds. The chapter sits silent until about 126 seconds, then the gateway returns HTTP 524 with an empty body. The same chapter sent again does the same thing. A third try a minute later returns HTTP 503 at 60 seconds, also with no JSON body.
Expected Behavior
The chapter call returns the finished chapter once the model is done, including when that takes several minutes.
When a call fails, the body is the JSON error described in the Provider API docs, with the upstream message and a request id.
Actual Behavior
Endpoint: POST https://api.commandcode.ai/provider/v1/chat/completions
Same key, same model, same minute. Times are 2026-10-04, America/Denver (MDT, UTC-6). stream was omitted, so it is false. One user message. No tools, no images.
Two short calls succeeded:
The review already took 48 seconds, and 3229 of its 3556 completion tokens were reasoning, for about 220 words of visible text.
The chapter used the same body shape. The prompt was about 3100 words / 19 KB and asked for about 4000 words. Three attempts:
Client error text:
No partial text was returned. An earlier run the same evening did the same thing: a spec of about 50 seconds succeeded, and the chapter returned 524.
Steps to reproduce the issue
Request body:
{ "model": "MiniMaxAI/MiniMax-M3", "messages": [{ "role": "user", "content": "<about 3100 words of story context and a chapter brief>" }], "max_tokens": 50000, "reasoning_effort": "low" }The client is PowerShell Invoke-WebRequest with no client-side timeout, so the status codes came from the gateway.
Command Code Version
1.74.1
Operating System
Windows
Terminal/IDE
Windows Terminal
Shell
pwsh
Session file (optional)
No response
Fix prompt (optional)
No response
Additional context
This matches Cloudflare error 524. The proxy connects, and the origin produces no response bytes before the roughly 100 second limit. With stream false, the socket stays silent until the full completion is ready. A chapter, on top of reasoning tokens, does not finish inside that window. The 48 second review is already halfway there because most of its completion tokens were reasoning.
The two 524s were likely still generating upstream when the proxy closed the client connection. Those tokens are wasted. The follow-up 503 looks like the gateway or the upstream was then overloaded. A request id on the error would make that checkable.