| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
I can confirm the same issue. This does not appear to be client-specific.
Environment:
Results:
{
"thinking": {
"type": "disabled"
}
}
The request returned HTTP 200, but thinking was still enabled. The response used all 2400 completion tokens as reasoning tokens and returned no visible content.
2. With:
{
"enable_thinking": false
}
The request also returned HTTP 200, but thinking was still enabled and reasoning tokens were still generated.
3. With:
{
"reasoning_effort": "none"
}
The API returned HTTP 400:
Invalid option: expected one of "low" | "medium" | "high" | "xhigh" | "max"
4. Sending both reasoning_effort: "none" and thinking.type: "disabled" also returned HTTP 400.
5. reasoning_effort: "low" is accepted, but it does not disable thinking. In one test, the response still contained 1736 reasoning tokens.
I also inspected the official command-code CLI package (1.39.2). It exposes reasoning effort levels but does not provide a supported off option.
Could you please confirm whether the CommandCode DeepSeek V4 Flash route currently forces reasoning mode? If disabling reasoning is supported, what is the exact request field or model ID that should be used?
No API keys, prompts, generated content, or private request data are included in this report.plz fix this.
same
Working on the disable support.
Post-fix retest (2026-09-09, Provider API): thinking still cannot be disabled or lowered below low, and low still emits substantial reasoning tokens on several models
Follow-up data after the fix announced in #697 (2026-09-07, "I have shipped a fix for this one. Please retry") — independent benchmark across 10 models via POST https://api.commandcode.ai/provider/v1/chat/completions (OpenAI format, streaming, prompt: Reply with exactly: OK, max_tokens: 256, 1 warmup + 5 measured runs per model, sequential, 50/50 requests succeeded). Windows 10, direct HTTP, no CLI involved.
1. There is still no off / minimal level. On Qwen/Qwen3.7-Flash:
{ "reasoning_effort": "minimal" }→ 400 Invalid option: expected one of "low"|"medium"|"high"|"xhigh"|"max" — unchanged post-fix. low is the floor.
2. Reasoning tokens actually emitted at reasoning_effort: "low" (median of 5 runs, from the final usage chunk):
| Model | Reasoning tok (median) | Total latency (median) |
|---|---|---|
| Qwen/Qwen3.7-Flash | 159 | 3661 ms |
| meta/muse-spark-1.2-contributor | 178 | 2624 ms |
| stepfun/Step-3.5-Flash | 122 | 3010 ms |
| tencent/hy3-paid | 23 | 2341 ms |
| Qwen/Qwen3.8-Flash | 22 | 1501 ms |
| xiaomi/mimo-v2.5 | 23 | 3125 ms |
| meituan/LongCat-2.0:free | 21 | 2423 ms |
| deepseek/deepseek-v4-flash-fast | 20 | 1261 ms |
| z-ai/glm-5.3-flash | 0 | 2020 ms |
| poolside/laguna-s-2.1-free | 0 | 979 ms |
So even at the lowest supported effort, a trivial one-line prompt still burns 120–180 reasoning tokens on Qwen 3.7 Flash / Muse Spark 1.2 Contributor / Step 3.5 Flash. Only z-ai/glm-5.3-flash and poolside/laguna-s-2.1-free actually emit zero reasoning tokens at low.
Request: either expose a true disable (reasoning_effort: "none" / "minimal", or forward thinking: {"type": "disabled"} to upstreams that support it natively), or document the per-model minimum-effort behavior so agents can stop paying the latency/token tax. Happy to share the benchmark script and raw per-request JSON if useful.
same
Retested today and it still isn't fixed. This can't be that hard to fix..
Still doesn't work on deepseek-v4.1-flash
It’s hard to imagine that an issue like this, opened a month ago, still hasn’t been resolved.
same issue here. still not fix..
OpenCode Go handles this correctly, and this alone is a major reason for me to choose OpenCode Go. Please fix this issue.
You can now turn off thinking on all deepseek models from /effort picker. Update to command-code@1.73.3+
Confirmed: setting "reasoning_effort": "off" successfully disables thinking
| Back | FazBrowse Home | New Git URL |
Summary
deepseek V4 flash connot close thinking mode
Expected Behavior
{"thinking":{"type":"disabled"}}
or
{"reasoning_effort":"none"}
can enable
Actual Behavior
enable_thinking:false、reasoning:false will be ignore
and reasoning_effort:"none":CommandCode response 400,only permit low/medium/high/xhigh/max。
Steps to reproduce the issue
enable_thinking:false、reasoning:false will be ignore
and reasoning_effort:"none":CommandCode response 400,only permit low/medium/high/xhigh/max。
Command Code Version
1.38.2
Operating System
macOS
Terminal/IDE
Terminal
Shell
zsh
Session file (optional)
No response
Fix prompt (optional)
No response
Additional context
OS: macOS Sequoia 15.6