FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

deepseek V4 flash connot close thinking mode · Issue #776 · CommandCodeAI/command-code · GitHub

Repository navigation

deepseek V4 flash connot close thinking mode #776

Description

Summary

deepseek V4 flash connot close thinking mode

Expected Behavior

{"thinking":{"type":"disabled"}}
or
{"reasoning_effort":"none"}
can enable

Actual Behavior

enable_thinking:false、reasoning:false will be ignore
and reasoning_effort:"none":CommandCode response 400,only permit low/medium/high/xhigh/max。

Steps to reproduce the issue

enable_thinking:false、reasoning:false will be ignore
and reasoning_effort:"none":CommandCode response 400,only permit low/medium/high/xhigh/max。

Command Code Version

1.38.2

Operating System

macOS

Terminal/IDE

Terminal

Shell

zsh

Session file (optional)

No response

Fix prompt (optional)

No response

Additional context

OS: macOS Sequoia 15.6

Activity

  1. added theissue type on Aug 31, 2026
  2. EighTwT2 commented on Sep 1, 2026

    I can confirm the same issue. This does not appear to be client-specific.

    Environment:

    • Endpoint: POST /provider/v1/chat/completions
    • Model: deepseek/deepseek-v4-flash
    • OpenAI-compatible request format

    Results:

    1. With:
    {
      "thinking": {
        "type": "disabled"
      }
    }
    The request returned HTTP 200, but thinking was still enabled. The response used all 2400 completion tokens as reasoning tokens and returned no visible content.
    2. With:
    {
      "enable_thinking": false
    }
    The request also returned HTTP 200, but thinking was still enabled and reasoning tokens were still generated.
    3. With:
    {
      "reasoning_effort": "none"
    }
    The API returned HTTP 400:
    Invalid option: expected one of "low" | "medium" | "high" | "xhigh" | "max"
    4. Sending both reasoning_effort: "none" and thinking.type: "disabled" also returned HTTP 400.
    5. reasoning_effort: "low" is accepted, but it does not disable thinking. In one test, the response still contained 1736 reasoning tokens.
    I also inspected the official command-code CLI package (1.39.2). It exposes reasoning effort levels but does not provide a supported off option.
    Could you please confirm whether the CommandCode DeepSeek V4 Flash route currently forces reasoning mode? If disabling reasoning is supported, what is the exact request field or model ID that should be used?
    No API keys, prompts, generated content, or private request data are included in this report.
  3. Hangsiin commented on Sep 2, 2026

    plz fix this.

  4. sakura-madoromi commented on Sep 5, 2026

    same

  5. self-assigned this
    on Sep 7, 2026
  6. ahmadbilaldev commented on Sep 7, 2026

    Working on the disable support.

  7. a404670637-ch commented on Sep 9, 2026

    Post-fix retest (2026-09-09, Provider API): thinking still cannot be disabled or lowered below low, and low still emits substantial reasoning tokens on several models

    Follow-up data after the fix announced in #697 (2026-09-07, "I have shipped a fix for this one. Please retry") — independent benchmark across 10 models via POST https://api.commandcode.ai/provider/v1/chat/completions (OpenAI format, streaming, prompt: Reply with exactly: OK, max_tokens: 256, 1 warmup + 5 measured runs per model, sequential, 50/50 requests succeeded). Windows 10, direct HTTP, no CLI involved.

    1. There is still no off / minimal level. On Qwen/Qwen3.7-Flash:

    { "reasoning_effort": "minimal" }

    → 400 Invalid option: expected one of "low"|"medium"|"high"|"xhigh"|"max" — unchanged post-fix. low is the floor.

    2. Reasoning tokens actually emitted at reasoning_effort: "low" (median of 5 runs, from the final usage chunk):

    Model Reasoning tok (median) Total latency (median)
    Qwen/Qwen3.7-Flash 159 3661 ms
    meta/muse-spark-1.2-contributor 178 2624 ms
    stepfun/Step-3.5-Flash 122 3010 ms
    tencent/hy3-paid 23 2341 ms
    Qwen/Qwen3.8-Flash 22 1501 ms
    xiaomi/mimo-v2.5 23 3125 ms
    meituan/LongCat-2.0:free 21 2423 ms
    deepseek/deepseek-v4-flash-fast 20 1261 ms
    z-ai/glm-5.3-flash 0 2020 ms
    poolside/laguna-s-2.1-free 0 979 ms

    So even at the lowest supported effort, a trivial one-line prompt still burns 120–180 reasoning tokens on Qwen 3.7 Flash / Muse Spark 1.2 Contributor / Step 3.5 Flash. Only z-ai/glm-5.3-flash and poolside/laguna-s-2.1-free actually emit zero reasoning tokens at low.

    Request: either expose a true disable (reasoning_effort: "none" / "minimal", or forward thinking: {"type": "disabled"} to upstreams that support it natively), or document the per-model minimum-effort behavior so agents can stop paying the latency/token tax. Happy to share the benchmark script and raw per-request JSON if useful.

  8. Clivos commented on Sep 15, 2026

    same

  9. silvertakana commented on Sep 18, 2026

    Retested today and it still isn't fixed. This can't be that hard to fix..

  10. TigerBeanst commented on Sep 25, 2026

    Still doesn't work on deepseek-v4.1-flash

  11. Clivos commented on Sep 27, 2026

    It’s hard to imagine that an issue like this, opened a month ago, still hasn’t been resolved.

  12. chanmankong commented on Sep 28, 2026

    same issue here. still not fix..

  13. Hangsiin commented on Sep 28, 2026

    OpenCode Go handles this correctly, and this alone is a major reason for me to choose OpenCode Go. Please fix this issue.

  14. ahmadbilaldev commented on Oct 1, 2026

    You can now turn off thinking on all deepseek models from /effort picker. Update to command-code@1.73.3+

  15. Oyu-NAna commented on Oct 2, 2026

    Confirmed: setting "reasoning_effort": "off" successfully disables thinking

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions


    Back | FazBrowse Home | New Git URL