Background (how this was found)
Normal CLI usage with meta/muse-spark-1.3-contributor (GOAT plan, cmd v1.79.2, macOS): prompt-cache hits are abnormal/unstable in long sessions (high then dropping to 0, flapping). Only after noticing this did I repro directly against the Provider API with /chat/completions and /responses using a fixed prefix — the API results below are the controlled confirmation, not the original symptom.
Repro (Provider API, controlled)
Fixed ~30k-char prefix, per-round only Round N: reply exactly RN, 1.5s gap, store:false on /responses:
# /responses per round:
{"model":"meta/muse-spark-1.3-contributor","instructions":"<FIXED ~30k chars>","input":"Round N: reply exactly RN","store":false,"max_output_tokens":50}
# /chat/completions per round:
{"model":"meta/muse-spark-1.3-contributor","messages":[{"role":"system","content":"<SAME FIXED TEXT>"},{"role":"user","content":"Round N: reply exactly RN"}],"temperature":0,"max_tokens":100,"stream":false}
# read cached tokens + system_fingerprint per round
Results (meta/muse-spark-1.3-contributor, input=6926)
| endpoint |
cached R1..Rn |
note |
| /responses, no prompt_cache_key (6 rounds) |
0 / 0 / 0 / 0 / 0 / 0 |
all miss in this run; hits are unstable across runs |
| /responses + fixed prompt_cache_key (6 rounds) |
0 / 6897 x5 |
~99.6% from R2 once pinned |
| /chat/completions (5 rounds) |
0 / 0 / 0 / 0 / 6897 |
fingerprints differ every round: fp_qsj404zp84, fp_cjeke1ia6o, fp_al1pyowolc, fp_pb2xpoxxs1, fp_ezxo85zttd |
Reads like a routing problem, not a fixed 0%: same-prefix turns scatter (fingerprint changes every round on chat; responses only sticks with a manual key). responses envelope also has no system_fingerprint to diagnose routing. And the CLI exposes no prompt_cache_key, so there is no in-product workaround.
Ask
- Sticky/session-aware routing (or implicit stable cache key) so plain same-prefix CLI turns hit stably without manual prompt_cache_key.
- Expose prompt_cache_key (or equivalent) in CLI/BYOK options for pinned sessions.
- Please also check other models for similar cache/endpoint inconsistencies (spot check hit one model returning 400 unsupported_model on /responses while serving fine on /chat/completions, and another flapping hit-then-miss); worth an audit across models/endpoints.
Env: cmd v1.79.2, macOS, GOAT plan, 2026-10-09.
Background (how this was found)
Normal CLI usage with meta/muse-spark-1.3-contributor (GOAT plan, cmd v1.79.2, macOS): prompt-cache hits are abnormal/unstable in long sessions (high then dropping to 0, flapping). Only after noticing this did I repro directly against the Provider API with /chat/completions and /responses using a fixed prefix — the API results below are the controlled confirmation, not the original symptom.
Repro (Provider API, controlled)
Fixed ~30k-char prefix, per-round only Round N: reply exactly RN, 1.5s gap, store:false on /responses:
Results (meta/muse-spark-1.3-contributor, input=6926)
Reads like a routing problem, not a fixed 0%: same-prefix turns scatter (fingerprint changes every round on chat; responses only sticks with a manual key). responses envelope also has no system_fingerprint to diagnose routing. And the CLI exposes no prompt_cache_key, so there is no in-product workaround.
Ask
Env: cmd v1.79.2, macOS, GOAT plan, 2026-10-09.