FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

Expose model costs, capabilities, and plan allowances on the Provider API · Issue #983 · CommandCodeAI/command-code · GitHub

Repository navigation

Expose model costs, capabilities, and plan allowances on the Provider API #983

Description

Feature Description

There's no way to get model costs, capabilities, or what your plan includes programmatically.

GET /provider/v1/models returns these fields per model, and none of them carry cost, capability or plan data:

id, object, created, owned_by, name, context_length, supported_endpoints

No pricing route is documented, and the ones I tried return 404:

GET /provider/v1/models/pricing  → 404 "is not a registered API route"
GET /provider/v1                 → 404   (no route index is published either)

The numbers are published on the docs site (per-1M-token rates, capability flags, and per-plan monthly allowances), so they're reachable from whatever serves those pages. They just aren't available as data.

Use Case

Any client built on the Provider API where cost matters:

  • a cost-aware model picker: "which models does my plan include, and what do they cost?"
  • budget guardrails before a long or streaming run
  • comparing a cheap model against a frontier one without leaving the tool

I ran into this myself. To answer "what can I call on my plan and what does it cost", I had to scrape /docs/plans/goat and ship a bundled snapshot of the result. That snapshot goes stale the moment a deal starts or a rate moves.

Additional Context

Related: #914 asks for these same rates inside the /model catalogue. That request is about the CLI surface; this one is about the Provider API. Same numbers, different consumer, so they'd ideally be served from one source.

Two scopes, worth separating because their auth needs differ.

Model-scoped: identical for every caller, so cacheable and needing no auth. Natural home is extending /provider/v1/models:

pricing:     { input, output, cache_read, cache_write, currency }
price_bands: [ { condition, input, output, cache_read, cache_write } ]
capabilities: [ "text", "vision", "reasoning" ]
reasoning_effort_values: [ "off", "low", "medium", "high", "xhigh", "max" ]
intelligence: 47.0
updated_at: "2026-10-03"

Account-scoped: depends on the authenticated caller, so it needs its own route. This is the half that answers "which models does my plan include, and how much of each do I get":

GET /provider/v1/entitlements
→ { plan, balance, included: [ { model, monthly_credits } ], limits: { five_hour, weekly, monthly } }

Three things that are awkward to obtain today, each encountered while building the client:

  1. Stable identity. The docs identify models by display name, which drifts: the plan table shows DeepSeek V4 Flash (latest) where the API's id is deepseek/deepseek-v4-flash. Keying the data on id would let clients join reliably instead of matching on a display name.

  2. Alternate price bands. Plenty of models have more than one price tier, and the API exposes none of it. The docs site already models this properly: its page payload carries a structured tiers array per model, so GPT-5.6 Sol has {"label":"Standard","context":"≤ 272K","rates":{"input":5,"output":30,...}} and {"label":"Long context","context":"> 272K","rates":{"input":10,"output":45,...}}, and DeepSeek V4 Pro carries a timeOfDay object with peak, offPeak, peakHoursPerDay and the UTC windows. That is the shape an API should be returning. But it only reaches clients as a client-side React payload, and the rendered table shows the first tier with a +N badge, so anything that wants to know what a request actually costs has to parse a Next.js flight blob.

  3. Per-model reasoning support. reasoning_effort accepts off|low|medium|high|xhigh|max, but off is rejected per model at runtime:

     Model "stealth/space-bunny-alpha" does not support reasoning_effort "off".
     Only models that can turn thinking off accept it.   code: unsupported_value
    

    A per-model reasoning_effort_values array would make that discoverable rather than a 400.

Rates change often (deals start and expire, DeepSeek moves between peak and off-peak by the hour), so updated_at with ETag/Cache-Control would let clients cache honestly instead of re-fetching or going stale.

How important is this to you?

Important for my workflow

Activity

  1. acedatacloud-dev commented on Oct 5, 2026

    I work on AceDataCloud's model catalog and billing APIs. The split proposed here is important: model-scoped pricing/capability metadata can be cached publicly, while plan entitlements and remaining allowance are account-specific and must be authenticated.

    As a concrete external example, our public GET /api/v1/models/?with_pricing=1 currently returns model IDs with pricing metadata and a current credit-to-USD rate; it is separate from account balance/permission data. That still is not an entitlement API, and a model being listed must not imply a caller can use it. A client test should join on stable model ID, preserve unknown rather than inventing zero for absent prices, expose tier conditions/time windows, and invalidate cached prices with updated_at/ETag as you suggest. Public endpoint: https://platform.acedata.cloud/api/v1/models/?with_pricing=1.

    Disclosure: AceDataCloud affiliation. This is a comparable catalog design, not a claim that CommandCode already supports our models or plan semantics.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions


      Back | FazBrowse Home | New Git URL