Feature Description
There's no way to get model costs, capabilities, or what your plan includes programmatically.
GET /provider/v1/models returns these fields per model, and none of them carry cost, capability or plan data:
id, object, created, owned_by, name, context_length, supported_endpoints
No pricing route is documented, and the ones I tried return 404:
GET /provider/v1/models/pricing → 404 "is not a registered API route"
GET /provider/v1 → 404 (no route index is published either)
The numbers are published on the docs site (per-1M-token rates, capability flags, and per-plan monthly allowances), so they're reachable from whatever serves those pages. They just aren't available as data.
Use Case
Any client built on the Provider API where cost matters:
- a cost-aware model picker: "which models does my plan include, and what do they cost?"
- budget guardrails before a long or streaming run
- comparing a cheap model against a frontier one without leaving the tool
I ran into this myself. To answer "what can I call on my plan and what does it cost", I had to scrape /docs/plans/goat and ship a bundled snapshot of the result. That snapshot goes stale the moment a deal starts or a rate moves.
Additional Context
Related: #914 asks for these same rates inside the /model catalogue. That request is about the CLI surface; this one is about the Provider API. Same numbers, different consumer, so they'd ideally be served from one source.
Two scopes, worth separating because their auth needs differ.
Model-scoped: identical for every caller, so cacheable and needing no auth. Natural home is extending /provider/v1/models:
pricing: { input, output, cache_read, cache_write, currency }
price_bands: [ { condition, input, output, cache_read, cache_write } ]
capabilities: [ "text", "vision", "reasoning" ]
reasoning_effort_values: [ "off", "low", "medium", "high", "xhigh", "max" ]
intelligence: 47.0
updated_at: "2026-10-03"
Account-scoped: depends on the authenticated caller, so it needs its own route. This is the half that answers "which models does my plan include, and how much of each do I get":
GET /provider/v1/entitlements
→ { plan, balance, included: [ { model, monthly_credits } ], limits: { five_hour, weekly, monthly } }
Three things that are awkward to obtain today, each encountered while building the client:
-
Stable identity. The docs identify models by display name, which drifts: the plan table shows DeepSeek V4 Flash (latest) where the API's id is deepseek/deepseek-v4-flash. Keying the data on id would let clients join reliably instead of matching on a display name.
-
Alternate price bands. Plenty of models have more than one price tier, and the API exposes none of it. The docs site already models this properly: its page payload carries a structured tiers array per model, so GPT-5.6 Sol has {"label":"Standard","context":"≤ 272K","rates":{"input":5,"output":30,...}} and {"label":"Long context","context":"> 272K","rates":{"input":10,"output":45,...}}, and DeepSeek V4 Pro carries a timeOfDay object with peak, offPeak, peakHoursPerDay and the UTC windows. That is the shape an API should be returning. But it only reaches clients as a client-side React payload, and the rendered table shows the first tier with a +N badge, so anything that wants to know what a request actually costs has to parse a Next.js flight blob.
-
Per-model reasoning support. reasoning_effort accepts off|low|medium|high|xhigh|max, but off is rejected per model at runtime:
Model "stealth/space-bunny-alpha" does not support reasoning_effort "off".
Only models that can turn thinking off accept it. code: unsupported_value
A per-model reasoning_effort_values array would make that discoverable rather than a 400.
Rates change often (deals start and expire, DeepSeek moves between peak and off-peak by the hour), so updated_at with ETag/Cache-Control would let clients cache honestly instead of re-fetching or going stale.
How important is this to you?
Important for my workflow
Feature Description
There's no way to get model costs, capabilities, or what your plan includes programmatically.
GET /provider/v1/models returns these fields per model, and none of them carry cost, capability or plan data:
No pricing route is documented, and the ones I tried return 404:
The numbers are published on the docs site (per-1M-token rates, capability flags, and per-plan monthly allowances), so they're reachable from whatever serves those pages. They just aren't available as data.
Use Case
Any client built on the Provider API where cost matters:
I ran into this myself. To answer "what can I call on my plan and what does it cost", I had to scrape /docs/plans/goat and ship a bundled snapshot of the result. That snapshot goes stale the moment a deal starts or a rate moves.
Additional Context
Related: #914 asks for these same rates inside the /model catalogue. That request is about the CLI surface; this one is about the Provider API. Same numbers, different consumer, so they'd ideally be served from one source.
Two scopes, worth separating because their auth needs differ.
Model-scoped: identical for every caller, so cacheable and needing no auth. Natural home is extending /provider/v1/models:
pricing: { input, output, cache_read, cache_write, currency } price_bands: [ { condition, input, output, cache_read, cache_write } ] capabilities: [ "text", "vision", "reasoning" ] reasoning_effort_values: [ "off", "low", "medium", "high", "xhigh", "max" ] intelligence: 47.0 updated_at: "2026-10-03"Account-scoped: depends on the authenticated caller, so it needs its own route. This is the half that answers "which models does my plan include, and how much of each do I get":
GET /provider/v1/entitlements → { plan, balance, included: [ { model, monthly_credits } ], limits: { five_hour, weekly, monthly } }Three things that are awkward to obtain today, each encountered while building the client:
Stable identity. The docs identify models by display name, which drifts: the plan table shows DeepSeek V4 Flash (latest) where the API's id is deepseek/deepseek-v4-flash. Keying the data on id would let clients join reliably instead of matching on a display name.
Alternate price bands. Plenty of models have more than one price tier, and the API exposes none of it. The docs site already models this properly: its page payload carries a structured tiers array per model, so GPT-5.6 Sol has {"label":"Standard","context":"≤ 272K","rates":{"input":5,"output":30,...}} and {"label":"Long context","context":"> 272K","rates":{"input":10,"output":45,...}}, and DeepSeek V4 Pro carries a timeOfDay object with peak, offPeak, peakHoursPerDay and the UTC windows. That is the shape an API should be returning. But it only reaches clients as a client-side React payload, and the rendered table shows the first tier with a +N badge, so anything that wants to know what a request actually costs has to parse a Next.js flight blob.
Per-model reasoning support. reasoning_effort accepts off|low|medium|high|xhigh|max, but off is rejected per model at runtime:
A per-model reasoning_effort_values array would make that discoverable rather than a 400.
Rates change often (deals start and expire, DeepSeek moves between peak and off-peak by the hour), so updated_at with ETag/Cache-Control would let clients cache honestly instead of re-fetching or going stale.
How important is this to you?
Important for my workflow