Context
1.0M
Input
$1.40
Output
$4.40
Cache read
$0.26
Cache write
—
Context
1.0M
Input
$1.40
Output
$4.40
Cache read
$0.26
Cache write
—
Context
262.1K
Input
$0.74
Output
$3.50
Cache read
$0.15
Cache write
—
Context
1M
Input
$0.32
Output
$0.8
Cache read
$0.08
Cache write
—
Context
1M
Input
$5
Output
$25
Cache read
$0.5
Cache write
$6.25
Context
256K
Input
$1
Output
$2
Cache read
$0.2
Cache write
—
Context
1M
Input
$1.50
Output
$9
Cache read
$0.15
Cache write
—
Context
1M
Input
$0.25
Output
$1.50
Cache read
$0.03
Cache write
—
Context
1M
Input
$1.25
Output
$2.50
Cache read
$0.2
Cache write
—
Context
262.1K
Input
$1.50
Output
$7.50
Cache read
—
Cache write
—
Context
1M
Input
$5
Output
$30
Cache read
$0.5
Cache write
—
Context
1M
Input
$30
Output
$180
Cache read
—
Cache write
—
Context
1.0M
Input
$0.14
Output
$0.28
Cache read
$0.028
Cache write
—
Context
1M
Input
$0.435
Output
$0.87
Cache read
$0.004
Cache write
—
Context
1M
Input
$5
Output
$25
Cache read
$0.5
Cache write
$6.25
Context
1.0M
Input
$2
Output
$12
Cache read
$0.2
Cache write
—
Context
262.1K
Input
$0.15
Output
$0.6
Cache read
$0.015
Cache write
—
Context
262.1K
Input
$0.14
Output
$0.4
Cache read
—
Cache write
—
Context
262.1K
Input
$0.25
Output
$0.9
Cache read
—
Cache write
—
Context
200K
Input
$0.3
Output
$1.20
Cache read
$0.06
Cache write
$0.375
Context
400K
Input
$0.2
Output
$1.25
Cache read
$0.02
Cache write
—
Context
1.1M
Input
$2.50
Output
$15
Cache read
$0.25
Cache write
—
Context
1.1M
Input
$30
Output
$180
Cache read
—
Cache write
—
Context
400K
Input
$1.75
Output
$14
Cache read
$0.175
Cache write
—
Context
400K
Input
$1.75
Output
$14
Cache read
$0.175
Cache write
—
Context
1M
Input
$3
Output
$15
Cache read
$0.3
Cache write
$3.75
1–25 of 48
Token prices are per 1M tokens; expand any model for full rates and capabilities, or open its detail page for the quickstart. The filters, sorting, search, pagination, and compare flow match the primary API catalog. Models marked DEDICATED are open-weight and available through Enterprise Inference as a single-tenant deployment.
Most models in this catalog run on the infrastructure of other providers; the Router is how you reach them safely. The models marked DEDICATED are open-weight, and Enterprise Inference runs them single-tenant on capacity reserved for you.
Every model in this catalog through one OpenAI-compatible endpoint, billed per token at the rates listed above. The Router reaches only upstream providers that are Zero Data Retention and do not train on your data.
Models marked DEDICATED are open-weight and can run in a deployment that is only yours — single-tenant, on capacity reserved for you, isolated from every other customer.
Every model above is one request away — OpenAI-compatible, with intelligent routing and prompt caching included.