Context
1M
Input
$2
Output
$6
Cache read
$0.25
Cache write
-
Context
1M
Input
$2
Output
$6
Cache read
$0.25
Cache write
-
Context
1M
Input
$5
Output
$30
Cache read
$0.5
Cache write
-
Context
1M
Input
$0.15
Output
$0.6
Cache read
$0.015
Cache write
-
Context
260K
Input
$0.04
Output
$0.15
Cache read
$0.004
Cache write
-
Context
1.1M
Input
$20
Output
$50
Cache read
$1
Cache write
$12.50
Context
1.1M
Input
$20
Output
$100
Cache read
$2
Cache write
$25
Context
1M
Input
$0.75
Output
$3.75
Cache read
$0.075
Cache write
-
Context
1.0M
Input
$2.10
Output
$6.60
Cache read
$0.21
Cache write
-
Context
991K
Input
$2
Output
$6
Cache read
$0.25
Cache write
$2.50
Context
1M
Input
$10
Output
$50
Cache read
$0.25
Cache write
$12.50
Context
1.0M
Input
$0.834
Output
$2.50
Cache read
$0.042
Cache write
-
Context
991K
Input
$0.15
Output
$0.47
Cache read
$0.016
Cache write
$0.2
Context
1M
Input
$0.15
Output
$0.5
Cache read
$0.03
Cache write
-
Context
1.0M
Input
$0.22
Output
$0.66
Cache read
$0.007
Cache write
-
Context
1M
Input
$1.40
Output
$4.40
Cache read
$0.14
Cache write
-
Context
1M
Input
$0.5
Output
$3
Cache read
$0.1
Cache write
$0.625
Context
1M
Input
$0.75
Output
$3.75
Cache read
$0.075
Cache write
-
Context
1M
Input
$0.66
Output
$1.98
Cache read
$0.066
Cache write
-
Context
500K
Input
$2
Output
$6
Cache read
$0.5
Cache write
-
Context
262.1K
Input
$0.05
Output
$0.2
Cache read
$0.01
Cache write
-
Context
131.1K
Input
$0.35
Output
$1.50
Cache read
$0.04
Cache write
-
Context
256K
Input
$0.021
Output
$0.063
Cache read
$0.004
Cache write
-
Context
1.0M
Input
$1.25
Output
$4.25
Cache read
$0.15
Cache write
-
Context
1.0M
Input
$0.1
Output
$0.2
Cache read
$0.002
Cache write
-
Context
256K
Input
$0.95
Output
$4
Cache read
$0.15
Cache write
-
1–25 of 233
Token prices are per 1M tokens; expand any model for full rates and capabilities, or open its detail page for the quickstart. The filters, sorting, search, pagination, and compare flow match the primary API catalog. Models marked DEDICATED are open-weight and available through Enterprise Inference as a single-tenant deployment.
Most models in this catalog run on the infrastructure of other providers; the Router is how you reach them safely. The models marked DEDICATED are open-weight, and Enterprise Inference runs them single-tenant on capacity reserved for you.
Every model in this catalog through one OpenAI-compatible endpoint, billed per token at the rates listed above. The Router reaches only upstream providers that are Zero Data Retention and do not train on your data.
Models marked DEDICATED are open-weight and can run in a deployment that is only yours: single-tenant, on capacity reserved for you, isolated from every other customer.
Every model above is one request away. OpenAI-compatible, with intelligent routing and prompt caching included.