Llama 3.3 70B Instruct API pricing
Live per-token API pricing across providers, synced from OpenRouter. Compare input, output, and cache-read rates, throughput, latency, and uptime in one place - then estimate what your workload actually costs.
Refreshed about 10 hours ago via OpenRouter - some pricing may be stale.
DeepInfra - $0.10/M input, $0.32/M output
Per 1M tokens, USD. Verify on the provider's site before committing.
Per-provider pricing
Live per-provider pricing, throughput and uptime - refreshed about 10 hours ago via OpenRouter. Click a column to sort.
some pricing may be stale - last verified 2026-10-08
| Provider | Type | Input $/M | Output $/M | Cache $/M | Tok/s | Latency | Uptime | Value |
|---|---|---|---|---|---|---|---|---|
|
Google
|
API | 0.72 | 0.72 | - | - | - | - | |
|
Together AI
stale
|
API | 0.88 | 0.88 | - | - | - | - | |
|
Fireworks AI
stale
|
API | 0.90 | 0.90 | - | - | - | - | |
|
Parasail
|
API | 0.22 | 0.50 | 0.110 | - | - | 99.99% | best uptime |
|
AkashML
|
API | 0.20 | 0.52 | 0.100 | - | - | 99.93% | |
|
CoreWeave
|
API | 0.71 | 0.71 | 0.710 | - | - | 99.88% | |
| API | 0.59 | 0.79 | 0.295 | - | - | 99.85% | ||
|
Novita
|
API | 0.14 | 0.40 | - | - | - | 99.76% | |
|
SambaNova
|
API | 0.45 | 0.90 | - | - | - | 99.74% | |
|
Together
|
API | 1.04 | 1.04 | - | - | - | 99.70% | |
|
DeepInfra
|
API | 0.10 | 0.32 | - | - | - | 98.88% | cheapest |
|
Cloudflare
|
API | 0.29 | 2.25 | - | - | - | 97.93% |
Default order: throughput among 95%+ uptime providers, then latency; subscriptions last. Sort by any column. Subscription rows show $/mo in the Value column - per-token columns are "-". Affiliate links are marked sponsored / nofollow. Confirm current pricing on the provider's site before committing.
Get this data as JSON
Same provider table, machine-readable. No auth, no rate limit beyond the edge cache.
curl -s https://tokenstead.ai/models/llama-3-3-70b-instruct/pricing.json
JSON: model metadata, cheapest_api, providers[], subscriptions[], verified_at. Unit is USD per 1M tokens.
Estimate what Llama 3.3 70B Instruct costs for your workload
Paste your prompt, set requests/day, and see monthly cost against self-hosting on your own GPU.
Inference cost over time
Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.