Claude Opus 4.8 cost & VRAM calculator
What does Claude Opus 4.8 cost to run for your workload, and can you run it on your own hardware? Set your workload below - we compute per-provider API cost live and tell you honestly whether local hardware can run it.
This is where owning stops making sense
Claude Opus 4.8 is a large-billion-parameter model. No rig an individual can buy runs it - so unlike a workstation GPU model, there's no break-even to compute. The honest answer for nearly everyone is the per-provider API cost below.
Proprietary dense model from Anthropic, the previous flagship before Opus 5. Parameter count is undisclosed. Native 1M-token context. Adaptive thinking with effort levels; strong agentic coding and long-horizon capability. Pricing: $5/1M input, $25/1M output. Agents on Rails benchmark (Aug 2026, Le Mans round). 79.4% accuracy on 63 runs - tied with GLM 5.3 and the fastest model at its score tier (3m 36s median, ~15% of Opus 5’s time). API recall 15.9%, the third-lowest in the field. Superseded by Opus 5 (92.1%) on this board, but the accuracy-per-minute is the best of any model above 79%.
Your workload
API cost for your workload
| Provider | Rate ($/1M) | Monthly cost |
|---|---|---|
| Anthropic may be stale | $5.00 in · $25.00 out | $0.12 cheapest |
Monthly cost is an estimate from list prices and your workload - verify against the provider before committing. Cached fraction applies the cache rate to that share of input.
Can you run it locally?
Claude Opus 4.8 has no published quantization that fits a rig one person can buy, so there's no local-hardware recommendation and no break-even to compute. The honest answer is the per-provider API cost above.
Proprietary dense model from Anthropic, the previous flagship before Opus 5. Parameter count is undisclosed. Native 1M-token context. Adaptive thinking with effort levels; strong agentic coding and long-horizon capability. Pricing: $5/1M input, $25/1M output. Agents on Rails benchmark (Aug 2026, Le Mans round). 79.4% accuracy on 63 runs - tied with GLM 5.3 and the fastest model at its score tier (3m 36s median, ~15% of Opus 5’s time). API recall 15.9%, the third-lowest in the field. Superseded by Opus 5 (92.1%) on this board, but the accuracy-per-minute is the best of any model above 79%.
See the model card for the full architecture notes and any cloud subscription plans.