LongCat 2.0
MoE premier1.6T total params, MoE with ~48B active per token. 1M native context via LongCat Sparse Attention (LSA); zero-compute experts dynamically route 33-56B per token.
Open weights under MIT license (also the model behind OpenRouter’s Owl Alpha). FP8 and INT8 quantized variants published alongside the base model.
- 1600.0B
- 1000k
- mit
- 🇨🇳 China
- Jul 2026
Scores
Score per dollar
287 pts per $/M input
general_score (86) divided by cheapest input price ($0.30/M). Higher is better value. See live pricing.
Save your hardware and every model page answers the real question: will it run on your machine, and how fast?
Join free - save your rig →Or run it in the cloud
Live per-provider pricing, throughput and uptime - refreshed about 17 hours ago via OpenRouter. Click a column to sort.
| Provider | Type | Input $/M | Output $/M | Cache $/M | Tok/s | Latency | Uptime | Value |
|---|---|---|---|---|---|---|---|---|
|
AtlasCloud
|
API | 0.30 | 1.20 | 0.006 | - | - | 100.00% | best uptime |
Default order: throughput among 95%+ uptime providers, then latency; subscriptions last. Sort by any column. Subscription rows show $/mo in the Value column - per-token columns are "-". Affiliate links are marked sponsored / nofollow. Confirm current pricing on the provider's site before committing.
Detailed API pricing page + JSON endpoint →
Inference cost over time
Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.