Models / Kimi K2.7 Code / Calculator

Kimi K2.7 Code cost & VRAM calculator

What does Kimi K2.7 Code cost to run for your workload, and can you run it on your own hardware? Set your workload below - we compute per-provider API cost live and tell you honestly whether local hardware can run it.

LOCAL ISN'T PRACTICAL

This is where owning stops making sense

Kimi K2.7 Code is a 1000-billion-parameter model. No rig an individual can buy runs it - so unlike a workstation GPU model, there's no break-even to compute. The honest answer for nearly everyone is the per-provider API cost below.

1T total params, 32B active per token (MoE) - 384 experts, 8 selected + 1 shared, 61 layers, 7168 attention hidden dimension, 64 heads, MLA, SwiGLU activation, 160K vocabulary. Built on Kimi K2.6 as a coding/agentic specialization, not a general chat replacement. Context and I/O: native 256K context (262,144 tokens via API); native vision and multimodal tool use (PNG/JPEG/WebP/GIF images, MP4/MOV/AVI video). Reasoning and tools: always-on thinking mode; Moonshot claims ~30% fewer thinking tokens than K2.6. Native function/tool calling with MCP support. Open weights under a Modified MIT license on HuggingFace at moonshotai/Kimi-K2.7-Code. The license requires products with >100M monthly active users or >$20M monthly revenue to display “Kimi K2.7 Code” prominently. Cloud-only for most users. A 1T-parameter MoE needs server-class hardware; there is no published Unsloth GGUF or consumer-grade quantization, so self-hosting is not practical on a workstation or DGX Spark. Run it via the Kimi API, OpenRouter, or vLLM/SGLang on a cluster. Cloud API: $0.95/1M input, $4.00/1M output, $0.19/1M cache-hit input. A kimi-k2.7-code-highspeed variant runs at ~180 tok/s (up to 260 tok/s in short contexts) for 2x the price. Benchmarks (Moonshot self-reported): Kimi Code Bench v2 62.0, Program Bench 53.6, MLS Bench Lite 35.1, Kimi Claw 24/7 Bench 46.9, MCP Atlas 76.0, MCP Mark Verified 81.1.

Your workload

Kimi K2.7 Code runs an always-on thinking mode. Reasoning (thinking) tokens are billed at the output rate ($15.00/M), so count them here to see the thinking portion of your bill.

API cost for your workload

Provider Rate ($/1M) Monthly cost
CoreWeave $0.71 in · $3.50 out · $0.15 cache $0.02 cheapest
StreamLake $0.71 in · $3.00 out · $0.14 cache $0.02 cheapest
Venice $0.75 in · $3.50 out · $0.16 cache $0.02 cheapest
ModelRun $0.85 in · $3.75 out · $0.16 cache $0.02 cheapest
Cloudflare $0.95 in · $4.00 out · $0.19 cache $0.02 cheapest
BaseTen $0.95 in · $4.00 out · $0.16 cache $0.02 cheapest
SiliconFlow $0.86 in · $3.80 out · $0.18 cache $0.02 cheapest
Moonshot AI $1.90 in · $8.00 out · $0.38 cache $0.05
Novita $0.91 in · $3.84 out · $0.18 cache $0.02 cheapest
GMICloud $0.95 in · $4.00 out · $0.19 cache $0.02 cheapest
Nebius $0.95 in · $4.00 out · $0.18 cache $0.02 cheapest
Inceptron $0.66 in · $3.30 out · $0.18 cache $0.02 cheapest
DeepInfra $0.68 in · $3.40 out · $0.14 cache $0.02 cheapest
Alibaba $0.95 in · $4.00 out · $0.19 cache $0.02 cheapest
Fireworks avoid $0.95 in · $4.00 out · $0.19 cache $0.02 cheapest

Monthly cost is an estimate from list prices and your workload - verify against the provider before committing. Cached fraction applies the cache rate to that share of input.

Can you run it locally?

NO - NO INDIVIDUAL RIG RUNS IT

Kimi K2.7 Code has no published quantization that fits a rig one person can buy, so there's no local-hardware recommendation and no break-even to compute. The honest answer is the per-provider API cost above.

1T total params, 32B active per token (MoE) - 384 experts, 8 selected + 1 shared, 61 layers, 7168 attention hidden dimension, 64 heads, MLA, SwiGLU activation, 160K vocabulary. Built on Kimi K2.6 as a coding/agentic specialization, not a general chat replacement. Context and I/O: native 256K context (262,144 tokens via API); native vision and multimodal tool use (PNG/JPEG/WebP/GIF images, MP4/MOV/AVI video). Reasoning and tools: always-on thinking mode; Moonshot claims ~30% fewer thinking tokens than K2.6. Native function/tool calling with MCP support. Open weights under a Modified MIT license on HuggingFace at moonshotai/Kimi-K2.7-Code. The license requires products with >100M monthly active users or >$20M monthly revenue to display “Kimi K2.7 Code” prominently. Cloud-only for most users. A 1T-parameter MoE needs server-class hardware; there is no published Unsloth GGUF or consumer-grade quantization, so self-hosting is not practical on a workstation or DGX Spark. Run it via the Kimi API, OpenRouter, or vLLM/SGLang on a cluster. Cloud API: $0.95/1M input, $4.00/1M output, $0.19/1M cache-hit input. A kimi-k2.7-code-highspeed variant runs at ~180 tok/s (up to 260 tok/s in short contexts) for 2x the price. Benchmarks (Moonshot self-reported): Kimi Code Bench v2 62.0, Program Bench 53.6, MLS Bench Lite 35.1, Kimi Claw 24/7 Bench 46.9, MCP Atlas 76.0, MCP Mark Verified 81.1.

See the model card for the full architecture notes and any cloud subscription plans.

Full model card API pricing table Generic token calculator