Models / Ember-1

An inference provider just shipped its own model. Ember-1 is Fireworks Research’s first model: Kimi K3 retrained to cut unnecessary reasoning while keeping the thinking that matters. It is the first shipped reasoning-token-diet model - same claimed quality as the base model, ~40% fewer tokens - and it was trained end to end on Fireworks’ own serverless training platform with more than 50 experiments and 200 evaluations. That is the signal: the hosting layer is becoming a model lab, because it sees exactly where token spend hurts.

Why retrain instead of turning effort down. K3’s reasoning effort settings trade quality for cost crudely: lowering effort gives up accuracy. The vendor’s experiments say 35-50% of K3’s reasoning is removable without touching answers, but only by training the model to keep productive self-reflection while escaping unproductive loops. The internal test table: K3 scores 0.751 on their harness with 49.3K output tokens; Ember-1 scores 0.753 with 29.9K (reasoning tokens down 71.3%, total down 39%).

The agentic math compounds. Reasoning models spend up to 90%+ of generated tokens on thinking, and in multi-turn agent work every turn replays prior reasoning to the model, so context and cost grow roughly quadratically with turns. Halving per-turn reasoning compounds across the loop. On the vendor’s benchmarks (all vendor-run until replicated): DeepSWE 75.2 vs K3-max’s 66.4 is the standout; SWE-bench Verified 92.2 vs 93.2 and Terminal-Bench 82.0 vs 80.9 hold quality; cost per task drops 5.9-51.9% depending on benchmark. Two customers’ live production A/B tests showed ~35% fewer tokens per task; one is moving Ember-1 into full production.

Availability and limits. API-only as a Research Preview on Fireworks Serverless, no open weights, which caps its modeldex value: you cannot run it locally or verify the vendor numbers yourself. The pricing context sharpens it - at K3’s rates ($3/M uncached input, $0.30/M cached, $15/M output) a 40% token cut on coding agents is a direct bill cut. Fireworks says Ember is the start of a series and is opening two-week research releases of future models.

agentic coding efficiency hosted
License
proprietary
Developer
Fireworks Research
Origin
🇺🇸 USA
Released
Sep 2026

Scores

Coding
92
Reasoning
88
Tool calling
88
General
85

Save your hardware and every model page answers the real question: will it run on your machine, and how fast?

Join free - save your rig →

Or run it in the cloud

No per-token API provider pricing tracked for Ember-1 yet. For flagship list prices, see the calculator.

PRICE HISTORY

Inference cost over time

Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.

Loading price history...