FreeToken: Berkeley edge-MoE engine runs DeepSeek at home

Published Oct 05, 2026

FreeToken: Berkeley’s edge-MoE engine serves DeepSeek on your desktop

Meta description (draft): FreeToken (Stoica’s group, Apache 2.0) serves DeepSeek V4 Flash at 22 to 25 tok/s on a 5090 and a 35B model at 39 tok/s on an 8GB laptop GPU. —

The local-MoE engine race has a lab-backed entrant now. FreeToken (FlashML-org on GitHub, Apache-2.0, 14,189 stars, pushed daily since July 20) comes from a Berkeley-affiliated group - Shuo Yang, Melissa Pan, Kurt Keutzer, Song Han, Matei Zaharia, Ion Stoica among the authors - with a systems paper, arXiv 2608.16157, that reads differently from the hobbyist-engineering approach of Strata: FreeToken co-designs the whole serving stack (layout, expert residency, CPU-GPU co-execution, caching) so the engine re-maps work onto whatever hardware it actually finds. The pitch: treat your personal machine not as a small GPU but as a unified, elastic inference platform.

The paper’s evaluated numbers, pulled from the full text (not the README): on an RTX 5090, 77 to 83 tok/s on Qwen3.6-35B-A3B and 22 to 25 tok/s on DeepSeek-V4-Flash (FP4), which is 1.5 to 2.3x the decode of the edges it compares to (llama.cpp, Ollama, KTransformers, MoE-Infinity) across four real agentic workloads. On an 8GB RTX 4060 laptop, it serves the 35B model at 39.3 tok/s - faster than the 33 tok/s median decode the Codex service itself posts. On one RTX PRO 6000, it serves the 753B GLM-5.2 at twice llama.cpp’s throughput. Across five consumer systems the lift is 1.3 to 2.1x.

Decode lift versus the field, with the two numbers that matter: 39.3 tok/s from an 8GB laptop, and tail TTFT that never exceeds 44s while every baseline blows past 150s somewhere.

Decode lift versus the field, with the two numbers that matter: 39.3 tok/s from an 8GB laptop, and tail TTFT that never exceeds 44s while every baseline blows past 150s somewhere.

The two findings that distinguish it

First: the agentic-degradation problem has an engineering answer. Every local engine dies the same way when an agent starts working: tool calls mutate context, prefill re-runs, caches thrash, TTFT climbs past the client’s timeout on the seventh tool call. FreeToken’s semantic anchor checkpoints let agentic context edits (tool calls, thinking blocks) skip redundant recomputation, and its decode rate stays within 12 percent of the single-turn setting while agentic workloads degrade every baseline it measured, each of which exceeds 150s worst-case TTFT in at least one setting. For people running Claude Code or Codex against local backends, that tail latency is the difference between usable and not.

Second: runtime elasticity. FreeToken re-allocates VRAM between expert cache and KV memory at runtime without restarting or reloading weights, and its bandwidth-adaptive policy picks where each expert runs (GPU, CPU, or streamed from RAM) based on what the machine actually exposes. The cost line it publishes honestly: swapping in DeepSeek-V4-Flash’s FP4 expert pool moves roughly 140GB - about 2 seconds on a PCIe 5.0 5090, 5 seconds on 4090/3090-class PCIe 4.0 boxes, 10 or more on the rest.

Four engines now serve server-class weights on consumer boxes: FreeToken (Jul), Strata (Sep 24), ds4 (mid-Sep), llama.cpp’s decision-model endpoint (Oct 2).

Four engines now serve server-class weights on consumer boxes: FreeToken (Jul), Strata (Sep 24), ds4 (mid-Sep), llama.cpp’s decision-model endpoint (Oct 2).

How to run it today

A desktop app for Windows and Linux (flashml.ai) with a GUI, or uv pip install "freetoken[accel]". Model support matters for this site’s lanes: DeepSeek-V4-Flash-0731, GLM-5.3-Flash NVFP4, GLM-5.2, Qwen3.8-Flash-Next in FP8 and NVFP4, plus the dense Qwen3.8-27B class and gpt-oss-120b - over 20 MoE models across MXFP4, NVFP4, FP8, and BF16, with OpenAI- and Anthropic-compatible endpoints so coding agents plug in unmodified. Where Strata wins on quant-picking for a single model, FreeToken wins on breadth of models and on stability under agentic load; the two are complementary, not competing, for most builds.

The headline numbers are the authors’ own benches (the repo publishes no third-party telemetry yet, unlike Strata’s community folder); DV4-Flash’s 22 to 25 tok/s is interactive for single-stream coding but not speed-demon territory, and the 4060-laptop figure is the friendliest cell in the paper. The direction, though, is unmistakable: two serious engines now agree that a consumer box is a platform for 100B-plus weights, and the paper’s authors are the people behind some of the most-deployed serving systems in the field.

Sources: FreeToken repo - arXiv 2608.16157 - flashml.ai - HN thread on Strata, the sibling approach

Related on this site: Strata runs Flash-Next on your gaming GPU - Qwen3.8-Flash-Next - DeepSeek V4.1 Flash - The only hard budget cap on agents is a box you own

Discussion

Be the first to comment

Start a discussion

Got a take on this, a rig to show off, or a benchmark that says otherwise? Sign up and start the thread - your comment publishes instantly once you're in.