Models / pplx-decider-v1.1-27b

pplx-decider-v1.1-27b

enthusiast

The search company’s router, released to everyone. Perplexity Labs shipped pplx-decider-v1 on October 1, 2026 and a v1.1 refresh on October 5: a 26.1B-parameter text-classification model, Apache 2.0, on the Qwen3.5 27B backbone the rest of the week’s decision models also tune. The contract is the category’s: text in, one label out of a fixed set, confidence number attached, one forward pass, no generated tokens. GGUF builds from the community quantizers (mradermacher covers 12B-class siblings of the same wave) appeared within days.

What Perplexity needs it for tells you where decision models are going. A search engine routes every query: web search vs math vs translation vs shopping vs a clarifying question, in milliseconds, on every request, before the expensive model ever loads. Decider is that router open-sourced - the internal traffic-shape classifier became the product. For local operators the same play applies at home: query routing for a local agent stack, intent gating for voice pipelines, ticket triage for a helpdesk, all on a card that fits in 16GB in 4-bit.

Where it sits in the wave. The October decision-model flood (Clef by Cloudflare, GEV by autotrust, Torchcast 12B) landed within seven days; Decider’s v1-to-v1.1 turnaround inside the same week signals fast iteration, and the Apache license keeps it forkable. Unlike Clef it is text-only - screenshots stay with the vision-tuned members of the family.

decision-model classification routing
Parameters
26.1B
Context
262k
License
apache 2.0
Developer
Perplexity
Origin
🇺🇸 USA
Released
Oct 2026

Guides covering pplx-decider-v1.1-27b

Save your hardware and every model page answers the real question: will it run on your machine, and how fast?

Join free - save your rig →

Run it locally

Per-quant memory needs and a static "can you run it?" reference - no rig entry required

Q4_K_M
14.9GB 16.0GB min 24.0GB rec
Balanced - the usual local sweet spot
FP16
52.2GB 53.0GB min 64.0GB rec
Full quality, largest

The reference hardware

Schematic of the NVIDIA Jetson Orin NX 16GB reference rig - 16GB unified memory, 102 GB/s aggregate bandwidth
Schematic of the Single GTX 1080 Ti (11GB) reference rig - 11GB VRAM, 484 GB/s aggregate bandwidth
Schematic of the 4x H100 80GB (320GB) reference rig - 320GB VRAM, 13400 GB/s aggregate bandwidth
Schematic of the NVIDIA DGX Station 748GB reference rig - 748GB unified memory, 8000 GB/s aggregate bandwidth
Schematic of the 8x RTX 3090 rack (192GB) reference rig - 192GB VRAM, 7489 GB/s aggregate bandwidth
Schematic of the 4x RTX 5090 (128GB) reference rig - 128GB VRAM, 7168 GB/s aggregate bandwidth
Schematic of the AMD Instinct MI300X (192GB) reference rig - 192GB VRAM, 5324 GB/s aggregate bandwidth
Schematic of the 4x RTX 4090 (96GB) reference rig - 96GB VRAM, 4032 GB/s aggregate bandwidth
Schematic of the 2x RTX 5090 (64GB) reference rig - 64GB VRAM, 3584 GB/s aggregate bandwidth
Schematic of the 2x RTX 3090 (48GB) reference rig - 48GB VRAM, 1872 GB/s aggregate bandwidth
Schematic of the Single RTX 5090 (32GB) reference rig - 32GB VRAM, 1792 GB/s aggregate bandwidth
Schematic of the RTX PRO 6000 Blackwell (96GB) reference rig - 96GB VRAM, 1792 GB/s aggregate bandwidth
Schematic of the Mac Studio M4 Ultra 192GB reference rig - 192GB unified memory, 1092 GB/s aggregate bandwidth
Schematic of the Mac Studio M4 Ultra 512GB reference rig - 512GB unified memory, 1092 GB/s aggregate bandwidth
Schematic of the Single RTX 4090 (24GB) reference rig - 24GB VRAM, 1008 GB/s aggregate bandwidth
Schematic of the MacBook Pro M5 Max 128GB reference rig - 128GB unified memory, 614 GB/s aggregate bandwidth
Schematic of the Dual EPYC 9004 + 768GB DDR5-4800 reference rig - 768GB unified memory, 460 GB/s aggregate bandwidth
Schematic of the DGX Spark 128GB unified reference rig - 128GB unified memory, 273 GB/s aggregate bandwidth
Schematic of the Ryzen AI Max+ 395 128GB reference rig - 128GB unified memory, 256 GB/s aggregate bandwidth
Schematic of the Jetson AGX Orin 64GB reference rig - 64GB unified memory, 204 GB/s aggregate bandwidth
Schematic of the Epyc + 512GB DDR4-3200 + 2x RTX 3090 reference rig - 560GB unified memory, 204 GB/s aggregate bandwidth
Schematic of the Epyc + 512GB DDR4-2400 + 2x RTX 3090 reference rig - 560GB unified memory, 153 GB/s aggregate bandwidth

22 reference configs, drawn in-house. Scroll for more.

Can you run it? - reference rigs

Rig Q4_K_M FP16
NVIDIA Jetson Orin NX 16GB tight no -> cloud
Single GTX 1080 Ti (11GB) offload no -> cloud
4x H100 80GB (320GB) fast 462.8t/s fast 138.5t/s
NVIDIA DGX Station 748GB fast 276.3t/s fast 82.7t/s
8x RTX 3090 rack (192GB) fast 258.7t/s fast 77.4t/s
4x RTX 5090 (128GB) fast 247.6t/s fast 74.1t/s
AMD Instinct MI300X (192GB) fast 183.9t/s fast 55.0t/s
4x RTX 4090 (96GB) fast 139.3t/s fast 41.7t/s
2x RTX 5090 (64GB) fast 123.8t/s fast 37.0t/s
2x RTX 3090 (48GB) fast 64.7t/s offload
Single RTX 5090 (32GB) fast 61.9t/s offload
RTX PRO 6000 Blackwell (96GB) fast 61.9t/s ok 18.5t/s
Mac Studio M4 Ultra 192GB fast 41.2t/s ok 12.3t/s
Mac Studio M4 Ultra 512GB fast 41.2t/s ok 12.3t/s
Single RTX 4090 (24GB) fast 34.8t/s no -> cloud
MacBook Pro M5 Max 128GB fast 23.1t/s slow 6.9t/s
Dual EPYC 9004 + 768GB DDR5-4800 ok 15.9t/s slow 4.8t/s
DGX Spark 128GB unified ok 9.4t/s slow 2.8t/s
Ryzen AI Max+ 395 128GB ok 8.8t/s slow 2.7t/s
Jetson AGX Orin 64GB slow 7.1t/s slow 2.1t/s
Epyc + 512GB DDR4-3200 + 2x RTX 3090 slow 7.1t/s slow 2.1t/s
Epyc + 512GB DDR4-2400 + 2x RTX 3090 slow 5.3t/s slow 1.6t/s

Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.

Download options

Q4_K_M community
14.9GB dl 16.0GB min 24.0GB rec
REC RAM vs largest quant
14.9GB q4 weights + KV 256MB/1k (64L x 4 kv x 256); 16GB card fits routing ctx, long inputs want 24GB
FP16 official
52.2GB dl 53.0GB min 64.0GB rec
REC RAM vs largest quant
52.2GB fp16 weights + KV 256MB/1k

Or run it in the cloud

No per-token API provider pricing tracked for pplx-decider-v1.1-27b yet. For flagship list prices, see the calculator.

PRICE HISTORY

Inference cost over time

Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.

Loading price history...