JEV-27B-VL
enthusiastThe decision model that can see. JEV-27B-VL adds a vision encoder to the Jev line: a robot-arm camera frame or a browser screenshot in, a probability for every allowed action out, one forward pass at up to a 262k context. 1.53 million downloads and 2,496 likes make it the most-adopted model of the October decision wave, Apache 2.0 on the Qwen3.5 27B backbone (64 layers, GQA 4x256).
The robot and the browser are the same workload. The card’s own demos show a physical pick-and-place at about 240 milliseconds per decision and headless-Chromium computer use finishing 95 percent of 60 multi-step tasks - both are “look at pixels, score the allowed actions” before a policy layer commits. For agent stacks the practical use is action-gating: the LLM proposes, JEV-VL scores the option set in one pass, the harness acts only on high-confidence scores and asks a bigger model otherwise. 15.9GB of q4 weights plus 256MB per 1,000 tokens of cache lands it on 24GB cards for real contexts.
Where it sits. Jev (flagship reasoning) proved typed decisions at model-scale; Jeff distilled it to 2B; the October wave industrialized the shape, and the -VL variant is the one that closes the loop on agents that act in the world instead of only in text.
- 27.8B
- 262k
- apache 2.0
- 🇺🇸 USA
- Sep 2026
Related models
Guides covering JEV-27B-VL
Save your hardware and every model page answers the real question: will it run on your machine, and how fast?
Join free - save your rig →Run it locally
Per-quant memory needs and a static "can you run it?" reference - no rig entry required
The reference hardware
22 reference configs, drawn in-house. Scroll for more.
Can you run it? - reference rigs
| Rig | Q4_K_M | FP16 |
|---|---|---|
| NVIDIA Jetson Orin NX 16GB | tight | no -> cloud |
| Single GTX 1080 Ti (11GB) | offload | no -> cloud |
| 4x H100 80GB (320GB) | fast 435.5t/s | fast 130.2t/s |
| NVIDIA DGX Station 748GB | fast 260.0t/s | fast 77.7t/s |
| 8x RTX 3090 rack (192GB) | fast 243.4t/s | fast 72.8t/s |
| 4x RTX 5090 (128GB) | fast 233.0t/s | fast 69.6t/s |
| AMD Instinct MI300X (192GB) | fast 173.1t/s | fast 51.7t/s |
| 4x RTX 4090 (96GB) | fast 131.0t/s | fast 39.2t/s |
| 2x RTX 5090 (64GB) | fast 116.5t/s | fast 34.8t/s |
| 2x RTX 3090 (48GB) | fast 60.9t/s | offload |
| Single RTX 5090 (32GB) | fast 58.2t/s | no -> cloud |
| RTX PRO 6000 Blackwell (96GB) | fast 58.2t/s | ok 17.4t/s |
| Mac Studio M4 Ultra 192GB | fast 38.7t/s | ok 11.6t/s |
| Mac Studio M4 Ultra 512GB | fast 38.7t/s | ok 11.6t/s |
| Single RTX 4090 (24GB) | fast 32.8t/s | no -> cloud |
| MacBook Pro M5 Max 128GB | fast 21.8t/s | slow 6.5t/s |
| Dual EPYC 9004 + 768GB DDR5-4800 | ok 15.0t/s | slow 4.5t/s |
| DGX Spark 128GB unified | ok 8.9t/s | slow 2.7t/s |
| Ryzen AI Max+ 395 128GB | ok 8.3t/s | slow 2.5t/s |
| Jetson AGX Orin 64GB | slow 6.7t/s | slow 2.0t/s |
| Epyc + 512GB DDR4-3200 + 2x RTX 3090 | slow 6.7t/s | slow 2.0t/s |
| Epyc + 512GB DDR4-2400 + 2x RTX 3090 | slow 5.0t/s | slow 1.5t/s |
Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.
Download options
Or run it in the cloud
No per-token API provider pricing tracked for JEV-27B-VL yet. For flagship list prices, see the calculator.
Inference cost over time
Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.