Ornith-1.5-35B-A3B
MoE enthusiast~36B total, ~3B active per token (MoE) - the local-favorite Ornith-1.5 size, released 2026-08-18. Activates only 3B parameters per token yet outperforms dense models like Gemma 4-31B and Qwen 3.6-35B on agentic coding. 262K context, MIT-licensed on HuggingFace at ornith-ai/Ornith-1.5-35B-A3B (GGUF, FP8, and NVFP4 quantizations).
- Coding (vendor self-reported): Terminal-Bench 2.1 67.8, SWE-bench Verified 79.0, SWE-bench Pro 59.6, NL2Repo 46.2.
- Reasoning: HLE 25.6 (no tools) / 33.4 (with tools), GPQA-Diamond 89.2.
- Agentic: MCP-Atlas 70.2, Toolathlon-Verified 48.7, ClawEval 72.5.
Local-friendly. Runs on enthusiast-class GPUs via GGUF quantizations; a strong open-weight pick for local agentic coding. Vendor benchmarks are claims pending independent replication.
- 36.0B
- 262k
- mit
- πΊπΈ USA
- Aug 2026
Benchmark scores
Vendor-reported - from the developer's own model card / tech report
Vendor-reported - from the developer's own model card / tech report
Ran this model on your own hardware? Join free and add your measured tok/s to the community numbers.
Related models
Save your hardware and every model page answers the real question: will it run on your machine, and how fast?
Join free - save your rig βRun it locally
Per-quant memory needs and a static "can you run it?" reference - no rig entry required
The reference hardware
22 reference configs, drawn in-house. Scroll for more.
Can you run it? - reference rigs
| Rig | Q4_K_M | Q5_K_M | Q6_K | Q8_0 | BF16 |
|---|---|---|---|---|---|
| NVIDIA Jetson Orin NX 16GB | no -> cloud | no -> cloud | no -> cloud | no -> cloud | no -> cloud |
| Single GTX 1080 Ti (11GB) | no -> cloud | no -> cloud | no -> cloud | no -> cloud | no -> cloud |
| 4x H100 80GB (320GB) | fast 1815.1t/s | fast 1690.2t/s | fast 1573.0t/s | fast 1364.3t/s | fast 901.3t/s |
| NVIDIA DGX Station 748GB | fast 1083.7t/s | fast 1009.1t/s | fast 939.1t/s | fast 814.5t/s | fast 538.1t/s |
| 8x RTX 3090 rack (192GB) | fast 1014.5t/s | fast 944.7t/s | fast 879.2t/s | fast 762.6t/s | fast 503.8t/s |
| 4x RTX 5090 (128GB) | fast 971.0t/s | fast 904.2t/s | fast 841.4t/s | fast 729.8t/s | fast 482.1t/s |
| AMD Instinct MI300X (192GB) | fast 721.3t/s | fast 671.7t/s | fast 625.1t/s | fast 542.1t/s | fast 358.2t/s |
| 4x RTX 4090 (96GB) | fast 546.2t/s | fast 508.6t/s | fast 473.3t/s | fast 410.5t/s | fast 271.2t/s |
| 2x RTX 5090 (64GB) | fast 485.5t/s | fast 452.1t/s | fast 420.7t/s | fast 364.9t/s | offload |
| 2x RTX 3090 (48GB) | fast 253.6t/s | fast 236.2t/s | fast 219.8t/s | fast 190.6t/s | offload |
| Single RTX 5090 (32GB) | fast 242.7t/s | fast 226.0t/s | fast 210.4t/s | offload | no -> cloud |
| RTX PRO 6000 Blackwell (96GB) | fast 242.7t/s | fast 226.0t/s | fast 210.4t/s | fast 182.5t/s | fast 120.5t/s |
| Mac Studio M4 Ultra 192GB | fast 161.4t/s | fast 150.3t/s | fast 139.8t/s | fast 121.3t/s | fast 80.1t/s |
| Mac Studio M4 Ultra 512GB | fast 161.4t/s | fast 150.3t/s | fast 139.8t/s | fast 121.3t/s | fast 80.1t/s |
| Single RTX 4090 (24GB) | fast 136.5t/s | offload | offload | offload | no -> cloud |
| MacBook Pro M5 Max 128GB | fast 90.7t/s | fast 84.5t/s | fast 78.6t/s | fast 68.2t/s | fast 45.1t/s |
| Dual EPYC 9004 + 768GB DDR5-4800 | fast 62.4t/s | fast 58.1t/s | fast 54.1t/s | fast 46.9t/s | fast 31.0t/s |
| DGX Spark 128GB unified | fast 37.0t/s | fast 34.4t/s | fast 32.1t/s | fast 27.8t/s | ok 18.4t/s |
| Ryzen AI Max+ 395 128GB | fast 34.7t/s | fast 32.3t/s | fast 30.1t/s | fast 26.1t/s | ok 17.2t/s |
| Jetson AGX Orin 64GB | fast 27.7t/s | fast 25.8t/s | fast 24.0t/s | fast 20.9t/s | no -> cloud |
| Epyc + 512GB DDR4-3200 + 2x RTX 3090 | fast 27.7t/s | fast 25.8t/s | fast 24.0t/s | fast 20.9t/s | ok 13.8t/s |
| Epyc + 512GB DDR4-2400 + 2x RTX 3090 | fast 20.8t/s | ok 19.4t/s | ok 18.0t/s | ok 15.6t/s | ok 10.3t/s |
Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.
Download options
Or run it in the cloud
No per-token API provider pricing tracked for Ornith-1.5-35B-A3B yet. For flagship list prices, see the calculator.
Inference cost over time
Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.