Ornith-1.5-9B
consumer~9B Dense - the edge-deployable Ornith-1.5, released 2026-08-18. Compact enough for phones (a quantized Ornith-1.5-9B-Mobile variant targets iOS/Android) yet matches or exceeds much larger models like Gemma 4-31B and Qwen 3.6-35B on agentic coding. 262K context, MIT-licensed on HuggingFace at ornith-ai/Ornith-1.5-9B (GGUF and MLX quantizations).
- Coding (vendor self-reported): Terminal-Bench 2.1 46.2, SWE-bench Verified 70.6, SWE-bench Pro 47.5, NL2Repo 32.4.
- Reasoning: HLE 20.2 (no tools) / 30.5 (with tools), GPQA-Diamond 86.4.
- Agentic: MCP-Atlas 54.2, Toolathlon-Verified 41.2, ClawEval 66.5.
Runs almost anywhere. Quantized builds run on phones, laptops, and small GPUs; the strongest open-weight coding model at this size. Vendor benchmarks are claims pending independent replication.
- 9.0B
- 262k
- mit
- πΊπΈ USA
- Aug 2026
Benchmark scores
Vendor-reported - from the developer's own model card / tech report
Vendor-reported - from the developer's own model card / tech report
Ran this model on your own hardware? Join free and add your measured tok/s to the community numbers.
Related models
Save your hardware and every model page answers the real question: will it run on your machine, and how fast?
Join free - save your rig βRun it locally
Per-quant memory needs and a static "can you run it?" reference - no rig entry required
The reference hardware
22 reference configs, drawn in-house. Scroll for more.
Can you run it? - reference rigs
| Rig | Q4_K_M | Q5_K_M | Q6_K | Q8_0 | BF16 |
|---|---|---|---|---|---|
| 4x H100 80GB (320GB) | fast 941.5t/s | fast 848.3t/s | fast 767.1t/s | fast 622.6t/s | fast 360.4t/s |
| NVIDIA DGX Station 748GB | fast 562.1t/s | fast 506.5t/s | fast 458.0t/s | fast 371.7t/s | fast 215.2t/s |
| 8x RTX 3090 rack (192GB) | fast 526.2t/s | fast 474.1t/s | fast 428.7t/s | fast 348.0t/s | fast 201.5t/s |
| 4x RTX 5090 (128GB) | fast 503.6t/s | fast 453.8t/s | fast 410.3t/s | fast 333.0t/s | fast 192.8t/s |
| AMD Instinct MI300X (192GB) | fast 374.1t/s | fast 337.1t/s | fast 304.8t/s | fast 247.4t/s | fast 143.2t/s |
| 4x RTX 4090 (96GB) | fast 283.3t/s | fast 255.3t/s | fast 230.8t/s | fast 187.3t/s | fast 108.5t/s |
| 2x RTX 5090 (64GB) | fast 251.8t/s | fast 226.9t/s | fast 205.2t/s | fast 166.5t/s | fast 96.4t/s |
| 2x RTX 3090 (48GB) | fast 131.6t/s | fast 118.5t/s | fast 107.2t/s | fast 87.0t/s | fast 50.4t/s |
| Single RTX 5090 (32GB) | fast 125.9t/s | fast 113.4t/s | fast 102.6t/s | fast 83.3t/s | fast 48.2t/s |
| RTX PRO 6000 Blackwell (96GB) | fast 125.9t/s | fast 113.4t/s | fast 102.6t/s | fast 83.3t/s | fast 48.2t/s |
| Mac Studio M4 Ultra 192GB | fast 83.7t/s | fast 75.4t/s | fast 68.2t/s | fast 55.4t/s | fast 32.0t/s |
| Mac Studio M4 Ultra 512GB | fast 83.7t/s | fast 75.4t/s | fast 68.2t/s | fast 55.4t/s | fast 32.0t/s |
| Single RTX 4090 (24GB) | fast 70.8t/s | fast 63.8t/s | fast 57.7t/s | fast 46.8t/s | fast 27.1t/s |
| MacBook Pro M5 Max 128GB | fast 47.1t/s | fast 42.4t/s | fast 38.3t/s | fast 31.1t/s | ok 18.0t/s |
| Single GTX 1080 Ti (11GB) | fast 34.0t/s | fast 30.6t/s | fast 27.7t/s | tight | no -> cloud |
| Dual EPYC 9004 + 768GB DDR5-4800 | fast 32.4t/s | fast 29.2t/s | fast 26.4t/s | fast 21.4t/s | ok 12.4t/s |
| DGX Spark 128GB unified | ok 19.2t/s | ok 17.3t/s | ok 15.6t/s | ok 12.7t/s | slow 7.3t/s |
| Ryzen AI Max+ 395 128GB | ok 18.0t/s | ok 16.2t/s | ok 14.7t/s | ok 11.9t/s | slow 6.9t/s |
| Jetson AGX Orin 64GB | ok 14.4t/s | ok 13.0t/s | ok 11.7t/s | ok 9.5t/s | slow 5.5t/s |
| Epyc + 512GB DDR4-3200 + 2x RTX 3090 | ok 14.4t/s | ok 13.0t/s | ok 11.7t/s | ok 9.5t/s | slow 5.5t/s |
| Epyc + 512GB DDR4-2400 + 2x RTX 3090 | ok 10.8t/s | ok 9.7t/s | ok 8.8t/s | slow 7.1t/s | slow 4.1t/s |
| NVIDIA Jetson Orin NX 16GB | slow 7.2t/s | slow 6.5t/s | slow 5.9t/s | slow 4.8t/s | no -> cloud |
Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.
Download options
Or run it in the cloud
No per-token API provider pricing tracked for Ornith-1.5-9B yet. For flagship list prices, see the calculator.
Inference cost over time
Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.