Muse Glimmer 30B
enthusiastMeta’s first open-weights model in the Muse family, released under Apache 2.0. Muse Glimmer 30B is a dense multimodal model with a built-in 1.8B vision encoder and a 128K+ context window. It is designed for coding, agentic workflows, and visual reasoning.
Local deployment. Ollama ships the model day-0 with an MLX-optimized tag for Apple Silicon: muse-glimmer:30b-mlx. Ollama reports that the MLX build runs 1.5x-1.8x faster with DFlash than the baseline implementation. Other common tags include muse-glimmer:30b, muse-glimmer:30b-q4_K_M, muse-glimmer:30b-q8_0, and muse-glimmer:30b-fp16.
Agent launch flows. The Ollama registry tags the model for direct launch with Claude Code, Codex, Pi, Hermes, and other agent tools that use the Ollama API.
Honest framing. Glimmer is positioned as a smaller, open alternative to the proprietary Muse Spark family. The 30B dense size means the full-precision model requires substantial unified memory, but quantized versions fit on high-end Apple Silicon Macs and modern NVIDIA/AMD GPUs. Independent coding and reasoning benchmarks are still rolling in; treat early claims as preliminary until replicated.
- 30.0B
- 128k
- apache 2.0
- 🇺🇸 USA
- Aug 2026
Related models
Guides covering Muse Glimmer 30B
Save your hardware and every model page answers the real question: will it run on your machine, and how fast?
Join free - save your rig →Run it locally
Per-quant memory needs and a static "can you run it?" reference - no rig entry required
The reference hardware
22 reference configs, drawn in-house. Scroll for more.
Can you run it? - reference rigs
| Rig | Q4_K_M | Q8_0 | FP16 |
|---|---|---|---|
| NVIDIA Jetson Orin NX 16GB | no -> cloud | no -> cloud | no -> cloud |
| Single GTX 1080 Ti (11GB) | no -> cloud | no -> cloud | no -> cloud |
| Single RTX 4090 (24GB) | tight | offload | no -> cloud |
| Single RTX 5090 (32GB) | tight | offload | no -> cloud |
| 4x H100 80GB (320GB) | fast 409.4t/s | fast 223.3t/s | fast 122.8t/s |
| NVIDIA DGX Station 748GB | fast 244.4t/s | fast 133.3t/s | fast 73.3t/s |
| 8x RTX 3090 rack (192GB) | fast 228.8t/s | fast 124.8t/s | fast 68.7t/s |
| 4x RTX 5090 (128GB) | fast 219.0t/s | fast 119.5t/s | fast 65.7t/s |
| AMD Instinct MI300X (192GB) | fast 162.7t/s | fast 88.7t/s | fast 48.8t/s |
| 4x RTX 4090 (96GB) | fast 123.2t/s | fast 67.2t/s | tight |
| 2x RTX 5090 (64GB) | fast 109.5t/s | fast 59.7t/s | offload |
| 2x RTX 3090 (48GB) | fast 57.2t/s | fast 31.2t/s | no -> cloud |
| RTX PRO 6000 Blackwell (96GB) | fast 54.8t/s | fast 29.9t/s | tight |
| Mac Studio M4 Ultra 192GB | fast 36.4t/s | ok 19.9t/s | ok 10.9t/s |
| Mac Studio M4 Ultra 512GB | fast 36.4t/s | ok 19.9t/s | ok 10.9t/s |
| MacBook Pro M5 Max 128GB | fast 20.5t/s | ok 11.2t/s | slow 6.1t/s |
| Dual EPYC 9004 + 768GB DDR5-4800 | ok 14.1t/s | slow 7.7t/s | slow 4.2t/s |
| DGX Spark 128GB unified | ok 8.3t/s | slow 4.6t/s | slow 2.5t/s |
| Ryzen AI Max+ 395 128GB | slow 7.8t/s | slow 4.3t/s | slow 2.4t/s |
| Jetson AGX Orin 64GB | slow 6.3t/s | slow 3.4t/s | no -> cloud |
| Epyc + 512GB DDR4-3200 + 2x RTX 3090 | slow 6.3t/s | slow 3.4t/s | slow 1.9t/s |
| Epyc + 512GB DDR4-2400 + 2x RTX 3090 | slow 4.7t/s | slow 2.6t/s | slow 1.4t/s |
Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.
Download options
Or run it in the cloud
Live per-provider pricing, throughput and uptime. Click a column to sort.
| Provider | Type | Input $/M | Output $/M | Cache $/M | Tok/s | Latency | Uptime | Value |
|---|---|---|---|---|---|---|---|---|
| Sub | - | - | - | - | - | - | $20.00/mo Pro | |
| Sub | - | - | - | - | - | - | $100.00/mo Max |
Default order: throughput among 95%+ uptime providers, then latency; subscriptions last. Sort by any column. Subscription rows show $/mo in the Value column - per-token columns are "-". Affiliate links are marked sponsored / nofollow. Confirm current pricing on the provider's site before committing.
Detailed API pricing page + JSON endpoint →
See who runs Meta in production →
Inference cost over time
Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.