Pareto
premierA composite model that hides its architecture. Pareto is a multimodal composite model (text+image input, text output) for research, coding, and agentic workflows. Unbiased AI does not disclose what is under it: no parameter count, no serving precision, no context-window statement from the vendor itself. The only architectural fact in public is OpenRouter’s API metadata: 262,144-token context window, 131,072-token maximum output, tokenizer “Other”. Treat the architecture as a closed composite - a routed stack, not a single open model.
Pricing, two lanes. The Pareto model on OpenRouter (unbiased/pareto) lists $2.50/M input and $7.50/M output (verified against the OpenRouter API 2026-09-23). Separately, the Pareto Inference serving platform publishes a live price feed for GLM 5.3 Flash on their own GPUs at https://paretoinference.com/api/landing-pricing - as of 2026-09-23 that feed reads $0.03/M input, $0.10/M output, $0.006/M cached input (“80% less than OpenRouter”), and Dylan Vu of Pareto Inference confirmed those direct prepaid rates by email (2026-09-21) with a revision coming. Two different products, two price sheets; check which lane you are buying before comparing.
Not open weights. No HuggingFace presence, no download, no license file. The model is reachable two ways: Pareto Inference’s own GPU serving (OpenAI-compatible API, prepaid credits) and OpenRouter (unbiased/pareto). Tool calling is supported; streaming is not documented in the OpenRouter parameter list.
What to treat as claims. Unbiased markets “5x the intelligence” and makes every AI model “compete for your traffic” - marketing framing, not a benchmark. Third-party eval (benchable.ai) places response times in the 43rd percentile with no independent intelligence-index run yet. The composite composition (which models route inside) is undisclosed.
- proprietary hosted
- 🇺🇸 USA
- Sep 2026
Save your hardware and every model page answers the real question: will it run on your machine, and how fast?
Join free - save your rig →Or run it in the cloud
No per-token API provider pricing tracked for Pareto yet. For flagship list prices, see the calculator.
Inference cost over time
Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.