Adoption tracker
Who runs which AI model
A curated, source-cited tracker of companies running AI workloads on open-weight models - Llama, Mistral, DeepSeek, Qwen, Kimi, and GLM - as frontier-model prices climb. Every row cites a real source and carries an honest status so a rumor is never shown as a confirmed move.
Statuses: confirmed official source (company blog or CEO statement) reported credible third-party, not officially confirmed testing evaluating, not deployed. Last updated 2026-09-18.
| Company | Model | Vendor | Status | Source | Noted |
|---|---|---|---|---|---|
|
Databricks
Yuchen Jin (Databricks): DeepSeek V4.1 Flash is his pick for easy and moderate tasks - fast and cheap. Databricks reports customers cutting spend dramatically by shifting 20% of internal coding traffic from Claude/GPT to open models.
|
DeepSeek V4.1 Flash | DeepSeek | confirmed | X post | 2026-09-18 |
|
Databricks
Yuchen Jin (Databricks): GLM-5.3 is cheap and fast, suited to high volumes of non-complex coding tasks. Part of Databricks engineers' shift to open models as daily drivers.
|
GLM-5.3 | Z.ai | confirmed | X post | 2026-09-18 |
|
Databricks
Databricks AI systems CTO Yuchen Jin: Kimi K3 is the best open-weight coding model in their experience. Databricks serves it at 239 tok/s, the largest open model they have hosted, and ranks #1 for K3 inference speed on Artificial Analysis.
|
Kimi K3 | Moonshot AI | confirmed | X post | 2026-09-18 |
|
Databricks
Databricks AI systems CTO Yuchen Jin: Kimi K3 is the best open-weight coding model in their experience. Databricks serves it at 239 tok/s, the largest open model they have hosted, and ranks #1 for K3 inference speed on Artificial Analysis.
|
Kimi K3 | Moonshot AI | confirmed | X post | 2026-09-18 |
|
Databricks
Yuchen Jin (Databricks): DeepSeek V4.1 Flash is his pick for easy and moderate tasks - fast and cheap. Databricks reports customers cutting spend dramatically by shifting 20% of internal coding traffic from Claude/GPT to open models.
|
DeepSeek V4.1 Flash | DeepSeek | confirmed | X post | 2026-09-18 |
|
Databricks
Yuchen Jin (Databricks): GLM-5.3 is cheap and fast, suited to high volumes of non-complex coding tasks. Part of Databricks engineers' shift to open models as daily drivers.
|
GLM-5.3 | Z.ai | confirmed | X post | 2026-09-18 |
|
37signals (DHH)
DHH and 37signals ship Omarchy, an Arch-based desktop marketed as beautiful, fun and agentic Linux. It is developed with AI agents as core contributors, ships OpenCode and Claude Code as default coding-agent harnesses, and recommends LM Studio and Ollama for running local open-weight models.
|
Omarchy: agentic Linux built with AI agents | workflow | confirmed | Official | 2026-09-18 |
|
OpenCode
Day-zero support in the OpenCode Go subscription on launch day, with a time-limited 4x usage promo for the first 72 hours (26,000 estimated requests per 5 hours on the $60 tier against 6,500 normally).
|
DeepSeek V4.1 Flash | DeepSeek | confirmed | X post | 2026-09-10 |
|
Microsoft
The Information: engineers evaluating Moonshot AI's Kimi K3 (2.8T open-weight, released 2026-07-16, $3/$15 per MTok) for Copilot features currently on GPT/Claude, citing strong coding benchmarks and ~60% lower inference cost. Not officially confirmed; evaluating, not deployed.
|
Kimi K3 | Moonshot AI | testing | News report | 2026-07-20 |
|
Microsoft
Bloomberg: Microsoft replacing OpenAI/Anthropic models with in-house MAI models in Excel and Outlook to cut inference spend; a tuned MAI variant claims GPT-5.4 parity at up to 10x efficiency. Proprietary (not open-weight), but the same frontier-API cost pressure.
|
MAI | Microsoft | reported | News report | 2026-07-07 |
|
Smartly
Self-hosted Llama 3.1 8B on Kubernetes automates support-ticket creation and resolution drafts for the ad-tech platform; 80% less time to create tickets.
|
Llama 3.1 8B | Meta | confirmed | Engineering blog | 2026-07-06 |
|
Caisse des Depots
Mistral Medium 3.5 (128B) for up to 100k French public-sector agents under a 4-year, EUR 140M framework; on-prem SecNumCloud option for sovereignty.
|
Mistral Medium 3.5 | Mistral AI | confirmed | News report | 2026-07-06 |
|
Capgemini
Self-hosted Codestral in its RAISE/SovBox coding assistant for regulated aerospace, defense and public-sector clients; code-completion accuracy 50% -> 90%.
|
Codestral | Mistral AI | confirmed | Engineering blog | 2026-07-06 |
|
SAP
Self-hosted, 100% European Codestral powers multilingual SBB rail support bots (1k -> 30k employees, 80% fewer repetitive queries) and an accounting accruals agent.
|
Codestral | Mistral AI | confirmed | Engineering blog | 2026-07-06 |
|
Unipol
Multimodal Mistral Pixtral handles image-based insurance claims on the same on-prem NAMI platform as its Llama and Granite workloads.
|
Mistral Pixtral | Mistral AI | confirmed | Engineering blog | 2026-07-06 |
|
Unipol
On-prem IBM Fusion HCI runs Llama (+ Granite + Mistral Pixtral) for the NAMI insurance platform; incident response time 20 min -> 90 sec, monitoring 26% -> 100%.
|
Llama | Meta | confirmed | Engineering blog | 2026-07-06 |
|
ANZ Bank
Ensayo AI platform fine-tunes Llama on API specs and incident history to accelerate software delivery; hybrid on-prem + cloud for data security.
|
Llama 3 | Meta | confirmed | Engineering blog | 2026-07-06 |
|
Goldman Sachs
Hosts Llama variants behind its firewall in the GS AI Platform router (alongside GPT-4o, Gemini, Claude); 1M+ prompts/month, ~20% dev productivity gain.
|
Llama | Meta | confirmed | News report | 2026-07-06 |
|
Perplexity
Serves Llama 3.1 8B / 70B / 405B on NVIDIA HGX H100 with Triton + TensorRT-LLM across 435M+ search queries/month; self-hosting cut cost vs third-party APIs.
|
Llama 3.1 405B | Meta | confirmed | Engineering blog | 2026-07-06 |
|
Uber Eats
Engineering docs describe a global search platform (restaurants, dishes, grocery) built on fine-tuned Qwen across every market. No direct engineering-blog URL found yet.
|
Qwen | Alibaba | reported | News report | 2026-06-29 |
|
Coinbase
Paired with GLM-5.2 in Coinbase's internal LLM gateway for routine tasks (code review, summarization, drafting).
|
Kimi K2.7 Code | Moonshot AI | confirmed | Company statement | 2026-06-28 |
|
Microsoft
Reported to be evaluating DeepSeek V4 for Copilot. Evaluating, not deployed.
|
DeepSeek V4 | DeepSeek | testing | News report | 2026-06-28 |
|
Snowflake
Reported to be testing GLM-5.2 as a cheaper alternative. Evaluating, not deployed.
|
GLM-5.2 | Zhipu AI | testing | News report | 2026-06-28 |
|
Coinbase
CEO Brian Armstrong: routed routine tasks to GLM-5.2 (Z.ai) + Kimi K2.7 Code via an internal gateway with model routing + caching; ~50% AI spend cut.
|
GLM-5.2 | Zhipu AI | confirmed | Company statement | 2026-06-28 |
|
Lindy
Moved 100% of production traffic from Claude to DeepSeek V4 (hosted by Atlas Cloud); ~90% inference cost cut. Still escalates to Claude Opus for edge cases.
|
DeepSeek V4 | DeepSeek | confirmed | Company statement | 2026-06-24 |
|
Trustpilot
Fine-tuned Gemma 2-9B runs on vLLM (A100) for real-time review topic classification, NER and sentiment on millions of reviews/day at a fraction of Gemini cost.
|
Gemma 2-9B | confirmed | Engineering blog | 2026-06-06 | |
|
Shopify
Fine-tuned Qwen3-32B into a tool-calling agent for Shopify Flow / Sidekick (2.2x faster, 68% cheaper). A separate Qwen3 pipeline replaced a GPT-5 one at 75x lower per-unit cost.
|
Qwen3-32B | Alibaba | confirmed | Engineering blog | 2026-04-15 |
|
Grupo Casas Bahia
Llama 3.3 70B on Databricks classifies 33,500 monthly reviews across 91 problem types for the Brazilian retailer; 14x productivity, ~4,000 staff-hours saved/yr.
|
Llama 3.3 70B | Meta | confirmed | News report | 2026-04-06 |
|
Cursor
Composer 2 is built on Kimi K2.5 (Moonshot) + RL fine-tuning, acknowledged by co-founder Aman Sanger. Authorized partnership via Fireworks AI.
|
Kimi K2.5 | Moonshot AI | confirmed | News report | 2026-03-20 |
|
BNY Mellon
First major bank on an on-prem NVIDIA DGX SuperPOD (H100); Llama used for code remediation in the Eliza 2.0 agent platform across 20k+ builders.
|
Llama | Meta | confirmed | News report | 2026-01-16 |
|
Airbnb
CEO Brian Chesky: Qwen powers customer service - 'fast and cheap'. Drew congressional scrutiny over national-security concerns.
|
Qwen | Alibaba | confirmed | Company statement | 2025-12-10 |
Curated and source-cited - not an automated scraper. We research company engineering blogs and X posts, then update this page by hand. Missing a company, or have a correction? tell us.