Adoption tracker

Who runs which AI model

A curated, source-cited tracker of companies running AI workloads on open-weight models - Llama, Mistral, DeepSeek, Qwen, Kimi, and GLM - as frontier-model prices climb. Every row cites a real source and carries an honest status so a rumor is never shown as a confirmed move.

Statuses: confirmed official source (company blog or CEO statement) reported credible third-party, not officially confirmed testing evaluating, not deployed. Last updated 2026-09-18.

Company Model Vendor Status Source Noted
Databricks
Yuchen Jin (Databricks): DeepSeek V4.1 Flash is his pick for easy and moderate tasks - fast and cheap. Databricks reports customers cutting spend dramatically by shifting 20% of internal coding traffic from Claude/GPT to open models.
DeepSeek V4.1 Flash DeepSeek confirmed X post 2026-09-18
Databricks
Yuchen Jin (Databricks): GLM-5.3 is cheap and fast, suited to high volumes of non-complex coding tasks. Part of Databricks engineers' shift to open models as daily drivers.
GLM-5.3 Z.ai confirmed X post 2026-09-18
Databricks
Databricks AI systems CTO Yuchen Jin: Kimi K3 is the best open-weight coding model in their experience. Databricks serves it at 239 tok/s, the largest open model they have hosted, and ranks #1 for K3 inference speed on Artificial Analysis.
Kimi K3 Moonshot AI confirmed X post 2026-09-18
Databricks
Databricks AI systems CTO Yuchen Jin: Kimi K3 is the best open-weight coding model in their experience. Databricks serves it at 239 tok/s, the largest open model they have hosted, and ranks #1 for K3 inference speed on Artificial Analysis.
Kimi K3 Moonshot AI confirmed X post 2026-09-18
Databricks
Yuchen Jin (Databricks): DeepSeek V4.1 Flash is his pick for easy and moderate tasks - fast and cheap. Databricks reports customers cutting spend dramatically by shifting 20% of internal coding traffic from Claude/GPT to open models.
DeepSeek V4.1 Flash DeepSeek confirmed X post 2026-09-18
Databricks
Yuchen Jin (Databricks): GLM-5.3 is cheap and fast, suited to high volumes of non-complex coding tasks. Part of Databricks engineers' shift to open models as daily drivers.
GLM-5.3 Z.ai confirmed X post 2026-09-18
37signals (DHH)
DHH and 37signals ship Omarchy, an Arch-based desktop marketed as beautiful, fun and agentic Linux. It is developed with AI agents as core contributors, ships OpenCode and Claude Code as default coding-agent harnesses, and recommends LM Studio and Ollama for running local open-weight models.
Omarchy: agentic Linux built with AI agents workflow confirmed Official 2026-09-18
OpenCode
Day-zero support in the OpenCode Go subscription on launch day, with a time-limited 4x usage promo for the first 72 hours (26,000 estimated requests per 5 hours on the $60 tier against 6,500 normally).
DeepSeek V4.1 Flash DeepSeek confirmed X post 2026-09-10
Microsoft
The Information: engineers evaluating Moonshot AI's Kimi K3 (2.8T open-weight, released 2026-07-16, $3/$15 per MTok) for Copilot features currently on GPT/Claude, citing strong coding benchmarks and ~60% lower inference cost. Not officially confirmed; evaluating, not deployed.
Kimi K3 Moonshot AI testing News report 2026-07-20
Microsoft
Bloomberg: Microsoft replacing OpenAI/Anthropic models with in-house MAI models in Excel and Outlook to cut inference spend; a tuned MAI variant claims GPT-5.4 parity at up to 10x efficiency. Proprietary (not open-weight), but the same frontier-API cost pressure.
MAI Microsoft reported News report 2026-07-07
Smartly
Self-hosted Llama 3.1 8B on Kubernetes automates support-ticket creation and resolution drafts for the ad-tech platform; 80% less time to create tickets.
Llama 3.1 8B Meta confirmed Engineering blog 2026-07-06
Caisse des Depots
Mistral Medium 3.5 (128B) for up to 100k French public-sector agents under a 4-year, EUR 140M framework; on-prem SecNumCloud option for sovereignty.
Mistral Medium 3.5 Mistral AI confirmed News report 2026-07-06
Capgemini
Self-hosted Codestral in its RAISE/SovBox coding assistant for regulated aerospace, defense and public-sector clients; code-completion accuracy 50% -> 90%.
Codestral Mistral AI confirmed Engineering blog 2026-07-06
SAP
Self-hosted, 100% European Codestral powers multilingual SBB rail support bots (1k -> 30k employees, 80% fewer repetitive queries) and an accounting accruals agent.
Codestral Mistral AI confirmed Engineering blog 2026-07-06
Unipol
Multimodal Mistral Pixtral handles image-based insurance claims on the same on-prem NAMI platform as its Llama and Granite workloads.
Mistral Pixtral Mistral AI confirmed Engineering blog 2026-07-06
Unipol
On-prem IBM Fusion HCI runs Llama (+ Granite + Mistral Pixtral) for the NAMI insurance platform; incident response time 20 min -> 90 sec, monitoring 26% -> 100%.
Llama Meta confirmed Engineering blog 2026-07-06
ANZ Bank
Ensayo AI platform fine-tunes Llama on API specs and incident history to accelerate software delivery; hybrid on-prem + cloud for data security.
Llama 3 Meta confirmed Engineering blog 2026-07-06
Goldman Sachs
Hosts Llama variants behind its firewall in the GS AI Platform router (alongside GPT-4o, Gemini, Claude); 1M+ prompts/month, ~20% dev productivity gain.
Llama Meta confirmed News report 2026-07-06
Perplexity
Serves Llama 3.1 8B / 70B / 405B on NVIDIA HGX H100 with Triton + TensorRT-LLM across 435M+ search queries/month; self-hosting cut cost vs third-party APIs.
Llama 3.1 405B Meta confirmed Engineering blog 2026-07-06
Uber Eats
Engineering docs describe a global search platform (restaurants, dishes, grocery) built on fine-tuned Qwen across every market. No direct engineering-blog URL found yet.
Qwen Alibaba reported News report 2026-06-29
Coinbase
Paired with GLM-5.2 in Coinbase's internal LLM gateway for routine tasks (code review, summarization, drafting).
Kimi K2.7 Code Moonshot AI confirmed Company statement 2026-06-28
Microsoft
Reported to be evaluating DeepSeek V4 for Copilot. Evaluating, not deployed.
DeepSeek V4 DeepSeek testing News report 2026-06-28
Snowflake
Reported to be testing GLM-5.2 as a cheaper alternative. Evaluating, not deployed.
GLM-5.2 Zhipu AI testing News report 2026-06-28
Coinbase
CEO Brian Armstrong: routed routine tasks to GLM-5.2 (Z.ai) + Kimi K2.7 Code via an internal gateway with model routing + caching; ~50% AI spend cut.
GLM-5.2 Zhipu AI confirmed Company statement 2026-06-28
Lindy
Moved 100% of production traffic from Claude to DeepSeek V4 (hosted by Atlas Cloud); ~90% inference cost cut. Still escalates to Claude Opus for edge cases.
DeepSeek V4 DeepSeek confirmed Company statement 2026-06-24
Trustpilot
Fine-tuned Gemma 2-9B runs on vLLM (A100) for real-time review topic classification, NER and sentiment on millions of reviews/day at a fraction of Gemini cost.
Gemma 2-9B Google confirmed Engineering blog 2026-06-06
Shopify
Fine-tuned Qwen3-32B into a tool-calling agent for Shopify Flow / Sidekick (2.2x faster, 68% cheaper). A separate Qwen3 pipeline replaced a GPT-5 one at 75x lower per-unit cost.
Qwen3-32B Alibaba confirmed Engineering blog 2026-04-15
Grupo Casas Bahia
Llama 3.3 70B on Databricks classifies 33,500 monthly reviews across 91 problem types for the Brazilian retailer; 14x productivity, ~4,000 staff-hours saved/yr.
Llama 3.3 70B Meta confirmed News report 2026-04-06
Cursor
Composer 2 is built on Kimi K2.5 (Moonshot) + RL fine-tuning, acknowledged by co-founder Aman Sanger. Authorized partnership via Fireworks AI.
Kimi K2.5 Moonshot AI confirmed News report 2026-03-20
BNY Mellon
First major bank on an on-prem NVIDIA DGX SuperPOD (H100); Llama used for code remediation in the Eliza 2.0 agent platform across 20k+ builders.
Llama Meta confirmed News report 2026-01-16
Airbnb
CEO Brian Chesky: Qwen powers customer service - 'fast and cheap'. Drew congressional scrutiny over national-security concerns.
Qwen Alibaba confirmed Company statement 2025-12-10

Curated and source-cited - not an automated scraper. We research company engineering blogs and X posts, then update this page by hand. Missing a company, or have a correction? tell us.