Ruby on Rails just published the second round of its Agents on Rails benchmark - the “Le Mans” round - doubling the field from 8 to 16 models. Every model ran the same 21 realistic Rails tasks three times (63 runs each) behind the same frozen lemans harness and miniswen agent. The result is the clearest public look yet at how frontier models, and one open-weight model you can run yourself, handle real Rails work.
The standout headline for Tokenstead’s audience: an open-weight model you can run on your own hardware finished mid-pack against cloud frontier models. Qwen3.8-27B scored 76.2% - matching Meta’s cloud-only Muse Spark 1.2 and beating GPT-5.6 Luna (73%) and Gemini 3.7 Flash (71.4%). The tradeoffs are time (27 minutes median vs 3 minutes for the fastest cloud models) and Rails API recall (7.9%, the worst in the field). But for a model that fits on a 24GB GPU or Mac, landing at 76% is the strongest argument yet for #supportlocal.
The 16-model leaderboard
Each model ran 63 passes (21 tasks x 3 runs) with the same harness and prompt. Accuracy is the share of runs that passed hidden behavior tests; refusals count as failures.
| Model | Accuracy | Median time | Mean cost/run | API recall |
|---|---|---|---|---|
| Claude Opus 5 | 92.1% | 9m 42s | $1.90 | 34.9% |
| Kimi K3 | 90.5% | 12m 45s | $1.09 | 31.7% |
| Claude Fable 5 | 90.5% | 6m 47s | $2.32 | 33.3% |
| GPT-5.6 Sol | 84.1% | 5m 04s | $0.52 | 31.7% |
| Grok 4.6 | 82.5% | 10m 50s | $0.78 | 33.3% |
| ox-alpha | 82.5% | 19m 26s | free preview | 28.6% |
| GLM 5.3 | 79.4% | 15m 41s | sub | 19.0% |
| Claude Opus 4.8 | 79.4% | 3m 36s | $0.71 | 15.9% |
| GPT-5.6 Terra | 77.8% | 3m 02s | $0.20 | 27.0% |
| Qwen3.8-27B (local) | 76.2% | 27m 04s | $1.23 | 7.9% |
| Muse Spark 1.2 | 76.2% | 15m 44s | $1.69 | 22.2% |
| GPT-5.6 Luna | 73.0% | 3m 19s | $0.014 | 25.4% |
| Gemini 3.7 Flash | 71.4% | 8m 47s | $0.28 | 27.0% |
| Claude Sonnet 5 | 69.8% | 8m 11s | $0.59 | 25.4% |
| GLM 5.2 | 66.7% | 6m 00s | $0.24 | 11.1% |
| DeepSeek V4 Flash | 65.1% | 6m 48s | $0.031 | 12.7% |
The local-vs-hosted story
The most useful read of this round is where open-weight, self-hostable models land against the cloud frontier:
- Qwen3.8-27B (76.2%) is the marquee local story - mid-pack, beating three cloud models, on hardware you may already own. The cost catches are time and API recall.
- GLM 5.2 (66.7%) and DeepSeek V4 Flash (65.1%) are the larger open-weight options; both sit lower but cost pennies per run.
- GLM 5.3 (79.4%) is the strongest open-source result yet, but its weights are still held for safety review (~Aug 28, 2026) - today it is cloud-only.
The pattern: local models now reach the lower-middle of the frontier field. The gap to the top cluster (Opus 5 / Kimi K3 / Fable 5 at 90%+) is still wide, but it is no longer a cliff. For a budget-aware Rails workflow, a local Qwen3.8-27B or GLM 5.2 closes most of the gap at a fixed hardware cost instead of a per-token bill.
What stands out in this round
OpenAI’s three models score in exact price order - Sol 84.1%, Terra 77.8%, Luna 73% - and all three run in the fastest third of the field. Terra at $0.20/run and 3m 02s median is the value pick: near-frontier accuracy at a fifth of the flagship’s cost.
Anthropic is the opposite story. Sonnet 5 (69.8%) is the weakest Anthropic result ever recorded - it reaches for the right Rails API more than Opus 4.8 (25.4% vs 15.9%) yet finishes six runs behind at twice the time. Better recall did not buy the result.
Grok 4.6 and the mystery ox-alpha tie at 82.5%. ox-alpha is an anonymous “stealth” model on OpenRouter - free during its preview, fingerprinting points to Z.ai’s GLM family but no lab has claimed it. If the attribution holds, a mid-tier frontier coder was quietly serving at $0.
Knowing Rails is what separates the models. API recall ranged from 7.9% (Qwen3.8-27B) to 34.9% (Opus 5). The best-recall models cluster at the top; the models that hand-roll fixes (DeepSeek, Qwen) sit at the bottom. When choosing a model for Rails work, the harness you pair it with matters almost as much as the model.
How to use this on Tokenstead
We have folded the Le Mans results into every model card. Look for the Agents on Rails benchmark (Aug 2026, Le Mans round) callout on the pages for Claude Opus 5, Kimi K3, Claude Fable 5, GPT-5.6 Sol, Grok 4.6, ox-alpha, GLM 5.3, Claude Opus 4.8, GPT-5.6 Terra, Qwen3.8-27B, Muse Spark 1.2, GPT-5.6 Luna, Gemini 3.7 Flash, Claude Sonnet 5, GLM 5.2, and DeepSeek V4 Flash 0731.
If you are choosing an agent harness for Rails work, the agent harnesses directory lists the tools that sit between an LLM and your codebase - including lemans, the harness behind every number in this report.
Practical takeaways
- Start with Luna or Terra for exploratory work. $0.014 to $0.20 per run for 73-78% is a remarkable baseline.
- Go local for fixed-cost Rails work. Qwen3.8-27B at 76.2% is the first open-weight model that genuinely competes mid-pack.
- Pay up only for high-stakes work. Opus 5, Kimi K3, and Fable 5 are the safety net when a missed fix is expensive.
- Watch the GLM open-weight release. GLM 5.3’s 79.4% at subscription pricing, plus its weights landing ~Aug 28, makes it the next local candidate to watch.
- Build harnesses that expose Rails APIs. The API-recall gap (7.9% to 34.9%) is the clearest design signal in the whole report.
The live leaderboard and raw runs are at rubyonrails.org/ai and github.com/rails/ai-evals; the round announcement is at rubyonrails.org/2026/8/24/agents-on-rails-lemans.
Discussion
Be the first to commentStart a discussion
Got a take on this, a rig to show off, or a benchmark that says otherwise? Sign up and start the thread - your comment publishes instantly once you're in.