Agent harness benchmark
Same model, same tasks, different harness. The agent harness you run - Claude Code, Aider, Kimi Code - changes the cost of a task by up to 9x while the success rate barely moves. This leaderboard runs one model across harnesses on a fixed task set and measures what actually matters: how many tasks pass, how many tokens it burns, how long it takes, and what it costs.
Task set
| Harness | Success | Median tokens | Time / task | Cost / task |
|---|---|---|---|---|
|
Kimi Code
composio-k3-28
|
22/28 (79%) | 61k | 297s | $0.22 |
|
Hermes Agent
composio-k3-28
|
21/28 (75%) | 67k | 179s | $0.28 |
|
Claude Code
composio-k3-28
|
20/28 (71%) | 340k | 348s | $2.00 |
Sorted by cost per task, cheapest first. The harness moves cost more than it moves success - that is the whole point. Each row carries the per-token prices used so you can re-derive cost from the token counts.
How we measure
- A fixed task set, identical for every harness. Each harness gets the same prompts and the same pass/fail verification - so a success on harness A means the same thing as a success on harness B. The task set is versioned (e.g.
composio-k3-28); when it changes, old and new runs are never silently compared. - Success, tokens, time, cost - not vibes. A task passes only when its verify check exits 0. Median tokens and median wall-clock time come straight from the harness's own accounting. Cost per task is tokens times the model's per-token price, which is stored on the row for reproducibility.
- Editorial, not crowd-sourced. "Tokens used" is gamifiable in a way decode tok/s (with a physics ceiling) is not, so there is no public submit form. Rows are seeded from published studies (Composio) and our own CLI runs, and each row names its source.
Run it yourself
tokenstead-harness-bench is a small CLI that drives a harness headless against a task set and emits the JSON row above. It is coming soon; this leaderboard will link the one-line install the day it ships. For now, browse the agent harnesses it will drive.