Leaderboard

Agent harness benchmark

Same model, same tasks, different harness. The agent harness you run - Claude Code, Aider, Kimi Code - changes the cost of a task by up to 9x while the success rate barely moves. This leaderboard runs one model across harnesses on a fixed task set and measures what actually matters: how many tasks pass, how many tokens it burns, how long it takes, and what it costs.

Model

Task set

Harness Success Median tokens Time / task Cost / task
Kimi Code
composio-k3-28
22/28 (79%) 61k 297s $0.22
Hermes Agent
composio-k3-28
21/28 (75%) 67k 179s $0.28
Claude Code
composio-k3-28
20/28 (71%) 340k 348s $2.00

Sorted by cost per task, cheapest first. The harness moves cost more than it moves success - that is the whole point. Each row carries the per-token prices used so you can re-derive cost from the token counts.

How we measure

  • A fixed task set, identical for every harness. Each harness gets the same prompts and the same pass/fail verification - so a success on harness A means the same thing as a success on harness B. The task set is versioned (e.g. composio-k3-28); when it changes, old and new runs are never silently compared.
  • Success, tokens, time, cost - not vibes. A task passes only when its verify check exits 0. Median tokens and median wall-clock time come straight from the harness's own accounting. Cost per task is tokens times the model's per-token price, which is stored on the row for reproducibility.
  • Editorial, not crowd-sourced. "Tokens used" is gamifiable in a way decode tok/s (with a physics ceiling) is not, so there is no public submit form. Rows are seeded from published studies (Composio) and our own CLI runs, and each row names its source.

Run it yourself

tokenstead-harness-bench is a small CLI that drives a harness headless against a task set and emits the JSON row above. It is coming soon; this leaderboard will link the one-line install the day it ships. For now, browse the agent harnesses it will drive.