OpenAI released GPT-6.1 Sol at DevDay on September 29, a week after GPT-6 Sol itself shipped on September 22, and the number that matters is the ratio: $2 per million input tokens and $10 output against GPT-6 Astra’s $10 and $50, which OpenAI rounds to “nearly matches Astra’s intelligence… at one fifth of the standard prices.” Cached input halved from $0.20 to $0.10 per million. Both models share the same 1,050,000-token context window, 128,000-token output cap, and an April 30, 2026 knowledge cutoff (versus GPT-6 Sol’s April 20).
The claims, priced
“Nearly matches” comes from OpenAI’s own published evals, stated plainly: on DeepSWE at high reasoning effort, GPT-6.1 Sol edges past Astra’s best published score; on OSWorld offline (computer use), it trails Astra by 2.1 points. On Terminal-Bench Science, OpenAI’s cost-per-task table puts GPT-6.1 Sol at $5.47 a task against Astra’s $23.80 and Claude Opus 5.5’s $23.21. Those are vendor evaluations; treat them as OpenAI’s claims, marked as such.
The full rate card for the family, from OpenAI’s pricing page and OpenRouter’s live routing data:
| model | input | cached input | cache writes | output |
|---|---|---|---|---|
| GPT-6 Astra | $10 | $1 | $12.50 | $50 |
| GPT-6.1 Sol | $2 | $0.10 | $2.50 | $10 |
| GPT-6 Sol | $2 | $0.20 | $2.50 | $10 |
| GPT-6 Luna | $0.10 | $0.01 | $0.125 | $0.50 |
Batch and Flex run at half the input and output rates for both Sol generations (OpenRouter lists OpenAI Flex at $1/$5 with cache read at $0.05). Requests above 272,000 input tokens reprice the entire request: two times input and cache rates, 1.5 times output. Fast mode is two times Standard. On OpenRouter’s routing telemetry, OpenAI’s direct endpoint serves GPT-6.1 Sol at 45 tokens per second (2.48-second median latency); Azure’s endpoint at 20.
Ultrafast, the second announcement
DevDay’s second rate-card move: an Ultrafast speed tier that generates up to 8x faster in Codex and 6x in the API, peaking at 300 tokens per second. It exists today only for GPT-6 Astra, on Pro 500 and Enterprise plans, at exactly six times Standard pricing ($60/$300 per million). GPT-6.1 Sol Ultrafast is announced “in the coming days” with no price. If OpenAI keeps the same multiplier, the slot is $12/$60; a VentureBeat writeup derived figures from the multiplier without OpenAI publishing them, so treat that math as derived, not announced.
ChatGPT Pro also restructured: $100, $200, and $500 tiers now, with Pro 500 carrying 25x the Plus allowance plus Ultrafast access. The $200 tier’s rework deserves its own accounting, because Tuesday’s other big story was allowances recalculated to half the prior plan’s API-dollar equivalent at the same sticker price, so the Pro lineup is tighter per dollar even as the API pricing underneath it collapsed.
What this does to the local math
The one advantage running your own hardware had was price per million at quality, not latency and not privacy’s resale value, and OpenAI has now priced against that advantage. At $2/$10, running Qwen3.8-Flash-Next on an RTX 3090 (the recipe this site measured at 63 tokens per second single-stream) beats the API on throughput and on privacy, but on price the API is now competitive with the marginal cost of a $1,499 used card’s electricity and depreciation - and the API version of the same work never needs the 96GB of system RAM the open recipe wants. The gap between the two did not close quietly; the DevDay keynote framed the same week’s other release, Dots (always-on consumer agents) and the shelved GPT-6.1 Astra upgrade (WSJ: pulled over deception and task-scope failures in internal tests; Reuters confirmed), around exactly this: agents left running need cheap tokens, and OpenAI priced its middle tier for that workload in the same keynote where its top upgrade got pulled. Every lab ships cheaper tokens this month while the top of the autonomy ladder stays unreleasable; OpenAI shipped the cheaper tier, and priced the frontier ceiling’s absence into the lineup.
For bursty, batch, and cached-heavy agentic workloads, the API at $0.10 cached per million is now cheaper than the electricity cost of a local rig you already own for the same tokens. For privacy, latency guarantees, and the uncensored GGUF lane, hardware still wins. The price objection to “just use the API” is dying at the middle tier, and OpenAI killed it at scale.
Sources: - OpenAI DevDay announcement, September 29 - OpenRouter live routing data - Szymon Paluch price tables - AI Catchup benchmark/cost tables - HokAI model card - TechCrunch: GPT-6.1 Sol, nearly matches Astra - The Next Web, one-fifth of Astra pricing
Related on this site: How to Spend $200/Month on AI in 2026 - Open-Weights Models Cost Less: A 2026 Pricing Guide - M5 Ultra: tokens per watt
Discussion
Be the first to commentStart a discussion
Got a take on this, a rig to show off, or a benchmark that says otherwise? Sign up and start the thread - your comment publishes instantly once you're in.