All five models passed all 17 hidden ticket tests in real app work.

A ticket run cost $0.195 to $0.958 and took 1.1 to 9.1 minutes, in a small TypeScript app. Ten models also explained one short Python file.

Real app work#

GPT-6.1 Sol was cheapest on the tickets test at $0.195 a run, and Claude Sonnet 5.5 was fastest at 1.1 minutes.

Each dot is a model's median run. Every model scored at or near full marks, so these tests show cost and time on a small job, not which model is best.

  • Claude models in Claude Code
  • GPT models in Codex CLI
Bug hunt
$0.00$0.25$0.50$0.75$1.00$1.250246810Cost per run at API pricesMinutes to finishGPT-6.1 SolGPT-6 SolGPT-5.6 SolOpus 5.5Sonnet 5.5, 19 of 20
Tickets
$0.00$0.25$0.50$0.75$1.00$1.250246810Cost per run at API pricesMinutes to finishGPT-6.1 SolSonnet 5.5GPT-6 SolGPT-5.6 SolOpus 5.5

Round 13 reran all 30 runs with the account extras off. Every median score held, costs moved by up to 61% and four of the five models took at least 1.4 times as long.

Source: skill-evals.brenhq.com, updated 6 October 2026. 3 runs per model on each test, graded by hidden tests, with the account extras off.

Bug hunt, each run's score and the mediansBest score first, then cheapest. Click a column to sort.
Bug hunt, each run's score and the medians
Model
GPT‑6.1 SolGPT‑6.1 SolCodex CLI 0.159.0, medium effort20, 20, 20$0.2238.2 min2182
GPT‑6 SolGPT‑6 SolCodex CLI 0.159.0, medium effort20, 20, 20$0.2808.5 min1264
GPT‑5.6 SolGPT‑5.6 SolCodex CLI 0.159.0, medium effort20, 20, 20$0.5748.6 min1454
Claude Opus 5.5Opus 5.5Claude Code 2.1.284, medium effort20, 20, 20$0.7822.6 min2584
Claude Sonnet 5.5Sonnet 5.5Claude Code 2.1.284, medium effort19, 19, 19$0.1510.4 min040
Tickets, each run's score and the mediansBest score first, then cheapest. Click a column to sort.
Tickets, each run's score and the medians
Model
GPT‑6.1 SolGPT‑6.1 SolCodex CLI 0.159.0, medium effort17, 17, 17$0.1957.6 min11166
Claude Sonnet 5.5Sonnet 5.5Claude Code 2.1.284, medium effort17, 17, 17$0.2661.1 min6185
GPT‑6 SolGPT‑6 SolCodex CLI 0.159.0, medium effort17, 17, 17$0.2726.9 min5168
GPT‑5.6 SolGPT‑5.6 SolCodex CLI 0.159.0, medium effort17, 17, 17$0.5419.1 min6206
Claude Opus 5.5Opus 5.5Claude Code 2.1.284, medium effort17, 17, 17$0.9583.2 min23252

Tests written counts the new test cases a run added. Lines changed counts the app code it added or removed. Both are medians of the runs. Neither is a score. More tests or fewer lines isn't better by itself, but both show how much a reviewer has to check.

How I tested

Every model got the same small TypeScript web app, with an API, a React front end and a written spec. Hidden tests the model never saw graded the result. The bug hunt plants 20 bugs across the front end, the API and shared logic. The tickets test asks for weekly repeating bookings, a fix for a double booking reported only by its symptom, and a fix for a security hole, checked by 17 hidden tests. Each model ran each test 3 times with no skill. Claude Opus 5.5 planted half the bugs and GPT-6.1 Sol the other half, and GPT-6.1 Sol reviewed the ticket tests for fairness. The apps and the tests stay private, so no model can train on them. The app is small, so a full pass says little about harder code.

The explain task#

Across 10 models on the explain task, the priciest cost 116x as much per run as the cheapest, and the slowest took 2.9x as long as the fastest§.

Each model ran with no skill in Claude Code or Codex CLI, read the same short Python file and explained it. Cost is at each model's published API prices.

  • Claude models in Claude Code
  • GPT models in Codex CLI

The table scrolls sideways on a small screen.

Model
Claude Code 2.1.284, medium effort8.1s33,00281%$0.0389
Codex CLI 0.159.0, medium effort10.4s30,81749%$0.00382
Claude Code 2.1.284, high effort11.8s71,76885%$0.0616
Codex CLI 0.159.0, medium effort14.5s§36,03057%$0.00182
Codex CLI 0.159.0, medium effort15.3s44,83762%$0.204
Claude Code 2.1.284, medium effort15.7s34,14572%$0.103
Claude Code 2.1.284, high effort17.3s36,35979%$0.212
Codex CLI 0.159.0, medium effort23.0s§36,76453%$0.0411
Codex CLI 0.159.0, medium effort23.6s§51,58771%$0.0364
Codex CLI 0.159.0, medium effort24.0s§41,32442%$0.110

Click a column to sort, or a model to highlight it on every chart. The Skills page measures how skills change these numbers. Every run is in results.json.

§ Took 1.7x to 2.1x as long in round 13 as in earlier timed runs.

§ GPT-6 Luna, GPT-5.6 Sol, GPT-6 Sol and GPT-6.1 Sol took 1.7x to 2.1x as long in round 13 as in their earlier timed runs, with answers of similar length. They wrote 14 to 20 tokens per second of waiting, down from 27 to 48, which points to OpenAI's service that day rather than more work. With their earlier times, the slowest model took 2.1x as long as the fastest, not 2.9x.

Source: skill-evals.brenhq.com, updated 6 October 2026. Fastest first. Time: average of 2 runs, on the release under each name. Tokens and cost: median of 2 no-skill runs from round 13, on Claude Code 2.1.284 and Codex CLI 0.159.0 with the account extras off.

Tokens, cache and output for each model
Tokens and cost on the explain taskEach bar is the median of 2 runs. Claude models in Claude Code come first, then GPT models in Codex CLI.

Total tokens#

Everything the model read and wrote in the whole run

Sonnet 5: 71,768 tokens, median of 2 runsSonnet 571.8kSonnet 5.5: 33,002 tokens, median of 2 runsSonnet 5.533.0kOpus 5.5: 34,145 tokens, median of 2 runsOpus 5.534.1kFable 5.1: 36,359 tokens, median of 2 runsFable 5.136.4kGPT-5.6 Luna: 30,817 tokens, median of 2 runsGPT-5.6 Luna30.8kGPT-6 Luna: 36,030 tokens, median of 2 runsGPT-6 Luna36.0kGPT-5.6 Sol: 41,324 tokens, median of 2 runsGPT-5.6 Sol41.3kGPT-6 Sol: 36,764 tokens, median of 2 runsGPT-6 Sol36.8kGPT-6.1 Sol: 51,587 tokens, median of 2 runsGPT-6.1 Sol51.6kGPT-6 Astra: 44,837 tokens, median of 2 runsGPT-6 Astra44.8k

Cost at API prices#

What the run's tokens cost at the model's published API prices, cached tokens at the cache price

Sonnet 5: $0.0616 at API prices, median of 2 runsSonnet 5$0.0616Sonnet 5.5: $0.0389 at API prices, median of 2 runsSonnet 5.5$0.0389Opus 5.5: $0.103 at API prices, median of 2 runsOpus 5.5$0.103Fable 5.1: $0.212 at API prices, median of 2 runsFable 5.1$0.212GPT-5.6 Luna: $0.00382 at API prices, median of 2 runsGPT-5.6 Luna$0.00382GPT-6 Luna: $0.00182 at API prices, median of 2 runsGPT-6 Luna$0.00182GPT-5.6 Sol: $0.110 at API prices, median of 2 runsGPT-5.6 Sol$0.110GPT-6 Sol: $0.0411 at API prices, median of 2 runsGPT-6 Sol$0.0411GPT-6.1 Sol: $0.0364 at API prices, median of 2 runsGPT-6.1 Sol$0.0364GPT-6 Astra: $0.204 at API prices, median of 2 runsGPT-6 Astra$0.204

Tokens read#

All the text sent to the model, counted each time it was sent, with Codex CLI's hidden first request

Sonnet 5: 70,921 tokens, median of 2 runsSonnet 570.9kSonnet 5.5: 32,116 tokens, median of 2 runsSonnet 5.532.1kOpus 5.5: 32,935 tokens, median of 2 runsOpus 5.532.9kFable 5.1: 35,210 tokens, median of 2 runsFable 5.135.2kGPT-5.6 Luna: 30,488 tokens, median of 2 runsGPT-5.6 Luna30.5kGPT-6 Luna: 35,872 tokens, median of 2 runsGPT-6 Luna35.9kGPT-5.6 Sol: 40,860 tokens, median of 2 runsGPT-5.6 Sol40.9kGPT-6 Sol: 36,441 tokens, median of 2 runsGPT-6 Sol36.4kGPT-6.1 Sol: 51,297 tokens, median of 2 runsGPT-6.1 Sol51.3kGPT-6 Astra: 44,550 tokens, median of 2 runsGPT-6 Astra44.5k

Read from cache#

Share of what it read that came from the cache, which is cheaper. Counts Codex CLI's hidden first request

Sonnet 5: 85% of prompt tokens, median of 2 runsSonnet 585%Sonnet 5.5: 81% of prompt tokens, median of 2 runsSonnet 5.581%Opus 5.5: 72% of prompt tokens, median of 2 runsOpus 5.572%Fable 5.1: 79% of prompt tokens, median of 2 runsFable 5.179%GPT-5.6 Luna: 49% of prompt tokens, median of 2 runsGPT-5.6 Luna49%GPT-6 Luna: 57% of prompt tokens, median of 2 runsGPT-6 Luna57%GPT-5.6 Sol: 42% of prompt tokens, median of 2 runsGPT-5.6 Sol42%GPT-6 Sol: 53% of prompt tokens, median of 2 runsGPT-6 Sol53%GPT-6.1 Sol: 71% of prompt tokens, median of 2 runsGPT-6.1 Sol71%GPT-6 Astra: 62% of prompt tokens, median of 2 runsGPT-6 Astra62%

Cache writes#

Text saved to the cache for reuse. Codex CLI always reports 0, so for Codex CLI I count what it read that did not come from the cache

Sonnet 5: 10,239 tokens, median of 2 runsSonnet 510.2kSonnet 5.5: 6,223 tokens, median of 2 runsSonnet 5.56.2kOpus 5.5: 9,266 tokens, median of 2 runsOpus 5.59.3kFable 5.1: 7,377 tokens, median of 2 runsFable 5.17.4kGPT-5.6 Luna: 15,640 tokens, median of 2 runsGPT-5.6 Luna15.6kGPT-6 Luna: 15,392 tokens, median of 2 runsGPT-6 Luna15.4kGPT-5.6 Sol: 23,324 tokens, median of 2 runsGPT-5.6 Sol23.3kGPT-6 Sol: 16,985 tokens, median of 2 runsGPT-6 Sol17.0kGPT-6.1 Sol: 14,945 tokens, median of 2 runsGPT-6.1 Sol14.9kGPT-6 Astra: 16,070 tokens, median of 2 runsGPT-6 Astra16.1k

Output tokens#

Everything it wrote: the answer, its tool calls and its hidden thinking

Sonnet 5: 847 tokens, median of 2 runsSonnet 5847Sonnet 5.5: 887 tokens, median of 2 runsSonnet 5.5887Opus 5.5: 1,210 tokens, median of 2 runsOpus 5.51.2kFable 5.1: 1,149 tokens, median of 2 runsFable 5.11.1kGPT-5.6 Luna: 330 tokens, median of 2 runsGPT-5.6 Luna330GPT-6 Luna: 158 tokens, median of 2 runsGPT-6 Luna158GPT-5.6 Sol: 464 tokens, median of 2 runsGPT-5.6 Sol464GPT-6 Sol: 324 tokens, median of 2 runsGPT-6 Sol324GPT-6.1 Sol: 290 tokens, median of 2 runsGPT-6.1 Sol290GPT-6 Astra: 287 tokens, median of 2 runsGPT-6 Astra287

Hidden thinking#

The part of the output spent thinking, which the answer doesn't show

Sonnet 5: 17 tokens, median of 2 runsSonnet 517Sonnet 5.5: 0 tokens, median of 2 runsSonnet 5.50Opus 5.5: 215 tokens, median of 2 runsOpus 5.5215Fable 5.1: 241 tokens, median of 2 runsFable 5.1241GPT-5.6 Luna: 98 tokens, median of 2 runsGPT-5.6 Luna98GPT-6 Luna: 0 tokens, median of 2 runsGPT-6 Luna0GPT-5.6 Sol: 138 tokens, median of 2 runsGPT-5.6 Sol138GPT-6 Sol: 100 tokens, median of 2 runsGPT-6 Sol100GPT-6.1 Sol: 0 tokens, median of 2 runsGPT-6.1 Sol0GPT-6 Astra: 0 tokens, median of 2 runsGPT-6 Astra0

The table scrolls sideways on a small screen.

Explain task, no skill, medians
ModelPassedTokens readFrom cacheCache writesOutputThinkingVisibleCost at API prices
Claude Sonnet 5Claude Code 2.1.284, high effort, Claude Code's default, 30 September 20262 of 270,92185%10,23984717830$0.0616
Claude Sonnet 5.5Claude Code 2.1.284, medium effort, Claude Code's default, 30 September 20262 of 232,11681%6,2238870887$0.0389
Claude Opus 5.5Claude Code 2.1.284, medium effort, Claude Code's default, 30 September 20262 of 232,93572%9,2661,210215995$0.103
Claude Fable 5.1Claude Code 2.1.284, high effort, Claude Code's default, 30 September 20262 of 235,21079%7,3771,149241908$0.212
GPT-5.6 LunaCodex CLI 0.159.0, medium effort, set by me, 30 September 20262 of 230,48849%15,64033098232$0.00382
GPT-6 LunaCodex CLI 0.159.0, medium effort, set by me, 30 September 20262 of 235,87257%15,3921580158$0.00182
GPT-5.6 SolCodex CLI 0.159.0, medium effort, set by me, 30 September 20262 of 240,86042%23,324464138327$0.110
GPT-6 SolCodex CLI 0.159.0, medium effort, set by me, 30 September 20262 of 236,44153%16,985324100224$0.0411
GPT-6.1 SolCodex CLI 0.159.0, medium effort, set by me, 30 September 20262 of 251,29771%14,9452900290$0.0364
GPT-6 AstraCodex CLI 0.159.0, medium effort, set by me, 30 September 20262 of 244,55062%16,0702870287$0.204
How the explain task works

Every run starts a fresh session on a clean copy of the same small Python project, an expense ledger, inside a container with the model's CLI at a pinned release and no instruction file. The prompt asks the model to explain in prose how the code computes monthly totals and what happens with an invalid date, without changing any file. It passes when the answer names ValueError and month.

The Skills page and real app work cover coding. I keep the exact prompts and test code private so no one can tune for them.

Time per run#

Total time runs from starting Claude Code or Codex CLI until it exits. Getting ready and running tools happen inside the CLI, not at the model. GPT-6.1 Sol has its own speed page, timed day by day since launch with its TPS.

Every model spent at least 62% of the run waiting on the model.

Seconds from starting the CLI until it exits, the average of 2 runs per model, fastest first.

  • Getting ready. Starting the CLI, before the first request reaches the model.
  • Waiting on the model. From each request going out until its reply finished, added up.
  • Running tools. Reading the file, passing results back and exiting.

I timed Claude Code's requests with a logging proxy and Codex CLI's with its own trace log. I split the time the same way for every model and ran one model at a time. § Took 1.7x to 2.1x as long in round 13 as in earlier timed runs. The note under the explain task table says why.

§ Took 1.7x to 2.1x as long in round 13 as in earlier timed runs.

Source: skill-evals.brenhq.com, updated 6 October 2026. 2 timed runs per model.

Show first token, when the answer starts, and every number
Time to first token#From the first request going out to the first piece of the reply. Hidden thinking counts as waiting. Fastest first.
When the answer starts showing#From sending the request that produced the written answer to its first words. Fastest first.

Claude Sonnet 5.5's written answer arrived all at once, in 0.1s, after the model finished writing it. The other models streamed theirs word by word.

The table scrolls sideways on a small screen.

Explain task, average of 2 runs per model
ModelPassedTotalGetting readyWaiting on the modelRunning toolsFirst tokenAnswer startsAnswer arrives overOutput tokensTokens per second waited
Claude Sonnet 5.5Claude Code 2.1.284, medium effort2 of 28.1s0.5s7.3s0.4s0.9s5.5s0.1s779107
GPT-5.6 LunaCodex CLI 0.159.0, medium effort2 of 210.4s2.0s7.0s1.4s1.0s1.5s1.6s28641
Claude Sonnet 5Claude Code 2.1.284, high effort2 of 211.8s0.4s11.0s0.4s1.5s1.0s5.8s78274
GPT-6 LunaCodex CLI 0.159.0, medium effort2 of 214.5s2.7s9.0s2.7s0.9s2.5s3.0s18220
GPT-6 AstraCodex CLI 0.159.0, medium effort2 of 215.3s1.9s11.9s1.5s1.9s2.0s6.1s24621
Claude Opus 5.5Claude Code 2.1.284, medium effort2 of 215.7s0.5s14.9s0.3s0.8s10.5s2.5s1,38193
Claude Fable 5.1Claude Code 2.1.284, high effort2 of 217.3s0.5s16.5s0.3s1.7s10.5s3.0s1,05965
GPT-6 SolCodex CLI 0.159.0, medium effort2 of 223.0s4.0s17.7s1.3s2.1s5.2s5.6s31218
GPT-6.1 SolCodex CLI 0.159.0, medium effort2 of 223.6s2.1s19.8s1.7s2.5s3.4s8.3s27014
GPT-5.6 SolCodex CLI 0.159.0, medium effort2 of 224.0s2.6s20.0s1.4s1.1s5.6s8.1s37419

Method and data#

Earlier runs: launch days

These ran before round 13, with the account extras on. Some numbers here carry stars, and they differ from the explain task table above.

Sonnet 5.5 against Sonnet 5 and Opus 5.5#

All three models ran the explain task on Sonnet 5.5's launch day, in one Claude Code release. Claude Code sends each model its own instructions. Its first request held 40,414 tokens for Sonnet 5 and 28,127 for Sonnet 5.5, measured in that day's caveman runs that were not light. With no skill Sonnet 5 also took three steps where Sonnet 5.5 took two. The bigger request and the extra step explain most of why Sonnet 5 used more tokens.

Show the charts and every number
Sonnet 5.5, Sonnet 5 and Opus 5.5 on the explain taskEach bar is the median of 2 runs, on Claude Code 2.1.284 at its default effort, medium for Sonnet 5.5 and Opus 5.5, high for Sonnet 5.

Total tokens#

Everything the model read and wrote in the whole run

Sonnet 5.5: 42,884 tokens, median of 2 runs, mostly light runsSonnet 5.542.9k*Sonnet 5: 115,842 tokens, median of 2 runsSonnet 5116kOpus 5.5: 36,465 tokens, median of 2 runs, mostly light runsOpus 5.536.5k*

Cost at API prices#

What the run's tokens cost at the model's published API prices, cached tokens at the cache price

Sonnet 5.5: $0.0602 at API prices, median of 2 runs, mostly light runsSonnet 5.5$0.0602*Sonnet 5: $0.110 at API prices, median of 2 runsSonnet 5$0.110Opus 5.5: $0.109 at API prices, median of 2 runs, mostly light runsOpus 5.5$0.109*

Tokens read#

All the text sent to the model, counted each time it was sent

Sonnet 5.5: 41,973 tokens, median of 2 runs, mostly light runsSonnet 5.542.0k*Sonnet 5: 115,001 tokens, median of 2 runsSonnet 5115kOpus 5.5: 34,984 tokens, median of 2 runs, mostly light runsOpus 5.535.0k*

Read from cache#

Share of what it read that came from the cache, which is cheaper

Sonnet 5.5: 75% of prompt tokens, median of 2 runs, mostly light runsSonnet 5.575%*Sonnet 5: 82% of prompt tokens, median of 2 runsSonnet 582%Opus 5.5: 74% of prompt tokens, median of 2 runs, mostly light runsOpus 5.574%*

Cache writes#

Text saved to the cache for reuse

Sonnet 5.5: 11,233 tokens, median of 2 runs, mostly light runsSonnet 5.511.2k*Sonnet 5: 20,610 tokens, median of 2 runsSonnet 520.6kOpus 5.5: 9,248 tokens, median of 2 runs, mostly light runsOpus 5.59.2k*

Output tokens#

Everything it wrote: the answer, its tool calls and its hidden thinking

Sonnet 5.5: 911 tokens, median of 2 runs, mostly light runsSonnet 5.5911*Sonnet 5: 842 tokens, median of 2 runsSonnet 5842Opus 5.5: 1,482 tokens, median of 2 runs, mostly light runsOpus 5.51.5k*

Hidden thinking#

The part of the output spent thinking, which the answer doesn't show

Sonnet 5.5: 37 tokens, median of 2 runs, mostly light runsSonnet 5.537*Sonnet 5: 27 tokens, median of 2 runsSonnet 527Opus 5.5: 424 tokens, median of 2 runs, mostly light runsOpus 5.5424*

* At least half the runs behind a starred number are light. Light runs start with fewer tokens of instructions and tools than the model's other runs, and some ended with a remark about the account's connectors, so their counts differ. What light runs are.

The table scrolls sideways on a small screen.

Same Claude Code release, explain task, no skill, medians
ModelPassedTokens readFrom cacheCache writesOutputThinkingVisibleCost at API prices
Claude Sonnet 5.5Claude Code 2.1.284, medium effort, Claude Code's default2 of 241,973*75%*11,233*911*37*875*$0.0602*
Claude Sonnet 5Claude Code 2.1.284, high effort, Claude Code's default2 of 2115,00182%20,61084227815$0.110
Claude Opus 5.5Claude Code 2.1.284, medium effort, Claude Code's default2 of 234,984*74%*9,248*1,482*424*1,058*$0.109*

GPT-6.1 Sol against GPT-6 Sol and Opus 5.5#

GPT-6.1 Sol ran on 29 September 2026, the day it came out, beside GPT-6 Sol and Opus 5.5. On this task it reported no hidden thinking at all.

Show how I ran it, the charts and every number

GPT-6.1 Sol and GPT-6 Sol ran the same task in Codex CLI 0.159.0, the first release that runs GPT-6.1 Sol on a ChatGPT login. I set both to medium effort. Opus 5.5 ran it the same day in Claude Code at its default effort, medium. Because each CLI adds its own instructions, compare what the models wrote, not what they read. Codex CLI's trace log shows it sent the effort setting on every request. For GPT-6.1 Sol, OpenAI's server reported 0 tokens of hidden thinking at medium effort, and at high effort in a check outside the counted runs. GPT-6 Sol reported hidden thinking for the same request.

GPT-6.1 Sol, GPT-6 Sol and Opus 5.5 on the explain taskEach bar is the median of 2 runs, the same day: the GPT models in Codex CLI 0.159.0 at medium effort, Opus 5.5 in Claude Code at its default medium effort.

Total tokens#

Everything the model read and wrote in the whole run

GPT-6.1 Sol: 49,053 tokens, median of 2 runsGPT-6.1 Sol49.1kGPT-6 Sol: 40,175 tokens, median of 2 runsGPT-6 Sol40.2kOpus 5.5: 46,516 tokens, median of 2 runs, mostly light runsOpus 5.546.5k*

Cost at API prices#

What the run's tokens cost at the model's published API prices, cached tokens at the cache price

GPT-6.1 Sol: $0.0366 at API prices, median of 2 runsGPT-6.1 Sol$0.0366GPT-6 Sol: $0.0427 at API prices, median of 2 runsGPT-6 Sol$0.0427Opus 5.5: $0.150 at API prices, median of 2 runs, mostly light runsOpus 5.5$0.150*

Tokens read#

All the text sent to the model, counted each time it was sent, with Codex CLI's hidden first request

GPT-6.1 Sol: 48,786 tokens, median of 2 runsGPT-6.1 Sol48.8kGPT-6 Sol: 39,896 tokens, median of 2 runsGPT-6 Sol39.9kOpus 5.5: 44,767 tokens, median of 2 runs, mostly light runsOpus 5.544.8k*

Read from cache#

Share of what it read that came from the cache, which is cheaper. Counts Codex CLI's hidden first request

GPT-6.1 Sol: 68% of prompt tokens, median of 2 runsGPT-6.1 Sol68%GPT-6 Sol: 56% of prompt tokens, median of 2 runsGPT-6 Sol56%Opus 5.5: 70% of prompt tokens, median of 2 runs, mostly light runsOpus 5.570%*

Cache writes#

Text saved to the cache for reuse. Codex CLI always reports 0, so for Codex CLI I count what it read that did not come from the cache

GPT-6.1 Sol: 15,314 tokens, median of 2 runsGPT-6.1 Sol15.3kGPT-6 Sol: 17,752 tokens, median of 2 runsGPT-6 Sol17.8kOpus 5.5: 13,649 tokens, median of 2 runs, mostly light runsOpus 5.513.6k*

Output tokens#

Everything it wrote: the answer, its tool calls and its hidden thinking

GPT-6.1 Sol: 267 tokens, median of 2 runsGPT-6.1 Sol267GPT-6 Sol: 279 tokens, median of 2 runsGPT-6 Sol279Opus 5.5: 1,750 tokens, median of 2 runs, mostly light runsOpus 5.51.7k*

Hidden thinking#

The part of the output spent thinking, which the answer doesn't show

GPT-6.1 Sol: 0 tokens, median of 2 runsGPT-6.1 Sol0GPT-6 Sol: 80 tokens, median of 2 runsGPT-6 Sol80Opus 5.5: 502 tokens, median of 2 runs, mostly light runsOpus 5.5502*

* At least half the runs behind a starred number are light. Light runs start with fewer tokens of instructions and tools than the model's other runs, and some ended with a remark about the account's connectors, so their counts differ. What light runs are.

The table scrolls sideways on a small screen.

Same day, explain task, no skill, medians
ModelPassedTokens readFrom cacheCache writesOutputThinkingVisibleCost at API prices
GPT-6.1 SolCodex CLI 0.159.0, medium effort, set by me2 of 248,78668%15,3142670267$0.0366
GPT-6 SolCodex CLI 0.159.0, medium effort, set by me2 of 239,89656%17,75227980199$0.0427
Claude Opus 5.5Claude Code 2.1.284, medium effort, Claude Code's default2 of 244,767*70%*13,649*1,750*502*1,248*$0.150*
Effort

Effort sets how much a model thinks before and between its steps.

ModelCLIEffortSet by
Claude Sonnet 5Claude Code 2.1.278, 2.1.283 and 2.1.284highClaude Code's defaultdefault
Claude Opus 5.5Claude Code 2.1.280, 2.1.283 and 2.1.284mediumClaude Code's defaultdefault
Claude Sonnet 5.5Claude Code 2.1.284mediumClaude Code's defaultdefault
Claude Fable 5.1Claude Code 2.1.283 and 2.1.284highClaude Code's defaultdefault
GPT-5.6 LunaCodex CLI 0.153.3lowset by meset by me
GPT-6 LunaCodex CLI 0.156.0lowset by meset by me
GPT-5.6 Luna, GPT-6 Luna, GPT-5.6 Sol, GPT-6 Sol and GPT-6 AstraCodex CLI 0.157.1mediumset by meset by me
GPT-6.1 Sol, GPT-6 Sol, GPT-5.6 Sol, GPT-5.6 Luna, GPT-6 Luna and GPT-6 AstraCodex CLI 0.159.0mediumset by meset by me

GPT-5.6 Luna, GPT-6 Luna, GPT-5.6 Sol, GPT-6 Sol and GPT-6 Astra appear more than once, because each row is one release and effort a model ran at.

Claude Code picks the effort for each Claude model, and I left it alone. I logged the effort in every request it sent on 28 September 2026, on Claude Code 2.1.284, the release every later Claude run used. I assume earlier releases sent the same default. Within that effort, each Claude model sets its own amount of thinking. Anthropic's own default for Sonnet 5.5 is high, but Claude Code sends medium. For the GPT models I set the effort on every run, and Codex CLI's replies confirmed it.

Definitions
  • Tokens are the chunks of text a model reads and writes, about three quarters of a word each. Tokens read counts all the text sent to the model, each time it was sent. An agent sends its instructions and the conversation again with every request. That total grows with each step.
  • The cache keeps text the model has already read, so the next request can reuse it faster and at a lower price. Read from cache is the share of what it read that came from there. Codex CLI's own count leaves out its hidden first request. Without it, Codex CLI's count gives 56% to 91% for these runs. Cache writes are text saved to the cache. Codex CLI always reports 0, so for Codex CLI I count what it read that did not come from the cache.
  • Codex CLI's hidden first request. At the start of every run Codex CLI sends a request that loads the model's instructions and tools and writes nothing. Codex CLI leaves it out of its own token count, so I measured it once for each model and Codex CLI release and add it back.
  • Output is everything the model wrote: its answer, its tool calls, such as asking to read a file, and its hidden thinking. Visible output is output minus thinking.
  • Time on a run this short is mostly fixed: a second or two to start, then one to three seconds before each reply begins. A short answer is not much faster than a long one. Two runs of the same model can differ by several seconds, and a gap of a second is a tie.
  • Account extras are what a subscription login loads by default. For Claude Code, they are the plugins, skills and connectors on the account's claude.ai. For Codex CLI, they are its connected apps and plugins. Round 13 turned them off. The launch-day runs under Earlier runs ran before that, with them on.
  • Cost at API prices is what a run's tokens would cost at the model's published API prices: input that missed the cache at the input price, cache reads at the cache price, Claude's cache writes at the one-hour rate Claude Code uses, and output. For every Claude run it equals Claude Code's own estimate. My runs used subscriptions, which don't bill per token. Runs share a cache, so a run's cost depends partly on the runs before it. The price table lists every rate.
  • Light runs are Claude Code runs that started with thousands fewer tokens of instructions and tools than the same model's other runs. They were a session's first run, before Claude Code loaded the account extras. They cost less and read fewer tokens. Some ended with a remark about the account's connectors. I removed that remark from the answers, but their token counts still include it. Round 13 reran this page's explain runs and real app work with the account extras off. A number where at least half the runs behind it are light carries a star. The Skills page's limits say how I find them.
  • Effort is picked by Claude Code for Claude models and set by me for GPT models. The effort table lists every model.
  • The token charts and tables show medians: the middle run, or the average of the two middle runs. The speed charts show the average of each model's timed runs. Every run is in the download below.

API prices used for cost#

The table scrolls sideways on a small screen.

API prices used for cost, US dollars per million tokens
ModelInputCache readCache writeOutput
Claude Sonnet 5$2.00$0.20$4.00$10.00
Claude Sonnet 5.5$2.00$0.20$4.00$10.00
Claude Opus 5.5$4.00$0.20$8.00$20.00
Claude Fable 5.1$10.00$0.25$20.00$50.00
GPT-5.6 Luna$0.20$0.02none$1.20
GPT-6 Luna$0.10$0.01none$0.50
GPT-5.6 Sol$4.00$0.40none$20.00
GPT-6 Sol$2.00$0.20none$10.00
GPT-6.1 Sol$2.00$0.10none$10.00
GPT-6 Astra$10.00$1.00none$50.00

Checked 29 September 2026. US dollars per million tokens, standard tier. OpenAI's rates are its short-context rates, which cover every request in these runs. OpenAI lists cache writes only for long context. Claude cache writes at the one-hour rate Claude Code uses, 2x input. Sources: OpenAI's pricing, Anthropic's pricing and Anthropic's caching docs.

Changes

Changes since the Models page began on 28 September. The Skills page has the full list.

  • Added the Models page at /numbers. It shows every model's input tokens, cache reads and writes, output, hidden thinking, time and output per second side by side, for round 8 and for each task and skill in rounds 1 to 7.
  • Codex CLI token counts now include the hidden first request Codex CLI leaves out of its own count, 9,751 to 12,371 tokens a run, measured for each model and CLI release. Codex CLI cache writes, which Codex CLI always reports as 0, are now the prompt tokens it did not read from cache.
  • Round 9 ran Claude Sonnet 5.5 on its launch day, with Sonnet 5 rerun on the same Claude Code release the same day, on the explain task.
  • Opus 5.5 ran the explain task twice on Claude Code 2.1.284, beside Sonnet 5.5 and Sonnet 5 on the same release. The Models page adds a total tokens chart and link anchors on every section and chart.
  • Added an effort table: the effort Claude Code sent for each Claude model, logged from its requests, and the effort I set for each Codex CLI model.
  • Timed every model on the explain task, two runs each, with the same method for Claude Code and Codex CLI: total time split into getting ready, waiting on the model and running tools, plus time to first token and when the written answer starts. The Models page now opens with a sortable leaderboard.
  • The site now has two pages. Skills tests each skill against the same task with no skill, and Models shows every model with no skill. The per-skill token charts moved to Skills and the effort table to Models, and old links forward to the new place.
  • The Sonnet 5.5 chart on the Models page now names the effort each model ran at: medium for Sonnet 5.5 and Opus 5.5, high for Sonnet 5, each Claude Code's default.
  • The Models page now opens with the leaderboard under a plain title and one sentence on the task. The table keeps time to finish, tokens used and share from cache, and fits a phone; first token and answer start are in the speed charts below it.

  • GPT-6.1 Sol came out during OpenAI's DevDay, and I ran it that day at medium effort on Codex CLI 0.159.0, the first release that runs it on a ChatGPT login: 2 explain runs with no skill, 2 with caveman and 2 timed runs. GPT-6 Sol reran on the same release and Opus 5.5 on Claude Code 2.1.284 the same day. All three are on the leaderboard and the overall chart with these runs.
  • Every run on the Models page now shows what its tokens cost at the model's published API prices, cached tokens at the cache rate, with the price table and its sources. For Claude it matches Claude Code's own estimate on all 194 Claude runs.
  • The Models page now says what the test is, under the leaderboard: the explain task, how it passes, and that it writes no code. The coding tasks are on the Skills page.
  • Added real app work to the Models page: a bug hunt with 20 planted bugs and a tickets test with 17 hidden tests, in a small TypeScript web app, 3 runs each for Opus 5.5, Sonnet 5.5, GPT-5.6 Sol, GPT-6 Sol and GPT-6.1 Sol, graded by hidden tests, with cost at API prices and time.
  • Correction: the effort table left out two settings the runs used, Claude Fable 5.1 on Claude Code 2.1.284 in the timed runs and GPT-5.6 Sol on Codex CLI 0.159.0 in the real app work. Both are listed now, and QA fails if a run's release is ever missing again.
  • Redesigned both pages. The Skills page opens with the count of claimed cuts that held, one chart of each claim against the measured cut, and each skill's claim, finding and one-word verdict. The Models page leads with real app work. Text runs the full page width, long detail sits in folds, charts save as images with CSV, and bar charts show bars only.

  • Correction: some Claude Code runs started with thousands fewer tokens of instructions and tools. They ran before Claude Code had synced the account's claude.ai plugins and skills and loaded its connectors, so they cost less and count fewer total tokens. A number where they make up at least half the runs now carries a star, a new limit explains how I found them, and I removed a remark about the connectors from the published answers. No claim check changes. The count of Claude runs in the 29 September cost entry is corrected from 254 to 194; it had counted Sonnet 5.5's runs twice.
  • One GPT-5.6 Sol tickets run edited an existing test file, which the prompt said not to do. The grader restores the original tests first, so its 17 of 17 stood. Round 13 replaced that run, and no rerun edited a test.
  • Round 13 reran the explain task for every model on the Models page, every Claude model's explain runs, all 30 real app work runs and the timed runs, all with the account extras off. Those numbers no longer carry stars, and Claude's explain-task token counts and costs dropped. caveman now cuts Sonnet 5.5's explanations 28%, where round 9 showed 13%, and Sonnet 5's 72%, still above its 65% claim. All five models still passed all 17 hidden ticket tests, and Opus 5.5 cost more, $0.96 a ticket run.
  • Redesigned the findings, tables and colors. Each finding has its own framed chart with horizontal bars and a saved image, the skill cards and real-work tables line up and sort like the Models page table, each vendor keeps one color family, and dark mode is warm.
  • Correction: the Skills page gave Sonnet 5.5's release day as 30 September, when it came out on 28 September. Two table notes still said cost covered Claude only, and the effort table left out the Codex CLI 0.159.0 row for three GPT models. All three are fixed, and QA now reads every run group.
  • The chart for any measure and task now uses the same framed horizontal bars as the findings, with a saved image. The Models page notes how much round 13 moved its numbers and has its own list of changes, and on a phone its table stays a table.
  • The chart for any measure and task is now its own section, Every measure and task. Charts with three skills stack at full width with a labeled scale, bars from failed runs show hollow, and the Models page notes every model the round 13 timing slowed.
  • Correction: 12 Claude coding runs also started light. Until now I had checked only the explain task. They now carry a star in the run list, and no median has half its runs light. The Models table title gives the time spread with and without the slow OpenAI day, and the cost panels say their range adds up each task's cheapest to costliest run.
  • On a phone, bar charts put each model's name above its bar. Saved images keep their panels side by side, the Models token charts are horizontal bars, and the Skills page adds a table of every model's cut under the top chart.
  • The GPT models that ran slower in round 13's timing carry a dagger in the explain table and the time chart, and the real app work note says their runs slowed the same day. Costs use one format on every table and chart.
  • Bar charts with the name above each bar leave a clear gap between rows, and their lines no longer cross a label. Grid cells that could be chance are outlined like the bars. The cut table names every model and marks the one cut that met its claim.
  • Costs use one format everywhere, three significant figures under ten cents and three decimals above. The cost-vs-time title and dots mark the slowed GPT models, and the run list adds round 10.
  • The cut table fits tablets and phones, where each claim becomes a block of model and cut pairs. Saved images explain their daggers, small cost bars get their own scale, and tables and charts in folds share one title style.
  • Round 8's table now stars its light medians. The real-work title and note use the table's cost figures, and the cost-vs-time title compares cost only. The page says a claim held, in one word, everywhere.
  • Output and thinking counts carry the light-run star too. The slowed GPT models are marked § instead of a dagger. The cut table says whether a cell was not compared or failed, and dates and model names no longer break across lines on phones.
  • The tick strip and the byline are gone: the chart under the intro shows the same checks, and its source line carries the date.
  • The real app work chart and the Models board views drop the rings and the scores on each dot, and the rerun note is two sentences.
  • The real app work chart says what its tests can't do: every model scored at or near full marks, so they show cost and time on a small job, not which model is best.
  • A dot on the real app work chart names its score when it fell short of full marks, as Claude Sonnet 5.5's bug hunt dot does at 19 of 20.

  • New page, GPT-6.1 Sol timed three days running: one short Python file explained on launch day, the day after, and on 1 October after OpenAI said it added capacity. The Models board keeps the 30 September runs.
  • The GPT-6.1 Sol page says what the test is up front and groups the chart by day. The ChatGPT apps and plugins setting moved to the method, in plain words.
  • The GPT-6.1 Sol page adds four runs from the afternoon of 1 October, after @sama said it should be much better: 13.5 to 16.8 s, faster than that morning and still slower than launch day.

  • The GPT-6.1 Sol page adds four runs from the afternoon of 3 October, after a 2 October post on X still called it slow: 15.6 to 26.6 s, slower by median than the afternoon of 1 October. Corrected 5 October: the two afternoons' runs overlap.

  • The GPT-6.1 Sol page adds four runs from the afternoon of 4 October: 13.2 to 22.9 s, all slower than launch day.

  • The GPT-6.1 Sol page adds four runs from 5 October, before OpenAI's speed post: 14.1 to 18.6 s, all slower than launch day.
  • The GPT-6.1 Sol page adds four runs from 5 October, minutes after OpenAI said its speed change would be felt within two hours: 14.3 to 19.5 s.
  • Readers were right that the GPT-6.1 Sol page's tokens a second was not TPS. It divides output by the time spent waiting on the model's replies, so the column is now tokens a second of waiting.
  • Headlines say a setup cut more, won or held its claim only when its runs sit apart from the other side's. Where runs overlap they now say by median, so three headlines changed, and the all-tasks totals add up each side's full run range.
  • The GPT-6.1 Sol page adds four runs from a 5 October retest, still inside the two hours OpenAI gave: 11.3 to 20.8 s, three of them inside launch day's range.
  • A TPS column briefly showed a rate from an OpenAI server-side timing field. That field times the engine, not delivery: on launch day it read 3.7 to 3.9 times the rate text reached my machine. It came down. A TPS proxy column now counts the pieces of answer text Codex logged a second, pauses included.
  • Two test runs of the updated timing script on 5 October are listed under How I ran it on the GPT-6.1 Sol page, not charted.
  • The GPT-6.1 Sol and ASD-STE100 pages said each run got a fresh container. Each run gets a fresh copy of the project; runs share the bench's container. The Sol page now says a run passes when its answer contains both "month" and "ValueError".
  • The GPT-6.1 Sol page's intro said posts on X on 3 October called it slow again. No such post could be found, so it now cites the 2 and 4 October posts that did.
  • The GPT-6.1 Sol page adds four runs from 5 October, after the two hours OpenAI gave: 11.4 to 17.7 s, streaming at 39.2 to 52.9 text pieces a second.
  • Readers asked how launch day was as quick with a slower stream. The GPT-6.1 Sol page now says why: the answer's stream is under half of every run, launch day's answers were shorter, and the runs since that made two model requests, like launch day, finished inside launch day's range.
  • Readers asked to sort the GPT-6.1 Sol run table. Click any column heading: Time puts the fastest first, TPS proxy the highest first, a second click reverses, and Day restores day order.
  • The GPT-6.1 Sol chart took more than a page. It is now one row per sitting in two panels, TPS proxy and seconds, each a median with launch day and the sittings since the 5 October retest picked out; the 5 October sittings are named by OpenAI's post, not the time of day; every run stays in the table. The intro is one line, and the background folds under How I ran it.
  • Readers asked whether the GPT-6.1 Sol page has true TPS. The table now has TPS by OpenAI's own token count for the eight runs since 5 October's retest: the answer's output tokens over the seconds from its first text event to its last, a rate observed at my machine. Earlier rows have no count and stay blank; two script test runs also logged it, at 36.2 and 43.4, and are in the CSV. The text-event rate is now labeled text pieces a second.
  • Readers couldn't find the GPT-6.1 Sol speed page from the main pages. It now has its own tab in the bar on every page (Sol on a phone), and the Models page's time section links it.

  • The GPT-6.1 Sol page adds four runs from 6 October: 10.3 to 15.7 s, the fastest run yet at 10.3 s, and 43.4 to 51.2 tokens a second by OpenAI's count. Every sitting from 5 October's retest on now counts as since the retest.
  • A reader asked for Claude Opus 5.5 and Claude Sonnet 5.5 beside the GPT-6.1 Sol test. The Sol page now shows them timed the same day and task, 4 runs each, start to finish in each model's own tool: Sonnet 5.5 7.8 to 9.8 s, Sol 10.3 to 15.7 s, Opus 5.5 15.2 to 17.7 s. No TPS for the Claude models: this bench's Claude Code capture does not give a streaming rate yet.
  • Readers loved the speed race videos on X, so the GPT-6.1 Sol page now has the same race, drawn live from every run: one lane per sitting, all leaving together at 5x with real seconds on the clock. It plays once when it scrolls into view and has a Replay button; with reduced motion it waits for Play.
  • Every section of the GPT-6.1 Sol page has a heading with a # link that copies its address, and an On this page row jumps between them: every sitting, the race, same day with other models, and how I ran it.
  • Readers said the race's dots were softer than the videos'. They are now drawn at full screen density with a sharp edge, a glow sized to the dot and the videos' white highlight.
  • The site now counts visits with Cloudflare Web Analytics, which sets no cookies. The footer, which said no analytics, now says so.
  • Each page now ships its text in the HTML, so search engines and AI tools that do not run scripts can read it. The site also has a sitemap, an llms.txt and a footer line linking every page.
  • Readers said the link cards on X did not show the site well. Each card now has the site's bar with its page marked, one finding and one chart, sized to read in a feed; the GPT-6.1 Sol card shows the streaming chart alone.
  • The GPT-6.1 Sol page's method now says the group since 5 October's retest was set after seeing the 5 October runs, and every later sitting joins it whatever it shows. The link cards have larger chart text.
Download the run records as JSON