All five models passed all 17 hidden ticket tests in real app work.
A ticket run cost $0.195 to $0.958 and took 1.1 to 9.1 minutes, in a small TypeScript app. Ten models also explained one short Python file.
Real app work#
GPT-6.1 Sol was cheapest on the tickets test at $0.195 a run, and Claude Sonnet 5.5 was fastest at 1.1 minutes.
Each dot is a model's median run. Every model scored at or near full marks, so these tests show cost and time on a small job, not which model is best.
- Claude models in Claude Code
- GPT models in Codex CLI
Round 13 reran all 30 runs with the account extras off. Every median score held, costs moved by up to 61% and four of the five models took at least 1.4 times as long.
Source: skill-evals.brenhq.com, updated 6 October 2026. 3 runs per model on each test, graded by hidden tests, with the account extras off.
| Model | |||||
|---|---|---|---|---|---|
| GPT‑6.1 SolGPT‑6.1 SolCodex CLI 0.159.0, medium effort | 20, 20, 20 | $0.223 | 8.2 min | 21 | 82 |
| GPT‑6 SolGPT‑6 SolCodex CLI 0.159.0, medium effort | 20, 20, 20 | $0.280 | 8.5 min | 12 | 64 |
| GPT‑5.6 SolGPT‑5.6 SolCodex CLI 0.159.0, medium effort | 20, 20, 20 | $0.574 | 8.6 min | 14 | 54 |
| Claude Opus 5.5Opus 5.5Claude Code 2.1.284, medium effort | 20, 20, 20 | $0.782 | 2.6 min | 25 | 84 |
| Claude Sonnet 5.5Sonnet 5.5Claude Code 2.1.284, medium effort | 19, 19, 19 | $0.151 | 0.4 min | 0 | 40 |
| Model | |||||
|---|---|---|---|---|---|
| GPT‑6.1 SolGPT‑6.1 SolCodex CLI 0.159.0, medium effort | 17, 17, 17 | $0.195 | 7.6 min | 11 | 166 |
| Claude Sonnet 5.5Sonnet 5.5Claude Code 2.1.284, medium effort | 17, 17, 17 | $0.266 | 1.1 min | 6 | 185 |
| GPT‑6 SolGPT‑6 SolCodex CLI 0.159.0, medium effort | 17, 17, 17 | $0.272 | 6.9 min | 5 | 168 |
| GPT‑5.6 SolGPT‑5.6 SolCodex CLI 0.159.0, medium effort | 17, 17, 17 | $0.541 | 9.1 min | 6 | 206 |
| Claude Opus 5.5Opus 5.5Claude Code 2.1.284, medium effort | 17, 17, 17 | $0.958 | 3.2 min | 23 | 252 |
Tests written counts the new test cases a run added. Lines changed counts the app code it added or removed. Both are medians of the runs. Neither is a score. More tests or fewer lines isn't better by itself, but both show how much a reviewer has to check.
How I tested
Every model got the same small TypeScript web app, with an API, a React front end and a written spec. Hidden tests the model never saw graded the result. The bug hunt plants 20 bugs across the front end, the API and shared logic. The tickets test asks for weekly repeating bookings, a fix for a double booking reported only by its symptom, and a fix for a security hole, checked by 17 hidden tests. Each model ran each test 3 times with no skill. Claude Opus 5.5 planted half the bugs and GPT-6.1 Sol the other half, and GPT-6.1 Sol reviewed the ticket tests for fairness. The apps and the tests stay private, so no model can train on them. The app is small, so a full pass says little about harder code.
The explain task#
Across 10 models on the explain task, the priciest cost 116x as much per run as the cheapest, and the slowest took 2.9x as long as the fastest§.
Each model ran with no skill in Claude Code or Codex CLI, read the same short Python file and explained it. Cost is at each model's published API prices.
- Claude models in Claude Code
- GPT models in Codex CLI
The table scrolls sideways on a small screen.
| Model | ||||
|---|---|---|---|---|
| Claude Code 2.1.284, medium effort | 8.1s | 33,002 | 81% | $0.0389 |
| Codex CLI 0.159.0, medium effort | 10.4s | 30,817 | 49% | $0.00382 |
| Claude Code 2.1.284, high effort | 11.8s | 71,768 | 85% | $0.0616 |
| Codex CLI 0.159.0, medium effort | 14.5s§ | 36,030 | 57% | $0.00182 |
| Codex CLI 0.159.0, medium effort | 15.3s | 44,837 | 62% | $0.204 |
| Claude Code 2.1.284, medium effort | 15.7s | 34,145 | 72% | $0.103 |
| Claude Code 2.1.284, high effort | 17.3s | 36,359 | 79% | $0.212 |
| Codex CLI 0.159.0, medium effort | 23.0s§ | 36,764 | 53% | $0.0411 |
| Codex CLI 0.159.0, medium effort | 23.6s§ | 51,587 | 71% | $0.0364 |
| Codex CLI 0.159.0, medium effort | 24.0s§ | 41,324 | 42% | $0.110 |
Click a column to sort, or a model to highlight it on every chart. The Skills page measures how skills change these numbers. Every run is in results.json.
§ Took 1.7x to 2.1x as long in round 13 as in earlier timed runs.
§ GPT-6 Luna, GPT-5.6 Sol, GPT-6 Sol and GPT-6.1 Sol took 1.7x to 2.1x as long in round 13 as in their earlier timed runs, with answers of similar length. They wrote 14 to 20 tokens per second of waiting, down from 27 to 48, which points to OpenAI's service that day rather than more work. With their earlier times, the slowest model took 2.1x as long as the fastest, not 2.9x.
Source: skill-evals.brenhq.com, updated 6 October 2026. Fastest first. Time: average of 2 runs, on the release under each name. Tokens and cost: median of 2 no-skill runs from round 13, on Claude Code 2.1.284 and Codex CLI 0.159.0 with the account extras off.
Tokens, cache and output for each model
Total tokens#
Everything the model read and wrote in the whole run
Cost at API prices#
What the run's tokens cost at the model's published API prices, cached tokens at the cache price
Tokens read#
All the text sent to the model, counted each time it was sent, with Codex CLI's hidden first request
Read from cache#
Share of what it read that came from the cache, which is cheaper. Counts Codex CLI's hidden first request
Cache writes#
Text saved to the cache for reuse. Codex CLI always reports 0, so for Codex CLI I count what it read that did not come from the cache
Output tokens#
Everything it wrote: the answer, its tool calls and its hidden thinking
Hidden thinking#
The part of the output spent thinking, which the answer doesn't show
The table scrolls sideways on a small screen.
| Model | Passed | Tokens read | From cache | Cache writes | Output | Thinking | Visible | Cost at API prices |
|---|---|---|---|---|---|---|---|---|
| Claude Sonnet 5Claude Code 2.1.284, high effort, Claude Code's default, 30 September 2026 | 2 of 2 | 70,921 | 85% | 10,239 | 847 | 17 | 830 | $0.0616 |
| Claude Sonnet 5.5Claude Code 2.1.284, medium effort, Claude Code's default, 30 September 2026 | 2 of 2 | 32,116 | 81% | 6,223 | 887 | 0 | 887 | $0.0389 |
| Claude Opus 5.5Claude Code 2.1.284, medium effort, Claude Code's default, 30 September 2026 | 2 of 2 | 32,935 | 72% | 9,266 | 1,210 | 215 | 995 | $0.103 |
| Claude Fable 5.1Claude Code 2.1.284, high effort, Claude Code's default, 30 September 2026 | 2 of 2 | 35,210 | 79% | 7,377 | 1,149 | 241 | 908 | $0.212 |
| GPT-5.6 LunaCodex CLI 0.159.0, medium effort, set by me, 30 September 2026 | 2 of 2 | 30,488 | 49% | 15,640 | 330 | 98 | 232 | $0.00382 |
| GPT-6 LunaCodex CLI 0.159.0, medium effort, set by me, 30 September 2026 | 2 of 2 | 35,872 | 57% | 15,392 | 158 | 0 | 158 | $0.00182 |
| GPT-5.6 SolCodex CLI 0.159.0, medium effort, set by me, 30 September 2026 | 2 of 2 | 40,860 | 42% | 23,324 | 464 | 138 | 327 | $0.110 |
| GPT-6 SolCodex CLI 0.159.0, medium effort, set by me, 30 September 2026 | 2 of 2 | 36,441 | 53% | 16,985 | 324 | 100 | 224 | $0.0411 |
| GPT-6.1 SolCodex CLI 0.159.0, medium effort, set by me, 30 September 2026 | 2 of 2 | 51,297 | 71% | 14,945 | 290 | 0 | 290 | $0.0364 |
| GPT-6 AstraCodex CLI 0.159.0, medium effort, set by me, 30 September 2026 | 2 of 2 | 44,550 | 62% | 16,070 | 287 | 0 | 287 | $0.204 |
How the explain task works
Every run starts a fresh session on a clean copy of the same small Python project, an expense ledger, inside a container with the model's CLI at a pinned release and no instruction file. The prompt asks the model to explain in prose how the code computes monthly totals and what happens with an invalid date, without changing any file. It passes when the answer names ValueError and month.
The Skills page and real app work cover coding. I keep the exact prompts and test code private so no one can tune for them.
Time per run#
Total time runs from starting Claude Code or Codex CLI until it exits. Getting ready and running tools happen inside the CLI, not at the model. GPT-6.1 Sol has its own speed page, timed day by day since launch with its TPS.
Every model spent at least 62% of the run waiting on the model.
Seconds from starting the CLI until it exits, the average of 2 runs per model, fastest first.
- Getting ready. Starting the CLI, before the first request reaches the model.
- Waiting on the model. From each request going out until its reply finished, added up.
- Running tools. Reading the file, passing results back and exiting.
I timed Claude Code's requests with a logging proxy and Codex CLI's with its own trace log. I split the time the same way for every model and ran one model at a time. § Took 1.7x to 2.1x as long in round 13 as in earlier timed runs. The note under the explain task table says why.
§ Took 1.7x to 2.1x as long in round 13 as in earlier timed runs.
Source: skill-evals.brenhq.com, updated 6 October 2026. 2 timed runs per model.
Show first token, when the answer starts, and every number
Claude Sonnet 5.5's written answer arrived all at once, in 0.1s, after the model finished writing it. The other models streamed theirs word by word.
The table scrolls sideways on a small screen.
| Model | Passed | Total | Getting ready | Waiting on the model | Running tools | First token | Answer starts | Answer arrives over | Output tokens | Tokens per second waited |
|---|---|---|---|---|---|---|---|---|---|---|
| Claude Sonnet 5.5Claude Code 2.1.284, medium effort | 2 of 2 | 8.1s | 0.5s | 7.3s | 0.4s | 0.9s | 5.5s | 0.1s | 779 | 107 |
| GPT-5.6 LunaCodex CLI 0.159.0, medium effort | 2 of 2 | 10.4s | 2.0s | 7.0s | 1.4s | 1.0s | 1.5s | 1.6s | 286 | 41 |
| Claude Sonnet 5Claude Code 2.1.284, high effort | 2 of 2 | 11.8s | 0.4s | 11.0s | 0.4s | 1.5s | 1.0s | 5.8s | 782 | 74 |
| GPT-6 LunaCodex CLI 0.159.0, medium effort | 2 of 2 | 14.5s | 2.7s | 9.0s | 2.7s | 0.9s | 2.5s | 3.0s | 182 | 20 |
| GPT-6 AstraCodex CLI 0.159.0, medium effort | 2 of 2 | 15.3s | 1.9s | 11.9s | 1.5s | 1.9s | 2.0s | 6.1s | 246 | 21 |
| Claude Opus 5.5Claude Code 2.1.284, medium effort | 2 of 2 | 15.7s | 0.5s | 14.9s | 0.3s | 0.8s | 10.5s | 2.5s | 1,381 | 93 |
| Claude Fable 5.1Claude Code 2.1.284, high effort | 2 of 2 | 17.3s | 0.5s | 16.5s | 0.3s | 1.7s | 10.5s | 3.0s | 1,059 | 65 |
| GPT-6 SolCodex CLI 0.159.0, medium effort | 2 of 2 | 23.0s | 4.0s | 17.7s | 1.3s | 2.1s | 5.2s | 5.6s | 312 | 18 |
| GPT-6.1 SolCodex CLI 0.159.0, medium effort | 2 of 2 | 23.6s | 2.1s | 19.8s | 1.7s | 2.5s | 3.4s | 8.3s | 270 | 14 |
| GPT-5.6 SolCodex CLI 0.159.0, medium effort | 2 of 2 | 24.0s | 2.6s | 20.0s | 1.4s | 1.1s | 5.6s | 8.1s | 374 | 19 |
Method and data#
Earlier runs: launch days
These ran before round 13, with the account extras on. Some numbers here carry stars, and they differ from the explain task table above.
Sonnet 5.5 against Sonnet 5 and Opus 5.5#
All three models ran the explain task on Sonnet 5.5's launch day, in one Claude Code release. Claude Code sends each model its own instructions. Its first request held 40,414 tokens for Sonnet 5 and 28,127 for Sonnet 5.5, measured in that day's caveman runs that were not light. With no skill Sonnet 5 also took three steps where Sonnet 5.5 took two. The bigger request and the extra step explain most of why Sonnet 5 used more tokens.
Show the charts and every number
Total tokens#
Everything the model read and wrote in the whole run
Cost at API prices#
What the run's tokens cost at the model's published API prices, cached tokens at the cache price
Tokens read#
All the text sent to the model, counted each time it was sent
Read from cache#
Share of what it read that came from the cache, which is cheaper
Cache writes#
Text saved to the cache for reuse
Output tokens#
Everything it wrote: the answer, its tool calls and its hidden thinking
Hidden thinking#
The part of the output spent thinking, which the answer doesn't show
* At least half the runs behind a starred number are light. Light runs start with fewer tokens of instructions and tools than the model's other runs, and some ended with a remark about the account's connectors, so their counts differ. What light runs are.
The table scrolls sideways on a small screen.
| Model | Passed | Tokens read | From cache | Cache writes | Output | Thinking | Visible | Cost at API prices |
|---|---|---|---|---|---|---|---|---|
| Claude Sonnet 5.5Claude Code 2.1.284, medium effort, Claude Code's default | 2 of 2 | 41,973* | 75%* | 11,233* | 911* | 37* | 875* | $0.0602* |
| Claude Sonnet 5Claude Code 2.1.284, high effort, Claude Code's default | 2 of 2 | 115,001 | 82% | 20,610 | 842 | 27 | 815 | $0.110 |
| Claude Opus 5.5Claude Code 2.1.284, medium effort, Claude Code's default | 2 of 2 | 34,984* | 74%* | 9,248* | 1,482* | 424* | 1,058* | $0.109* |
GPT-6.1 Sol against GPT-6 Sol and Opus 5.5#
GPT-6.1 Sol ran on 29 September 2026, the day it came out, beside GPT-6 Sol and Opus 5.5. On this task it reported no hidden thinking at all.
Show how I ran it, the charts and every number
GPT-6.1 Sol and GPT-6 Sol ran the same task in Codex CLI 0.159.0, the first release that runs GPT-6.1 Sol on a ChatGPT login. I set both to medium effort. Opus 5.5 ran it the same day in Claude Code at its default effort, medium. Because each CLI adds its own instructions, compare what the models wrote, not what they read. Codex CLI's trace log shows it sent the effort setting on every request. For GPT-6.1 Sol, OpenAI's server reported 0 tokens of hidden thinking at medium effort, and at high effort in a check outside the counted runs. GPT-6 Sol reported hidden thinking for the same request.
Total tokens#
Everything the model read and wrote in the whole run
Cost at API prices#
What the run's tokens cost at the model's published API prices, cached tokens at the cache price
Tokens read#
All the text sent to the model, counted each time it was sent, with Codex CLI's hidden first request
Read from cache#
Share of what it read that came from the cache, which is cheaper. Counts Codex CLI's hidden first request
Cache writes#
Text saved to the cache for reuse. Codex CLI always reports 0, so for Codex CLI I count what it read that did not come from the cache
Output tokens#
Everything it wrote: the answer, its tool calls and its hidden thinking
Hidden thinking#
The part of the output spent thinking, which the answer doesn't show
* At least half the runs behind a starred number are light. Light runs start with fewer tokens of instructions and tools than the model's other runs, and some ended with a remark about the account's connectors, so their counts differ. What light runs are.
The table scrolls sideways on a small screen.
| Model | Passed | Tokens read | From cache | Cache writes | Output | Thinking | Visible | Cost at API prices |
|---|---|---|---|---|---|---|---|---|
| GPT-6.1 SolCodex CLI 0.159.0, medium effort, set by me | 2 of 2 | 48,786 | 68% | 15,314 | 267 | 0 | 267 | $0.0366 |
| GPT-6 SolCodex CLI 0.159.0, medium effort, set by me | 2 of 2 | 39,896 | 56% | 17,752 | 279 | 80 | 199 | $0.0427 |
| Claude Opus 5.5Claude Code 2.1.284, medium effort, Claude Code's default | 2 of 2 | 44,767* | 70%* | 13,649* | 1,750* | 502* | 1,248* | $0.150* |
Effort
Effort sets how much a model thinks before and between its steps.
| Model | CLI | Effort | Set by |
|---|---|---|---|
| Claude Sonnet 5 | Claude Code 2.1.278, 2.1.283 and 2.1.284 | high | Claude Code's defaultdefault |
| Claude Opus 5.5 | Claude Code 2.1.280, 2.1.283 and 2.1.284 | medium | Claude Code's defaultdefault |
| Claude Sonnet 5.5 | Claude Code 2.1.284 | medium | Claude Code's defaultdefault |
| Claude Fable 5.1 | Claude Code 2.1.283 and 2.1.284 | high | Claude Code's defaultdefault |
| GPT-5.6 Luna | Codex CLI 0.153.3 | low | set by meset by me |
| GPT-6 Luna | Codex CLI 0.156.0 | low | set by meset by me |
| GPT-5.6 Luna, GPT-6 Luna, GPT-5.6 Sol, GPT-6 Sol and GPT-6 Astra | Codex CLI 0.157.1 | medium | set by meset by me |
| GPT-6.1 Sol, GPT-6 Sol, GPT-5.6 Sol, GPT-5.6 Luna, GPT-6 Luna and GPT-6 Astra | Codex CLI 0.159.0 | medium | set by meset by me |
GPT-5.6 Luna, GPT-6 Luna, GPT-5.6 Sol, GPT-6 Sol and GPT-6 Astra appear more than once, because each row is one release and effort a model ran at.
Claude Code picks the effort for each Claude model, and I left it alone. I logged the effort in every request it sent on 28 September 2026, on Claude Code 2.1.284, the release every later Claude run used. I assume earlier releases sent the same default. Within that effort, each Claude model sets its own amount of thinking. Anthropic's own default for Sonnet 5.5 is high, but Claude Code sends medium. For the GPT models I set the effort on every run, and Codex CLI's replies confirmed it.
Definitions
- Tokens are the chunks of text a model reads and writes, about three quarters of a word each. Tokens read counts all the text sent to the model, each time it was sent. An agent sends its instructions and the conversation again with every request. That total grows with each step.
- The cache keeps text the model has already read, so the next request can reuse it faster and at a lower price. Read from cache is the share of what it read that came from there. Codex CLI's own count leaves out its hidden first request. Without it, Codex CLI's count gives 56% to 91% for these runs. Cache writes are text saved to the cache. Codex CLI always reports 0, so for Codex CLI I count what it read that did not come from the cache.
- Codex CLI's hidden first request. At the start of every run Codex CLI sends a request that loads the model's instructions and tools and writes nothing. Codex CLI leaves it out of its own token count, so I measured it once for each model and Codex CLI release and add it back.
- Output is everything the model wrote: its answer, its tool calls, such as asking to read a file, and its hidden thinking. Visible output is output minus thinking.
- Time on a run this short is mostly fixed: a second or two to start, then one to three seconds before each reply begins. A short answer is not much faster than a long one. Two runs of the same model can differ by several seconds, and a gap of a second is a tie.
- Account extras are what a subscription login loads by default. For Claude Code, they are the plugins, skills and connectors on the account's claude.ai. For Codex CLI, they are its connected apps and plugins. Round 13 turned them off. The launch-day runs under Earlier runs ran before that, with them on.
- Cost at API prices is what a run's tokens would cost at the model's published API prices: input that missed the cache at the input price, cache reads at the cache price, Claude's cache writes at the one-hour rate Claude Code uses, and output. For every Claude run it equals Claude Code's own estimate. My runs used subscriptions, which don't bill per token. Runs share a cache, so a run's cost depends partly on the runs before it. The price table lists every rate.
- Light runs are Claude Code runs that started with thousands fewer tokens of instructions and tools than the same model's other runs. They were a session's first run, before Claude Code loaded the account extras. They cost less and read fewer tokens. Some ended with a remark about the account's connectors. I removed that remark from the answers, but their token counts still include it. Round 13 reran this page's explain runs and real app work with the account extras off. A number where at least half the runs behind it are light carries a star. The Skills page's limits say how I find them.
- Effort is picked by Claude Code for Claude models and set by me for GPT models. The effort table lists every model.
- The token charts and tables show medians: the middle run, or the average of the two middle runs. The speed charts show the average of each model's timed runs. Every run is in the download below.
API prices used for cost#
The table scrolls sideways on a small screen.
| Model | Input | Cache read | Cache write | Output |
|---|---|---|---|---|
| Claude Sonnet 5 | $2.00 | $0.20 | $4.00 | $10.00 |
| Claude Sonnet 5.5 | $2.00 | $0.20 | $4.00 | $10.00 |
| Claude Opus 5.5 | $4.00 | $0.20 | $8.00 | $20.00 |
| Claude Fable 5.1 | $10.00 | $0.25 | $20.00 | $50.00 |
| GPT-5.6 Luna | $0.20 | $0.02 | none | $1.20 |
| GPT-6 Luna | $0.10 | $0.01 | none | $0.50 |
| GPT-5.6 Sol | $4.00 | $0.40 | none | $20.00 |
| GPT-6 Sol | $2.00 | $0.20 | none | $10.00 |
| GPT-6.1 Sol | $2.00 | $0.10 | none | $10.00 |
| GPT-6 Astra | $10.00 | $1.00 | none | $50.00 |
Checked 29 September 2026. US dollars per million tokens, standard tier. OpenAI's rates are its short-context rates, which cover every request in these runs. OpenAI lists cache writes only for long context. Claude cache writes at the one-hour rate Claude Code uses, 2x input. Sources: OpenAI's pricing, Anthropic's pricing and Anthropic's caching docs.
Changes
Changes since the Models page began on 28 September. The Skills page has the full list.
- Added the Models page at /numbers. It shows every model's input tokens, cache reads and writes, output, hidden thinking, time and output per second side by side, for round 8 and for each task and skill in rounds 1 to 7.
- Codex CLI token counts now include the hidden first request Codex CLI leaves out of its own count, 9,751 to 12,371 tokens a run, measured for each model and CLI release. Codex CLI cache writes, which Codex CLI always reports as 0, are now the prompt tokens it did not read from cache.
- Round 9 ran Claude Sonnet 5.5 on its launch day, with Sonnet 5 rerun on the same Claude Code release the same day, on the explain task.
- Opus 5.5 ran the explain task twice on Claude Code 2.1.284, beside Sonnet 5.5 and Sonnet 5 on the same release. The Models page adds a total tokens chart and link anchors on every section and chart.
- Added an effort table: the effort Claude Code sent for each Claude model, logged from its requests, and the effort I set for each Codex CLI model.
- Timed every model on the explain task, two runs each, with the same method for Claude Code and Codex CLI: total time split into getting ready, waiting on the model and running tools, plus time to first token and when the written answer starts. The Models page now opens with a sortable leaderboard.
- The site now has two pages. Skills tests each skill against the same task with no skill, and Models shows every model with no skill. The per-skill token charts moved to Skills and the effort table to Models, and old links forward to the new place.
- The Sonnet 5.5 chart on the Models page now names the effort each model ran at: medium for Sonnet 5.5 and Opus 5.5, high for Sonnet 5, each Claude Code's default.
- The Models page now opens with the leaderboard under a plain title and one sentence on the task. The table keeps time to finish, tokens used and share from cache, and fits a phone; first token and answer start are in the speed charts below it.
- GPT-6.1 Sol came out during OpenAI's DevDay, and I ran it that day at medium effort on Codex CLI 0.159.0, the first release that runs it on a ChatGPT login: 2 explain runs with no skill, 2 with caveman and 2 timed runs. GPT-6 Sol reran on the same release and Opus 5.5 on Claude Code 2.1.284 the same day. All three are on the leaderboard and the overall chart with these runs.
- Every run on the Models page now shows what its tokens cost at the model's published API prices, cached tokens at the cache rate, with the price table and its sources. For Claude it matches Claude Code's own estimate on all 194 Claude runs.
- The Models page now says what the test is, under the leaderboard: the explain task, how it passes, and that it writes no code. The coding tasks are on the Skills page.
- Added real app work to the Models page: a bug hunt with 20 planted bugs and a tickets test with 17 hidden tests, in a small TypeScript web app, 3 runs each for Opus 5.5, Sonnet 5.5, GPT-5.6 Sol, GPT-6 Sol and GPT-6.1 Sol, graded by hidden tests, with cost at API prices and time.
- Correction: the effort table left out two settings the runs used, Claude Fable 5.1 on Claude Code 2.1.284 in the timed runs and GPT-5.6 Sol on Codex CLI 0.159.0 in the real app work. Both are listed now, and QA fails if a run's release is ever missing again.
- Redesigned both pages. The Skills page opens with the count of claimed cuts that held, one chart of each claim against the measured cut, and each skill's claim, finding and one-word verdict. The Models page leads with real app work. Text runs the full page width, long detail sits in folds, charts save as images with CSV, and bar charts show bars only.
- Correction: some Claude Code runs started with thousands fewer tokens of instructions and tools. They ran before Claude Code had synced the account's claude.ai plugins and skills and loaded its connectors, so they cost less and count fewer total tokens. A number where they make up at least half the runs now carries a star, a new limit explains how I found them, and I removed a remark about the connectors from the published answers. No claim check changes. The count of Claude runs in the 29 September cost entry is corrected from 254 to 194; it had counted Sonnet 5.5's runs twice.
- One GPT-5.6 Sol tickets run edited an existing test file, which the prompt said not to do. The grader restores the original tests first, so its 17 of 17 stood. Round 13 replaced that run, and no rerun edited a test.
- Round 13 reran the explain task for every model on the Models page, every Claude model's explain runs, all 30 real app work runs and the timed runs, all with the account extras off. Those numbers no longer carry stars, and Claude's explain-task token counts and costs dropped. caveman now cuts Sonnet 5.5's explanations 28%, where round 9 showed 13%, and Sonnet 5's 72%, still above its 65% claim. All five models still passed all 17 hidden ticket tests, and Opus 5.5 cost more, $0.96 a ticket run.
- Redesigned the findings, tables and colors. Each finding has its own framed chart with horizontal bars and a saved image, the skill cards and real-work tables line up and sort like the Models page table, each vendor keeps one color family, and dark mode is warm.
- Correction: the Skills page gave Sonnet 5.5's release day as 30 September, when it came out on 28 September. Two table notes still said cost covered Claude only, and the effort table left out the Codex CLI 0.159.0 row for three GPT models. All three are fixed, and QA now reads every run group.
- The chart for any measure and task now uses the same framed horizontal bars as the findings, with a saved image. The Models page notes how much round 13 moved its numbers and has its own list of changes, and on a phone its table stays a table.
- The chart for any measure and task is now its own section, Every measure and task. Charts with three skills stack at full width with a labeled scale, bars from failed runs show hollow, and the Models page notes every model the round 13 timing slowed.
- Correction: 12 Claude coding runs also started light. Until now I had checked only the explain task. They now carry a star in the run list, and no median has half its runs light. The Models table title gives the time spread with and without the slow OpenAI day, and the cost panels say their range adds up each task's cheapest to costliest run.
- On a phone, bar charts put each model's name above its bar. Saved images keep their panels side by side, the Models token charts are horizontal bars, and the Skills page adds a table of every model's cut under the top chart.
- The GPT models that ran slower in round 13's timing carry a dagger in the explain table and the time chart, and the real app work note says their runs slowed the same day. Costs use one format on every table and chart.
- Bar charts with the name above each bar leave a clear gap between rows, and their lines no longer cross a label. Grid cells that could be chance are outlined like the bars. The cut table names every model and marks the one cut that met its claim.
- Costs use one format everywhere, three significant figures under ten cents and three decimals above. The cost-vs-time title and dots mark the slowed GPT models, and the run list adds round 10.
- The cut table fits tablets and phones, where each claim becomes a block of model and cut pairs. Saved images explain their daggers, small cost bars get their own scale, and tables and charts in folds share one title style.
- Round 8's table now stars its light medians. The real-work title and note use the table's cost figures, and the cost-vs-time title compares cost only. The page says a claim held, in one word, everywhere.
- Output and thinking counts carry the light-run star too. The slowed GPT models are marked § instead of a dagger. The cut table says whether a cell was not compared or failed, and dates and model names no longer break across lines on phones.
- The tick strip and the byline are gone: the chart under the intro shows the same checks, and its source line carries the date.
- The real app work chart and the Models board views drop the rings and the scores on each dot, and the rerun note is two sentences.
- The real app work chart says what its tests can't do: every model scored at or near full marks, so they show cost and time on a small job, not which model is best.
- A dot on the real app work chart names its score when it fell short of full marks, as Claude Sonnet 5.5's bug hunt dot does at 19 of 20.
- New page, GPT-6.1 Sol timed three days running: one short Python file explained on launch day, the day after, and on 1 October after OpenAI said it added capacity. The Models board keeps the 30 September runs.
- The GPT-6.1 Sol page says what the test is up front and groups the chart by day. The ChatGPT apps and plugins setting moved to the method, in plain words.
- The GPT-6.1 Sol page adds four runs from the afternoon of 1 October, after @sama said it should be much better: 13.5 to 16.8 s, faster than that morning and still slower than launch day.
- The GPT-6.1 Sol page adds four runs from the afternoon of 3 October, after a 2 October post on X still called it slow: 15.6 to 26.6 s, slower by median than the afternoon of 1 October. Corrected 5 October: the two afternoons' runs overlap.
- The GPT-6.1 Sol page adds four runs from the afternoon of 4 October: 13.2 to 22.9 s, all slower than launch day.
- The GPT-6.1 Sol page adds four runs from 5 October, before OpenAI's speed post: 14.1 to 18.6 s, all slower than launch day.
- The GPT-6.1 Sol page adds four runs from 5 October, minutes after OpenAI said its speed change would be felt within two hours: 14.3 to 19.5 s.
- Readers were right that the GPT-6.1 Sol page's tokens a second was not TPS. It divides output by the time spent waiting on the model's replies, so the column is now tokens a second of waiting.
- Headlines say a setup cut more, won or held its claim only when its runs sit apart from the other side's. Where runs overlap they now say by median, so three headlines changed, and the all-tasks totals add up each side's full run range.
- The GPT-6.1 Sol page adds four runs from a 5 October retest, still inside the two hours OpenAI gave: 11.3 to 20.8 s, three of them inside launch day's range.
- A TPS column briefly showed a rate from an OpenAI server-side timing field. That field times the engine, not delivery: on launch day it read 3.7 to 3.9 times the rate text reached my machine. It came down. A TPS proxy column now counts the pieces of answer text Codex logged a second, pauses included.
- Two test runs of the updated timing script on 5 October are listed under How I ran it on the GPT-6.1 Sol page, not charted.
- The GPT-6.1 Sol and ASD-STE100 pages said each run got a fresh container. Each run gets a fresh copy of the project; runs share the bench's container. The Sol page now says a run passes when its answer contains both "month" and "ValueError".
- The GPT-6.1 Sol page's intro said posts on X on 3 October called it slow again. No such post could be found, so it now cites the 2 and 4 October posts that did.
- The GPT-6.1 Sol page adds four runs from 5 October, after the two hours OpenAI gave: 11.4 to 17.7 s, streaming at 39.2 to 52.9 text pieces a second.
- Readers asked how launch day was as quick with a slower stream. The GPT-6.1 Sol page now says why: the answer's stream is under half of every run, launch day's answers were shorter, and the runs since that made two model requests, like launch day, finished inside launch day's range.
- Readers asked to sort the GPT-6.1 Sol run table. Click any column heading: Time puts the fastest first, TPS proxy the highest first, a second click reverses, and Day restores day order.
- The GPT-6.1 Sol chart took more than a page. It is now one row per sitting in two panels, TPS proxy and seconds, each a median with launch day and the sittings since the 5 October retest picked out; the 5 October sittings are named by OpenAI's post, not the time of day; every run stays in the table. The intro is one line, and the background folds under How I ran it.
- Readers asked whether the GPT-6.1 Sol page has true TPS. The table now has TPS by OpenAI's own token count for the eight runs since 5 October's retest: the answer's output tokens over the seconds from its first text event to its last, a rate observed at my machine. Earlier rows have no count and stay blank; two script test runs also logged it, at 36.2 and 43.4, and are in the CSV. The text-event rate is now labeled text pieces a second.
- Readers couldn't find the GPT-6.1 Sol speed page from the main pages. It now has its own tab in the bar on every page (Sol on a phone), and the Models page's time section links it.
- The GPT-6.1 Sol page adds four runs from 6 October: 10.3 to 15.7 s, the fastest run yet at 10.3 s, and 43.4 to 51.2 tokens a second by OpenAI's count. Every sitting from 5 October's retest on now counts as since the retest.
- A reader asked for Claude Opus 5.5 and Claude Sonnet 5.5 beside the GPT-6.1 Sol test. The Sol page now shows them timed the same day and task, 4 runs each, start to finish in each model's own tool: Sonnet 5.5 7.8 to 9.8 s, Sol 10.3 to 15.7 s, Opus 5.5 15.2 to 17.7 s. No TPS for the Claude models: this bench's Claude Code capture does not give a streaming rate yet.
- Readers loved the speed race videos on X, so the GPT-6.1 Sol page now has the same race, drawn live from every run: one lane per sitting, all leaving together at 5x with real seconds on the clock. It plays once when it scrolls into view and has a Replay button; with reduced motion it waits for Play.
- Every section of the GPT-6.1 Sol page has a heading with a # link that copies its address, and an On this page row jumps between them: every sitting, the race, same day with other models, and how I ran it.
- Readers said the race's dots were softer than the videos'. They are now drawn at full screen density with a sharp edge, a glow sized to the dot and the videos' white highlight.
- The site now counts visits with Cloudflare Web Analytics, which sets no cookies. The footer, which said no analytics, now says so.
- Each page now ships its text in the HTML, so search engines and AI tools that do not run scripts can read it. The site also has a sitemap, an llms.txt and a footer line linking every page.
- Readers said the link cards on X did not show the site well. Each card now has the site's bar with its page marked, one finding and one chart, sized to read in a feed; the GPT-6.1 Sol card shows the streaming chart alone.
- The GPT-6.1 Sol page's method now says the group since 5 October's retest was set after seeing the 5 October runs, and every later sitting joins it whatever it shows. The link cards have larger chart text.