## Round 23: Claude Haiku 5.5 against GPT-6 Luna at medium and xhigh (logged 2026-10-07, before its counted runs) - Site numbering: round 25. Reason: Anthropic released Claude Haiku 5.5 on 7 October at the same list price per token as GPT-6 Luna for prompts up to 100k tokens. Bren asked for the full suite on both, at medium and at xhigh effort. - Setting: Claude Code 2.1.284 with `--effort medium` or `--effort xhigh` (its default for Haiku 5.5 is high, seen in the smoke runs, while Opus 5.5 and Sonnet 5.5 default to medium); Codex CLI 0.159.0 with model_reasoning_effort medium or xhigh (Luna's default is medium). Apps, plugins and MCP servers off (scripts/run.sh defaults). Image rebuilt to add claude-haiku-5-5 to bench.py HOST_OF. - Smoke runs, not counted, reported under the method: runs-timing/timing-smoke-haiku55-20261007.jsonl (2 runs, Haiku at its default) and runs-timing/timing-smoke-h2h-20261007.jsonl (2 runs per model and effort, plus 2 Sonnet 5.5 runs that test the burst check). - Measurement change, made after the smoke runs and before any counted run: usage_proxy.js and timing.py now count the bursts an answer's text arrives in (text more than 5 ms after the previous text starts a new burst) and the longest gap. timing.py adds answer_visible_out (the answer reply's output tokens minus its thinking or reasoning tokens, each vendor's own count) and answer_visible_tps (that over the seconds from the first text event to the last). - Streams, every run counted and reported, failed ones included: 1. Speed: scripts/timing.py on the explain task, one run at a time, alone (no other stream running), 6 runs per model and effort, as three calls of 2 per cell in the order haiku-medium, luna-medium, haiku-xhigh, luna-xhigh, three times, into runs-timing/timing-round25.jsonl. 2. Skills: bench.py, tasks T1 to T4, conditions control, placebo, caveman, ponytail and karpathy, 3 reps, so 60 runs per model and effort (240), into runs-round25-haiku-medium, runs-round25-haiku-xhigh, runs-round25-luna-medium and runs-round25-luna-xhigh, seeds 2150 to 2153. 3. ASD-STE100: the ste condition on T3, 3 reps per model and effort (12 runs), into runs-round25-ste--, seed 2154. The no-skill side is stream 2's T3 control runs (same day, release and setting). 4. Real app work: B1 and B2, 3 reps per model and effort (24 runs), one container per run, seed 4000 as before, graded by the same hidden tests into runs-quality. - Measures, fixed now: - Speed: start-to-finish seconds; visible TPS. A run's TPS counts only if its answer text arrived in at least 20 bursts; otherwise it is shown as held, not a streaming rate. Anthropic and OpenAI count tokens differently, so each TPS is in its own vendor's tokens, and the page says so. - Skills, ASD-STE100 and real app work: the site's existing measures and rules for those pages. - Cost at each vendor's list price for prompts up to 100k tokens (Haiku 5.5: $0.10 input, $0.01 cache read, $0.125 cache write, $0.50 output; GPT-6 Luna short context: $0.10, $0.01, $0.125 cache write, $0.50). Codex's usage log has no cache-write count, so Luna's cost leaves cache writes out, and the page says so. - Comparisons use the site's overlap rule: faster, slower, more or fewer outright only when no runs overlap, otherwise by median. - Review before publishing: Fable 5.1, GPT-6.1 Sol in Codex and Grok 4.7 in Cursor review the findings, the page and the video (Bren). - Amendment, logged 2026-10-07 during the speed stream and before the skills streams: Haiku 5.5's cost uses the cache-write rate Claude Code's requests pay, the one-hour rate of $0.20 per million tokens (Anthropic's prompt-caching docs: 2x input), as every earlier Claude round does; the $0.125 above is the five-minute rate. Anthropic prices a request over 100k prompt tokens at the higher tier ($0.50 input, $0.05 cache read, $1 one-hour cache write, $2.50 output), so any Haiku request over 100k is priced at that tier from its own usage. For Luna, OpenAI's pricing page now lists a short-context cache-write price ($0.125). Whether it applies to Codex's prompts is checked at source by the reviewers before any Luna cost is published, and the page says which rule was used. - Amendment, logged 2026-10-07 after the speed stream and before any result is published: one GPT-6 Luna medium speed run failed (it answered in 1 request without reading the file, 31 visible tokens, 3.8 s). The plan did not say how a failed run counts in a speed comparison, so this is decided after seeing it: speed comparisons (medians, ranges, the overlap rule, the video's lane ranges) use passed runs; the failed run is listed with its time and drawn as a hollow dot, and the page says one Luna medium run failed. Every other stream keeps its page's existing rules. - Amendment, logged 2026-10-07 during the streams: Claude Code's changelog shows Haiku 5.5 support was added in 2.1.293 ("Added Claude Haiku 5.5 (claude-haiku-5-5)"); this round pins 2.1.284, which accepts the model but prices it at Opus 5.5's rates and runs it at high effort by default. Check: claude-haiku-5-5's speed test rerun on Claude Code 2.1.293 (6 runs at medium, 6 at xhigh, explain task, into runs-timing/timing-round25-cc293.jsonl), compared with the 2.1.284 speed runs by the overlap rule, plus 2.1.293's default effort for Haiku 5.5 (2 runs, no --effort). If the medians or the default differ, Haiku 5.5's streams are rerun on 2.1.293 and both sets are reported; if not, the 2.1.284 runs stand and the page says which release ran them and why. This check runs while the skills and real-work streams are still running, so it is not alone the way the counted speed stream was, and the page says so. - Result of the 2.1.293 check: Haiku 5.5 on Claude Code 2.1.293 took 6.6 to 9.0 s at medium (median 7.6) and 11.4 to 14.9 s at xhigh (median 14.4), against 2.1.284's 6.4 to 7.9 s (7.3) and 8.9 to 16.1 s (15.2). The runs overlap at both efforts, so the 2.1.284 runs stand. Visible TPS by median: 225 and 246 on 2.1.293, 216 and 257 on 2.1.284. Without --effort, 2.1.293 ran Haiku 5.5 at medium (2 runs), so on current releases both tools default to medium. Every check run is kept in runs-timing/timing-round25-cc293*.jsonl. - Amendment, logged 2026-10-07 after the speed stream, at Bren's request: one extra GPT-6 Luna medium speed run (the cell's seventh), so the Luna medium lane of the X video has six passed runs like the others. The failed run stays in the data, the page and the table; the video leaves it out and its lane says so. This run went while the skills and real-work streams were still running, so it was not alone the way the counted stream was. - Decisions logged 2026-10-07 after review by Fable 5.1 and GPT-6.1 Sol, before the page is published: - The 2.1.293 check's trigger is waived, and the page says so. Its default effort did differ (high on 2.1.284, medium on 2.1.293), but every counted run set --effort explicitly. Thinking tokens in the answer reply match across releases by median (503 against 500 at medium, 1,955 against 2,087 at xhigh), and the speed runs overlap. Only speed was checked, so skills and real-work token counts on 2.1.293 could differ. - Haiku 5.5 runs whose last prompt passed 100k tokens cannot be priced request by request as planned, because the run records keep only the last request's usage. The page shows a range for them, from every request at the lower tier to every request at the higher one. Future rounds log per-request usage. - GPT-6 Luna's costs on this page are a range: uncached prompt tokens from the $0.10 input rate to the $0.125 short-context cache-write rate OpenAI lists, because Codex logs no cache-write count. - The X reply video (v3) shows the six passed Luna medium runs without the lane note planned above, at Bren's request. The page lists all seven runs, and a reply under the video discloses the failed run and its rerun. - Added after review, logged 2026-10-07: a cell's TPS median is shown only when at least three of its runs streamed in 20 or more bursts (Fable 5.1), so GPT-6 Luna xhigh, with two, shows none. The four Haiku 5.5 real-work runs priced exactly had last prompts of 72,332 to 99,754 tokens and no compaction in their output, so every earlier request was smaller. - Check, logged 2026-10-07 after publishing, at Bren's request: GPT-6 Luna on Codex CLI 0.161.0 (newest), against the round's 0.159.0. On 0.159.0, 4 of 6 xhigh and 3 of 6 medium speed runs delivered their answer in 6 to 17 bursts (no streaming rate), and 5 Luna medium runs answered without reading the file. Check: scripts/timing.py, 6 runs at medium and 6 at xhigh on 0.161.0, alone, into runs-timing/timing-round25-cx161.jsonl. If 0.161.0 streams in 20+ bursts on most runs or stops the no-read answers, Luna's speed stream is rerun on 0.161.0 as its own logged set and the page says which release each number comes from; every run stays reported. - Result of the Codex 0.161.0 check: the same pattern as 0.159.0. At medium, 5 of 6 runs passed (one answered without reading the file again); 3 streamed in 68 to 76 bursts at 57.7 to 59.6 TPS and 2 arrived in 7 or 8 bursts. At xhigh, all 6 passed; 2 streamed (59.1 and 60.9 TPS) and 4 arrived in 6 to 11 bursts. Seconds overlapped 0.159.0's at both efforts (medium 5.2 to 8.5 s, xhigh 6.0 to 11.0 s). So the CLI release is not the cause, and Luna's speed stream is not rerun. The page shows GPT-6 Luna xhigh's TPS from the four runs that streamed across both releases (59.1 to 61.5, median 61.0), saying so, and lists the check runs. - Amendment, logged 2026-10-07 before any of the runs it names, at Bren's request ("we want clean runs"): - Why GPT-6 Luna's TPS was thin: every Luna speed run whose answer arrived in 6 to 11 bursts (15 of 25 across 0.159.0 and 0.161.0, both failed runs aside) got a per-turn timing event from OpenAI with only pre_inference_ms and total_turn_time_s. Every run that streamed in 61 to 90 bursts also got the engine fields (engine_service_tbt about 18.5 ms). The bunched runs were faster before any text arrived (answer_start 1.1 to 1.7 s at xhigh, against 2.8 to 4.1 s) and over the whole turn. So OpenAI served those runs another way; it is not Codex or this machine. Keeping only the runs that streamed left Luna's TPS from its slower runs. - TPS stream (new): scripts/timing.py with TIMING_TASK=T3-long, the explain task asking for a walkthrough of at least 600 words, so an answer that arrives in bunches still arrives in 20 or more. 6 runs per model and effort, three calls of 2 per cell in the order haiku-medium, luna-medium, haiku-xhigh, luna-xhigh, alone, Claude Code 2.1.284 and Codex CLI 0.159.0, same settings as the speed stream, into runs-timing/timing-round25-long.jsonl. Measure: answer_visible_tps, counted only with 20 or more bursts; each cell's median and range over its passed runs. The page's TPS comes from this stream; the explain task keeps seconds, waiting and answer length. Luna runs are also split by whether OpenAI's timing event carried engine fields, and the page says how many came each way. - Failed runs: every failed Luna explain run replied that no file-reading tool was available and asked for the file. Check, not counted: scripts/tools_probe.py runs Luna at medium on the explain task with Codex's trace log, and keeps only the tool names in each request Codex sends and whether the run passed, to see whether those replies come with the tools missing from the request. - Reporting: results, charts, the race and the run table use passed runs only. The method says how many runs failed, where, and what they replied; every run stays in the data files. - Amendment, logged 2026-10-07 before the TPS stream: with Codex's full trace log on, Codex 0.159.0 now keeps running at full CPU after its turn ends (seen in the probe today, then in a plain timing.py run; the counted speed runs ended normally, 5 to 11 s). The TPS stream and the probe therefore log only tungstenite::protocol at trace (TIMING_RUST_LOG), the WebSocket message lines timing.py reads, so the timestamps are the same ones from a smaller log. Probe so far: 12 runs at medium, all passed, every request echoed back with the same 11 tools. - Result of the tools probe (runs-timing/tools-probe-round25.jsonl): 36 GPT-6 Luna runs at medium on Codex 0.159.0. 33 passed. 3 failed the same way as before, replying that the tools available could not read the file, with no tool call. On every request, the failed ones included, OpenAI's response.created event listed the same 11 tools, Codex's shell (functions.exec) among them, and the failed runs' first real prompts were 12,290 to 12,294 tokens, the same as the passing runs' 12,290 to 12,296. So the failures are the model's, not a missing tool. All 36 runs came without OpenAI's engine timing. - Result of the TPS stream (runs-timing/timing-round25-long.jsonl): all 24 runs passed and every answer arrived in 20 or more bursts (Haiku 5.5 247 to 372, GPT-6 Luna 53 to 79), so every run counts. Visible TPS by median: Haiku 5.5 239.2 at medium (228.8 to 250.9) and 297.4 at xhigh (284.8 to 319.5); GPT-6 Luna 121.9 at medium (116.2 to 142.3) and 125.6 at xhigh (109.3 to 143.1). No runs overlap at either effort. Every GPT-6 Luna run in this stream came without OpenAI's engine timing, the way that arrived in bunches on the explain task, so the explain task's Luna runs that streamed (57.7 to 61.5) are the other way, and the page says both. - Correction, logged 2026-10-07 after review (Fable 5.1, GPT-6.1 Sol, Grok 4.7): the amendment above that starts "Why GPT-6 Luna's TPS was thin" miscounted. It is 13 of the 23 passed Luna speed runs across 0.159.0 and 0.161.0 (14 of all 25) that arrived in 6 to 11 bursts without the engine fields. Its "faster" and "slower runs" hold at xhigh, where the bunched runs reached their first words sooner (1.1 to 1.7 s against 2.8 to 4.1 s) and finished sooner; at medium the two groups' times overlap. The TPS stream's Luna runs all came without the engine fields, so it does not time the other way; the page says so, and gives that way's explain-task TPS (57.7 to 61.5) only in the method.