Karpathy's ASD-STE100 tip, measured

Karpathy's tip: "Ask your LLM to explain something in ASD-STE100", a controlled language from aerospace maintenance manuals with "heavy constraints on clean writing style that I often find a lot more readable." I gave six models his tip as a one-line instruction and measured what it changed on the explain task the Skills page uses.

Asking for ASD-STE100 shortened answers on three of six models, clearly on two, and GPT-6 Luna failed every run with it.

Change in visible output tokens against no instruction, median of 3 runs a side, on the explain task: each model reads one 53-line Python file and explains it.

50% fewer40% fewer30% fewer20% fewer10% fewerNo instruction10% more20% moreClaude Sonnet 5: 652 visible output tokens with the instruction, 956 withoutClaude Sonnet 532% fewerClaude Opus 5.5: 1102 visible output tokens with the instruction, 988 withoutClaude Opus 5.512% more, could be chanceClaude Sonnet 5.5: 733 visible output tokens with the instruction, 864 withoutClaude Sonnet 5.515% fewerGPT-5.6 Luna: 250 visible output tokens with the instruction, 236 withoutGPT-5.6 Luna6% more, could be chanceGPT-6 Luna: 52 visible output tokens with the instruction, 186 withoutGPT-6 Lunafailed every run with the instructionGPT-6.1 Sol: 247 visible output tokens with the instruction, 301 withoutGPT-6.1 Sol18% fewer, could be chance

Could be chance: some runs with the instruction were longer than some without it, or the other way round. GPT-6 Luna is left out: all three of its runs with the instruction failed the task, and all three without it passed.

The table scrolls sideways on a small screen.

Visible output and sentence length with and without the ASD-STE100 instruction
ModelOutput, no instructionOutput, ASD-STE100ChangeShort sentences, no instructionShort sentences, ASD-STE100Passed
Claude Sonnet 5Claude Code 2.1.284, default effort95665232% fewer53%100%6 of 6
Claude Opus 5.5Claude Code 2.1.284, default effort9881,10212% more95%100%6 of 6
Claude Sonnet 5.5Claude Code 2.1.284, default effort86473315% fewer100%100%6 of 6
GPT-5.6 LunaCodex CLI 0.159.0, low effort2362506% more100%100%6 of 6
GPT-6 LunaCodex CLI 0.159.0, low effort18652not compared100%100%3 of 6
GPT-6.1 SolCodex CLI 0.159.0, medium effort30124718% fewer100%100%6 of 6

Output counts visible output tokens, what the model writes minus its hidden thinking. Short sentences is the median share of an answer's sentences with 25 words or fewer, STE's limit for descriptive writing, counted roughly: code blocks left out, a sentence ending at a full stop or a line break. Passed counts runs with and without the instruction whose answer named ValueError and month.

Source: skill-evals.brenhq.com, updated 2 October 2026. 18 runs with the instruction and 18 without, on 2 October and in round 13, account apps and plugins off.

caveman shortened answers the most on four of five models with ASD-STE100 results, and ASD-STE100 shortened them more than my control file on two of five by median, clearly on one.

Change in visible output tokens against no instruction on the explain task, median of 3 runs a side. Each setup is against no-instruction runs on its own CLI release and setting, so these are separate experiments side by side, not one race under one setting. My control file is 365 words of generic project guidelines I wrote: not a token-saving skill, and not a neutral placebo either, since it asks for a review pass and edge-case handling.

GPT-6 Luna has no bars: all three of its runs with ASD-STE100 failed the task, and one of its three no-skill runs on the Skills page failed, so the skills have nothing to compare against.

ASD-STE100

20% moreNo instruction20% less40% less60% lessASD-STE100 on Claude Sonnet 5: 652 tokens, against 956 tokens with no instructionClaude Sonnet 532% fewerASD-STE100 on Claude Opus 5.5: 1,102 tokens, against 988 tokens with no instructionClaude Opus 5.512% more, could be chanceASD-STE100 on Claude Sonnet 5.5: 733 tokens, against 864 tokens with no instructionClaude Sonnet 5.515% fewerASD-STE100 on GPT-5.6 Luna: 250 tokens, against 236 tokens with no instructionGPT-5.6 Luna6% more, could be chanceASD-STE100 on GPT-6 Luna: runs failedGPT-6 LunafailedASD-STE100 on GPT-6.1 Sol: 247 tokens, against 301 tokens with no instructionGPT-6.1 Sol18% fewer, could be chance

caveman

20% moreNo instruction20% less40% less60% lesscaveman on Claude Sonnet 5: 272 tokens, against 956 tokens with no instructionClaude Sonnet 572% fewercaveman on Claude Opus 5.5: 728 tokens, against 988 tokens with no instructionClaude Opus 5.526% fewercaveman on Claude Sonnet 5.5: 621 tokens, against 864 tokens with no instructionClaude Sonnet 5.528% fewercaveman on GPT-5.6 Luna: 223 tokens, against 226 tokens with no instructionGPT-5.6 Luna1% fewer, could be chancecaveman on GPT-6 Luna: runs failedGPT-6 Lunafailedcaveman on GPT-6.1 Sol: 184 tokens, against 286 tokens with no instructionGPT-6.1 Sol36% fewer

ponytail

20% moreNo instruction20% less40% less60% lessponytail on Claude Sonnet 5: 733 tokens, against 956 tokens with no instructionClaude Sonnet 523% fewer, could be chanceponytail on Claude Opus 5.5: 967 tokens, against 988 tokens with no instructionClaude Opus 5.52% fewer, could be chanceponytail on Claude Sonnet 5.5: 774 tokens, against 864 tokens with no instructionClaude Sonnet 5.510% fewer, could be chanceponytail on GPT-5.6 Luna: 265 tokens, against 226 tokens with no instructionGPT-5.6 Luna17% moreponytail on GPT-6 Luna: runs failedGPT-6 Lunafailedponytail on GPT-6.1 Sol: 211 tokens, against 286 tokens with no instructionGPT-6.1 Sol26% fewer

karpathy-skills

20% moreNo instruction20% less40% less60% lesskarpathy-skills on Claude Sonnet 5: 765 tokens, against 956 tokens with no instructionClaude Sonnet 520% fewer, could be chancekarpathy-skills on Claude Opus 5.5: 1,046 tokens, against 988 tokens with no instructionClaude Opus 5.56% more, could be chancekarpathy-skills on Claude Sonnet 5.5: 857 tokens, against 864 tokens with no instructionClaude Sonnet 5.51% fewer, could be chancekarpathy-skills on GPT-5.6 Luna: 248 tokens, against 226 tokens with no instructionGPT-5.6 Luna10% morekarpathy-skills on GPT-6 Luna: runs failedGPT-6 Lunafailedkarpathy-skills on GPT-6.1 Sol: 253 tokens, against 286 tokens with no instructionGPT-6.1 Sol12% fewer, could be chance

My control file

20% moreNo instruction20% less40% less60% lessMy control file on Claude Sonnet 5: 732 tokens, against 956 tokens with no instructionClaude Sonnet 523% fewer, could be chanceMy control file on Claude Opus 5.5: 980 tokens, against 988 tokens with no instructionClaude Opus 5.51% fewer, could be chanceMy control file on Claude Sonnet 5.5: 788 tokens, against 864 tokens with no instructionClaude Sonnet 5.59% fewer, could be chanceMy control file on GPT-5.6 Luna: 220 tokens, against 226 tokens with no instructionGPT-5.6 Luna3% fewer, could be chanceMy control file on GPT-6 Luna: runs failedGPT-6 LunafailedMy control file on GPT-6.1 Sol: 225 tokens, against 286 tokens with no instructionGPT-6.1 Sol21% fewer

Source: skill-evals.brenhq.com. ASD-STE100 ran on 2 October in Claude Code 2.1.284 and Codex CLI 0.159.0, with the account apps and plugins off. Its no-instruction runs are fresh ones for GPT-5.6 Luna, GPT-6 Luna and GPT-6.1 Sol and, for Claude Sonnet 5, Claude Opus 5.5 and Claude Sonnet 5.5, their 30 September runs on the same release and setting. The skills and my control file are the Skills page's runs: Claude Sonnet 5, Claude Opus 5.5 and Claude Sonnet 5.5 on 30 September in Claude Code 2.1.284; GPT-6.1 Sol on 29 September in Codex CLI 0.159.0; GPT-5.6 Luna and GPT-6 Luna on 22 September in Codex CLI 0.153.3 and Codex CLI 0.156.0. The Claude ones had the account apps and plugins off, the GPT ones on.

How I ran it#

The instruction is one line, installed as CLAUDE.md for Claude Code and AGENTS.md for Codex CLI, the same way the Skills page installs every skill: "When you explain something, write the explanation in ASD-STE100, the Simplified Technical English controlled language specification." Each run starts the CLI on a fresh copy of a small Python project in the bench's Docker container and asks the model to explain how its 53-line file adds up monthly totals and what happens with a bad date. The account's apps and plugins are off.

The Claude models' runs without the instruction are their round 13 explain runs, on the same Claude Code release and setting. The GPT models got new runs without the instruction beside their STE runs, because their earlier ones had the apps and plugins on. Both Lunas ran at low effort and GPT-6.1 Sol at medium, as on the Skills page.

Karpathy's claim is readability, which this does not measure. The page measures length and one STE rule a script can count. The STE dictionary of approved words is ASD copyright and is not checked. A run with the instruction cost $0.00126 to $0.0897 at API prices. The run records are in results.json, under ste_runs and round13_runs.