I’ve been a ride-or-die Claude man for eight months now. Since Florian Brand Claudepilled me in January, I’ve used it almost every day- until Astra’s launch.
As of today I cannot think of a single metric whereby Fable 5.1 is doing better than Astra, except vibes.
Here’s my token usage since August:
I was up to just over 2 billion Claude tokens in the week of August 31 to September 6, and this week that number verges on zero.
At the same time, I’m getting just as much work done with a quarter of the tokens. Codex is just way, way more efficient in the work that I’m doing with it; seems to re-read things less, and makes more progress with fewer tokens.
I also have a habit of recording both major and minor incidents when AI goes wrong. A minor one would be sloppy copy or oversteering. A major one would be an example of Claude or Codex doing something that would endanger the viability of a project, or something so mind-bogglingly stupid that it should be rightfully called out.
Those are recorded as incident reports; here’s a graph of them by week.
For every 100 messages I sent, Claude had 1.7 formal reports. That's a lot. It also averaged 3.8 minor issues.
Across the same period, Codex had 0.55 formal reports per 100 messages, less than a third of Claude’s rate (buy about the same number of minor issues). Those mostly derive from a harness that was too tight (than you Seb) and a lack of experience working together on open-ended tasks.
This model is very, very good. It’s a pleasure to work with. Better at most tasks than Fable 5.1 and genuinely feels like a step improvement. I’m delighted by the good work that OpenAI has done and although I’m scared for the future of AI, I am definitely enjoying its benefits.
Hinton once said: “the prospect of discovery is too sweet”. I feel that way about Astra. It’s very, very good and if you haven’t really explored it yet, today is the day.
sources & notes (AI generated)
The tools and the switch. OpenAI announced GPT-6 Astra on September 3, 2026; Anthropic announced Claude Fable 5.1 on September 1. These charts compare my use of Claude Code and Codex across changing models, rather than isolating those two model releases. Astra announcement; Claude Fable; Codex; Claude Code.
My earlier Claude enthusiasm. “Dispatches from Claude Psychosis: Episode 1,” January 23, 2026, records my initial adoption of Claude Code. The account of Florian introducing me to it is my recollection. Earlier post; Florian Brand.
Token usage and efficiency. My retained local records, August 4 through the September 9, 2026 snapshot, Pacific time: Claude Code 4,747,853,304 processed tokens; Codex 865,200,671. These include cached input and output, and exclude subagents and automated sessions. Claude’s August 31–September 6 total was 2,091,983,149; September 7–9 was 41,745,473. First and last weeks are partial, and some transcripts are missing. The assessment that I am doing just as much work with a quarter of the tokens is my own experience of the switch.
Formal reports and minor corrections. The private incident files and submitted-message histories give Claude Code 29 formal reports and 65 provisional minor corrections across 1,704 messages; Codex has 3 and 24 across 546 messages. Per 100 messages, those are 1.70 and 3.81 for Claude, and 0.55 and 4.40 for Codex. Messages include approvals and steering. “Formal” means a written report; severity varies. These are reporting rates, and adding them does not establish the percentage of messages that failed.
Counting limits. Codex compiled the provisional minor counts from corrective chat interactions. Retrieval is incomplete and 40 recent candidates still need contextual review. Known report-related feedback and repeated complaints were excluded. Two shared reports and one ChatGPT Desktop report sit outside the direct counts; assigning both shared reports to Codex would make its formal rate 0.92 per 100 messages. The September 9 Chrome-associated Codex report is included, with its root cause unresolved. Raw chats and incident files remain private; the charts show aggregates.
Hinton quotation. Raffi Khatchadourian, “The Doomsday Invention,” The New Yorker, November 15, 2015. Hinton was answering Nick Bostrom’s question about why he continued AI research despite his concern about its misuse. Original reporting.
Images. The cover was generated with OpenAI’s image-generation tool from the token chart. The three charts were prepared by Codex from the local records described above.





