Contents

Today's post is a bit of an experiment.

I already have AI write things here, that part isn't new. Usually, though, I hand it something I actually touched and experienced, and it shapes that into prose. This time I skipped that step too: "collect the AI news going around today and draft a blog post." What came out, I tidied up and published.

What I actually did

Simple enough.

  1. Ask the AI to pull together today's AI news and the reaction on X
  2. Narrow the bullet list down to things I can actually get my hands on
  3. Have the AI draft it in this blog's tone
  4. Read it myself at the end and cut whatever is inflated

Honestly, at step 3 the draft came back with several clickbait-flavoured title options. Things like "the AI that OpenAI itself called dangerously strong". Reading the substance, they weren't wrong, but published as-is the numbers would have taken on a life of their own, so those went.

The main event: Claude Fable 5.1

On 1 September, Anthropic announced Claude Fable 5.1 (Mythos 5.1 for research and specialist use). The base model is positioned as a continuous improvement on Fable 5, but a few numbers stand out.

  • Terminal-Bench 4.0 rose from 42.0% to 55.8%
  • Better stability on long-running tasks, and better reporting when it gets stuck
  • API cache-read pricing dropped from 1 dollar to 0.25 dollars per million tokens, a 75% cut

Personally the price cut lands harder than the benchmark. The longer a task runs, the more often the cache gets re-read, so making that cheaper translates directly into "I can leave it running for a long time without thinking about it."

The other thing I enjoyed: alongside the release, the official account announced that usage limits had been reset. Shipping a new model and resetting the limits on the same day is generous.

Other news that came out alongside it

Fable 5.1 is the lead, but a few other things landed at the same time, so here they are as notes.

Codex 0.152.0 (OpenAI's coding tool) also shipped on 1 September. Less a headline feature than a set of quiet improvements for handing an AI long stretches of work: Vim-style search, better handling of long-running commands. Not flashy, but the kind of update that shows up for people who use it hard.

The bigger item is OpenAI's next model, "Astra". It isn't publicly available yet — "coming soon" — but OpenAI says it scored highly on cybersecurity-related evaluations. That's a statement about capability and about how access gets restricted, so I'll write about it once it's out and I've touched it.

Honestly, how did it go, having an AI summarise the news?

Both good parts and parts I couldn't use as-is.

The good part is the speed of collection. Release notes and reactions on X scattered across places, put into one list quickly, is genuinely helpful. Much faster than opening every source myself.

The part I couldn't use was the temperature of the headlines. "X is insane", "it's become the strongest": phrasing that would get clicks but oversells what's underneath. Even when the numbers come from the sources, the certainty of the phrasing can't be taken at face value, so a human check was needed there.

So is this way of working worth it?

As a division of labour — the AI handles finding topics and gathering primary sources, I decide how to write it and how far to commit — yes, definitely. In this case, using the Fable 5.1 price cut and the benchmark numbers as they came, and only cutting the title and the hype myself, was about the right distance.

Leave everything to AI and you get a post where the facts are right but the atmosphere is overcooked. Research it all myself and this post doesn't go out today. Somewhere in between is where it sits, for now.