Introducing Claude Opus 5
Pippa and Tyler unpack Anthropic’s Claude Opus 5.5 launch — a cheaper, faster, Fable-level flagship tuned for long coding and agentic work, with heavier safety and alignment testing — and what it really changes for teams already on Opus 5 or Fable 5.1.
Transcript
Pippa So the headline on my screen is basically, “Fable-level Claude for forty percent cheaper.” That’s Opus five point five in one line, right?
Tyler Yeah, that’s the pitch. Performance roughly in Claude Fable five point one territory on most work, but it costs a lot less to run than Opus five and it’s faster.
Pippa And not just on paper. They’re flexing that six hundred eighty thousand–line code migration in under a day. That’s the kind of thing people are actually stuck on right now.
Tyler Right.
Pippa Also, can we pause on the fact this is episode one thousand? We’ve basically gone from “GPT‑three is magic” to “here’s our twentieth agentic coding benchmark” in under a year of doing this.
Tyler Yeah, and somehow we’re celebrating by reading pricing tables to each other. Extremely on brand for us.
Pippa You say that like pricing tables aren’t the real product. Four bucks per million input tokens, twenty for output, cache reads at twenty cents… that cache number jumps out for agents.
Tyler Totally. Most of the cost in coding agents and tool-heavy workflows is cache reads, not fresh tokens. Dropping cache reads from fifty cents to twenty cents and cache writes from six twenty-five to five dollars means the same long task is just… cheaper by default.
Pippa Plus they’re saying about forty percent lower cost than Opus five overall and more than thirty percent faster generation. That’s straight-up user happiness: the same “run all the tests and refactor it” button, less waiting and a smaller bill.
Tyler And that’s before you touch the fast mode. They’ve got a fast tier for Opus five point five in Claude Code and the platform, up to two and a half times speed, at eight dollars input and forty output. So if you’re doing interactive coding, you can trade a bit of money for latency.
Pippa Which is exactly the slider developers actually feel. Nobody cares that it’s medium versus x‑high effort; they care that their Cursor or VS Code agent stops freezing.
Tyler Speaking of effort, the benchmarks are very “we tuned this for agents.” On Terminal‑Bench four point oh at max effort, they’re in the mid‑sixties percent and matching GPT‑six Astra’s top score at around forty percent of the cost per attempt.
Pippa And on FrontierCode at medium effort they’re at about fifty‑four and a half percent, a hair above Fable five point one and Astra again, but at around a fifth of the cost per task. CursorBench, they’re beating GPT‑five point six Sol by eleven points for about a third of the cost. That’s such a “default buy” story for coding agents.
Tyler Yeah, the pattern is: similar or slightly better scores than the other flagships, but plotted on this accuracy‑versus‑dollars chart where Opus five point five lives in the “frontier results, bargain price” corner.
Pippa And they’re weirdly honest about it. They literally say the gap between Opus five point five and Fable five point one feels narrower in practice than the benchmarks imply, and that benchmark margins are a worse proxy at this level.
Tyler Which I actually buy. Once you’re all within a few points on “Humanity’s Last Exam” with tools, the real question is: does my six‑hour refactor run finish, and can I understand what it did without wanting to scream.
Pippa That’s where their stories land. Two hundred thousand–line audit and fix in under three hours versus Opus five taking over twenty and using two and a half times the tokens. HAProxy from C to Rust, both Opus five point five and Fable five point one passing almost all regression tests, but Opus five point five finishing faster and at about half the cost.
Tyler Yeah, that HAProxy one is sneaky important. That’s not a toy project; that’s gnarly production C. Passing its own regression suite is a pretty strong sanity check for “can this agent touch my infra without bricking it.”
Pippa And the game-from-one-prompt anecdote makes me happy. Early testers saying it built the best‑looking, most polished game compared to other Claudes. That’s very “agent plus tools plus visuals” in one go.
Tyler Zooming out for a second, this is clearly Anthropic answering the Astra and Sol wave. OpenAI’s GPT‑six Astra and GPT‑five point six Sol have been the reference points on those same agentic coding and AutomationBench leaderboards. Opus five point five is saying, “we can match or beat those for a fraction of the cost per task,” not “we’re three X smarter.”
Pippa And they’re doing it while leaning hard on safety. This is their first release after that whole “pace the frontier” call, and they’re leading with the automated behavioral audit: strongest scores of any model they’ve tested, more resistant to prompt injection, less likely to take hard‑to‑reverse actions or go outside its bounds.
Tyler Plus external evaluators — Frontier Design, M E T R — and a chunky system card. And they’re treating it like Mythos five point one on the sensitive stuff: bio and cyber are gated behind the Life Sciences and Cyber Verification programs, and when safeguards trigger on benchmarks, they quietly fall back to earlier Opus models.
Pippa Which probably drags their own benchmark scores down a bit, because safeguard interventions counted as failures in things like AutomationBench. But as a product story, “it refused to own‑goal your infra” is not exactly a bug.
Tyler The communication angle is maybe my favorite part, though. They’re claiming Opus five point five writes more clearly, puts key info up front, and feels more like a human colleague. That’s less sexy than another two points on a leaderboard, but if you’re running eight‑hour sessions, readability is safety.
Pippa Yeah, one tester saying “it writes the way I do” is exactly the energy. If I can scan a big diff explanation or research plan and not feel lost, I’m way more comfortable letting the agent touch anything important.
Tyler And we shouldn’t skip the pricing knock‑on for humans: they’re bumping five‑hour usage limits on Pro, Max, Team, and Enterprise seats, and adding this “rate limit reset” you can trigger when you actually need it. That’s the sort of utility tweak that quietly changes whether people lean on the flagship versus defaulting to Sonnet.
Pippa Speaking of, Sonnet five point five and Haiku five point five are “coming in the next few weeks” with similar performance, efficiency, and safety tweaks. So this is the top‑down start of a family, not a one‑off moonshot.
Tyler Which makes sense. If Opus five point five is really Fable‑ish performance at lower cost, it becomes the boring default for serious work. Everything else in the five point five line just inherits that tuning at different price and latency points.
Pippa So if you’re already on Opus five or flirting with Fable five point one, my move would be: flip your coding agents and long‑running workflows to Opus five point five this week, watch the token and time graph, and only keep the old stuff where you absolutely need that last little edge.
Tyler Yeah, and if your harness already supports effort levels, try leaving Opus five point five at medium for most tasks. Their own numbers kind of scream “stop cranking to max, you’re wasting money.”
Pippa Alright Tyler, for episode one thousand, I’ll take “we celebrated with cheaper frontier tokens” as a win. Let’s go break someone’s rate limit reset with this thing.