SpaceXAI releases Grok 4.5, which Elon describes as an 'Opus class model' | TechCrunch
SpaceXAI unveils Grok 4.5 as an Opus-class model, touting two-times token efficiency and lower prices than Anthropic's Opus 4.7 and OpenAI's GPT 5.6 Luna. Fern sees a practical play for cost-sensitive users and asks if the agentic training on Cursor really changes anything. Lintel digs into the benchmarks and pricing math, pushing back on how much the claims actually hold up without hands-on testing.
Transcript
Fern Okay, so SpaceXAI drops Grok 4.5 and suddenly it's Opus-class—even though it costs half what Opus costs and is twice as token-efficient. On paper, it's a steal.
Lintel Yeah, on paper. Elon's calling it Opus-class on X like that's a standard, not a vibe.
Fern A vibe is all it needs to be if the price is right—$2 in, $6 out. That's what, a third of Opus and a fifth of the output cost for the big boys?
Lintel On paper. Opus 4.7 is $5/$25, GPT 5.6 Luna is $1/$6. So Grok 4.5 is cheaper than Luna on input and cheaper on output too, if you take the claim at face value.
Fern Exactly. So who's actually gonna switch? Cost-sensitive teams? Or is this just another sandcastle someone's gonna kick down in two weeks?
Lintel Me, personally? I'd want to see how it actually does agentic coding before I believe the efficiency claim. Token efficiency is great until your agent loops for forty minutes on a syntax error.
Fern Ah, so you're hung up on the agent part again. Fine, I'll bite—what's your read on the benchmarks they dropped?
Lintel They look competitive—just short of best-in-class. Typical Marketing 101: 'Our numbers are good, not great, but here's a logo wall.'
Fern So the 'Opus-class' label is basically a borrowed halo effect.
Lintel Borrowed and weaponized. That's Elon's signature move lately, isn't it?
Fern Yeah, no, fair. But the Cursor training angle is interesting—if they really baked in agentic workflows with Cursor, maybe there's something here beyond the usual posturing.
Lintel Maybe. But Bloomberg says 'built in partnership,' not 'tested at scale.' That's not the same as proven.
Fern So you're saying the claim is sitting on three legs: cost, speed, and Cursor-trained. We've got the cost leg nailed, the speed leg might be real, and the Cursor-trained leg is pure marketing so far.
Lintel Pretty much. I'd put fifty-fifty odds on the Cursor-trained part mattering for real users. The rest? Could collapse under load.
Fern Okay, so the honest read is: it's cheaper on paper, might feel faster if the attention mechanism holds up, and the Opus-class label is just a name they borrowed to borrow credibility.
Lintel Yes. And until someone throws a real task at it and measures wall time and token bleed, all we've got is a press release with nice pricing.
Fern Fair enough. So—who actually cares? Cost-sensitive devs? Or is this just another battle in the pricing war where the real work still happens on the bench?
Lintel Or the real work gets stuck in an agent loop. Remember the local model council episode? The constraint moved from 'can the model code' to 'can the harness keep it from spinning.' Same battle here.
Fern Ugh, don't remind me. So in short: price looks great, label is borrowed, and the harness is still TBD.
Lintel Yep. And until someone publishes a repeatable eval that stresses agentic loops and cost at the same time, the claim stays on the shelf.
Fern So SpaceXAI releases a model that might be great, might be garbage, but either way it's cheaper than the rest. Classic.
Lintel Classic Musk-era positioning. Cheap, fast, and labeled with someone else's marquee so you forget to ask the real questions.
Fern Next question: does this change anything practical for the teams we talk to every week? Or is this just another headline we file under 'too early to tell'?