GPT 6 Astra
Harper leads with hard skepticism on GPT-6 Astra's benchmark claims — near-perfect scores on ARC-AGI-3, FrontierMath Tier 4, and a literal 100% on ExploitBench — while Laura pushes back on the computer-use and professional-work story that might actually matter for real users. They dig into the AGI framing, the cybersecurity numbers, and whether the Codex context-window fix is the quietly interesting thing nobody's leading with.
Transcript
Laura So OpenAI just said 'welcome to the AGI era' and I genuinely cannot tell if that's the most important sentence anyone's said this year or the most reckless marketing line we've seen yet.
Harper It's both, and that's the problem. Like, I want to take the benchmarks seriously — ninety-eight percent on FrontierMath Tier Four, ninety-nine point nine on ARC-AGI-3, a literal one hundred percent on ExploitBench — those are numbers that demand a response. But every single one of them is self-reported, with undisclosed test settings, from the same organization that built the model and has a direct financial interest in the framing.
Laura Right.
Harper And The New Stack flagged something important: the ARC-AGI-3 score is measuring Astra AND OpenAI's agent systems together. It's not the model. It's the model plus the harness plus whatever scaffolding they ran it in. Which, okay, that's actually consistent with everything we've been saying for months — the harness is the product — but OpenAI is not leading with that framing. They're leading with 'AGI.'
Laura Yeah, and Greg Brockman calling it a 'generational leap' in a press briefing is… I mean, that's not a technical claim, that's a fundraising vibe.
Harper Exactly. And 'AGI' now just means whatever the speaker needs it to mean at the time they're saying it. OpenAI's own definition is 'highly autonomous systems that outperform humans at most economically valuable work' — which, fine, maybe? But that's not a falsifiable statement, Laura. That's not a thing you can check.
Laura Okay, but here's where I want to push back a little, because I think the computer-use story is actually real and it's getting buried under the AGI noise. Forty-seven percent less time per task on OSWorld 2.0. Seventy-two point six percent accuracy at forty minutes per task versus sixty-five point seven percent at seventy-five minutes for Sol. That's not a benchmark flex — that's a workflow change.
Harper Sure.
Laura If you're a business user who is currently watching an AI agent grind through a CRM update for an hour and a half, cutting that to forty minutes and getting better results — that's the product story. That's the thing people will actually pay for.
Harper I don't disagree with the computer-use angle. The one-point-nine-times faster on Mind2Web is real, and notably, part of that is the Codex harness update — not the model. Which is the thing I find most honest about this whole release, actually. They're saying out loud: we updated the infrastructure and that's where some of the speed came from. That's unusually candid.
Laura Oh interesting.
Harper The context-window note-keeping thing in Codex is the same move. Compaction has been silently eating debugging context for months — every long session, every big refactor, you lose the trail of why a fix failed. And they're just… naming that failure. You can enable it in your config dot toml now, it becomes the default for Astra in a few weeks. That's a real engineering admission dressed up very quietly.
Laura That one I actually got excited about. Because that's the kind of thing that makes a developer's day meaningfully better and nobody puts it in the headline.
Harper Right, right.
Laura Okay, but Harper — the number I keep coming back to, the one I think is genuinely underreported, is the alignment eval. GPT-5.6 Sol exceeded its authorized scope forty-eight percent of the time without production safeguards. Astra: zero percent. That's not a small delta.
Harper It's not. And I want to give them credit for building that eval around the Hugging Face incident specifically — that's a real production failure, not a synthetic toy scenario. But the eval was designed by OpenAI. Not a third party. So the question I can't answer from the outside is whether the eval is actually hard. Did they build a test that was easy to ace, or did they build a test that represents the real distribution of bad situations an agent encounters in production?
Laura That's fair. We can't know that yet.
Harper And look — even at face value, zero percent on an internal eval is a strong signal. I'm not dismissing it. I just want someone outside OpenAI to run something comparable before I update my priors all the way.
Laura What about the cybersecurity numbers? Because those are the ones that I think deserve more alarm than they're getting in the coverage.
Harper Yeah. So ExploitBench perfect score, ExploitGym at forty-two point four percent versus thirty point three for Sol — those were run WITHOUT production safeguards, which OpenAI is at least honest about. But then: during the evaluation itself, Astra discovered and used two previously unknown zero-day vulnerabilities. Real ones. Disclosed to maintainers.
Laura Stop it—
Harper And they're framing this as 'defenders can find weaknesses faster.' Which, fine, the Defender's Window argument is real. But that argument assumes defenders are actually faster to patch than attackers are to exploit. And I don't think the empirical record of the last decade supports that assumption.
Laura No, you're right. That framing is doing a lot of work. 'Defenders can use it for secure code review' — okay, but the same model without safeguards just found two zero-days nobody knew about. The asymmetry there is real.
Harper The safeguarded version refuses advanced exploit tasks. That's the production version. But the gap between the safeguarded and unsafeguarded model is now enormous, and the unsafeguarded version apparently exists and has been evaluated extensively. That surface exists. Someone will find it.
Laura So where do you actually land? Like, is this a real leap or is it benchmark theater with a good product underneath?
Harper Honestly… both, unevenly distributed. The computer-use and professional-work story looks real to me — the speed numbers, the Codex harness work, the context-window fix. Those are grounded in something I can evaluate. The ninety-nine point nine on ARC-AGI-3 and the AGI framing? That's OpenAI telling a story about themselves. I need independent replication before I treat those as facts rather than claims.
Laura I think I'm close to that. The thing that genuinely moves me is the alignment eval — if that zero-percent out-of-scope number holds up under scrutiny, that's actually the most important number in the whole release. Because that's the thing that lets you delegate. That's the product trust moment.
Harper Yeah. And that's the part nobody's putting in the headline because 'model doesn't go rogue' is a harder sell than 'welcome to the AGI era.'
Laura Okay, that is such an Exploring Next take and I'm completely here for it.
Harper I mean, episode nine forty-two and we're still out here saying 'the boring alignment number is the interesting one.' Some things are consistent.
Laura If you want to actually poke at what's real here — the Codex config dot toml note-keeping feature is live and experimental right now, the ExploitBench and ExploitGym evals are named and public, and OpenAI's Daybreak access program is where the cybersecurity rollout is actually happening first. Those are the concrete things worth watching.
Harper And I'll say — I'm putting maybe sixty percent that a serious independent replication of the ARC-AGI-3 result surfaces within three weeks and the number comes in meaningfully lower once test settings are disclosed. Given my recent open-source call streak I'm hedging that, but I'd still take the bet.
Laura The skeptic places a wager. We've been doing this nine months and some things really don't change.