A Scorecard for the AI Age
OpenAI’s scorecard argues AI value must be measured in useful work per dollar, not just token cost. Cooper sees a practical product story; Miles pokes at the metrics and pushes for mechanistic honesty. The two hash out whether the framework holds up and what it changes day-to-day.
Transcript
Cooper —okay, but what if the thing we’re shipping is garbage and we just called it ‘work’ to justify the bill? That headline metric—useful work per dollar—sounds like the kind of thing CFOs eat up. And honestly, we’ve been saying the same thing since November. I’m into it.
Miles Wait— what does "useful work" even mean when the work is spread across spreadsheets, emails, and the person who actually hit "save" at 3 a.m.?
Cooper It means ‘done’ inside the workflow—customer issue closed, code change landed in main, contract reviewed on time. The post walks through a finance forecast example: ChatGPT Work moves data, reconciles tabs, keeps slides honest. One workflow, one definition of done. Then count what actually cleared quality and ask what it cost. That’s the scorecard.
Miles Sure, sure—tell a CFO their team’s midnight saves are now ‘useful work’ and watch the smiles disappear. Have you tried measuring ‘done’ across a real org? ‘Quality meets the bar’ is squishy. A human still eyeballs half the outputs. Where’s the hard mechanism behind the metric?
Cooper The mechanism is the model tier and the product scaffolding that lets you push the context it needs without copying files at gunpoint. GPT-5.6 drops last week with Sol, Terra, Luna—Terra’s the balanced mid-tier, Luna’s the fast cheap one. The post even cites DeepSWE v1.1: Sol set a new high at 72.7% at 36.2% lower estimated API cost. Fewer output tokens, one pass, fewer retries. That’s the proof.
Miles Again: 36.2% lower estimated API cost? ‘Estimated.’ And Sol’s the only tier that actually hit 72.7%? What happened to ‘dependable’? This read feels sanded smooth for C-suites; the real cost is the human review budget no one’s budgeting.
Cooper Uh—we’re in July 2026, Miles. ChatGPT Work shipped this week and it’s got the governance layer enterprises actually care about. Security, privacy, workspace management—so orgs can give AI access to real workflows without the IT bloodbath. That’s the boring part that unlocks the ‘done’ part.
Miles I don’t disagree the security story matters; I’m saying the post—and the scorecard—still hand-wave the operational side. ‘Needs correction’ is the polite euphemism for ‘half the contract clauses are still wrong.’ Where’s the honest taxonomy of failure modes?
Cooper It’s in the four questions: dependability tracks ready-to-use vs needs correction vs needs escalation. That’s explicit. The post even tells teams to define boundaries—what data AI can touch, which systems it can change, when a human must approve. That’s the floor, not the ceiling.
Miles Fine, floor is set. But the ceiling is the compute flywheel—training pushes frontier, inference pushes product, better product drives adoption, revenue funds next-gen research. Neat story, but only if every loop closes. OpenAI’s own infrastructure stack has to pull its weight across products, APIs, and Enterprise. One weak link and the flywheel stalls.
Cooper Which is why the tiered family isn’t just marketing—it’s the lever. Luna handles high-volume fast flows; Terra steps up when depth matters; Sol steps in when you need reasoning to land first try. The post’s numbers are still early, but the mechanism is there: purpose-built hardware, smarter routing, efficient inference. If the flywheel turns, each dollar produces more value at scale.
Miles I’ll grant you the framework is directionally right—measure work, measure cost per outcome, track dependability, watch value grow. But the moment you hand a CFO an Excel full of these inputs, every cell becomes a negotiation and every assumption becomes a fight. The post sells the dream; reality sells the spreadsheet.
Cooper —okay, fair. The dream is you hand them a live dashboard that traps real ‘done’ metrics inside the workflow, so the data stops being fiction. That’s still a product problem, but at least it’s a solvable one.
Miles Solvable until the workflow mutates tomorrow and the dashboard still only speaks last week’s definition of done.
Cooper Ha! Okay—true. Which means the scorecard has to keep tightening its own screws. But I’ll take solvable over unsolvable every time.