AA Briefcase: Agentic Knowledge Work Benchmark | Artificial Analysis
Justy and Cody dig into AA-Briefcase, Artificial Analysis's new agentic benchmark that tests models on real knowledge-work deliverables — spreadsheets, presentations, memos — across four multi-week scenarios. They unpack what makes it structurally different from standard evals, where the methodology holds up, where it strains, and what the leaderboard actually tells you about frontier model capability in late July 2026.
Transcript
Justy Okay, episode eight hundred. Which is a number that should mean something, and yet here we are, talking about a benchmark.
Cody As is tradition. What's the benchmark?
Justy AA-Briefcase, from Artificial Analysis. And the thing that got me is the framing — it's not grading models on whether they predicted the right token. It's asking them to actually PRODUCE a deliverable. A spreadsheet, a slide deck, a memo. Real work product.
Cody Okay, so walk me through the grading, because that's where I want to poke. Three layers, right? Binary rubric checks, analytical quality pairwise, and presentation pairwise.
Cody The binary rubric is the part I actually respect most. Pass or fail — did the model find the requirement that was buried across source files, did it cite the right evidence, did it resolve a planted conflict between sources? That's genuinely hard to fake. But the pairwise stuff — analytical quality, presentation — that's Elo, which means your score depends on who else is in the pool. If the pool shifts, the rankings shift. It's relative, not absolute.
Justy That's fair. Though I'd push back a little — like, what's the alternative for grading whether a memo is analytically rigorous? Multiple choice? At some point you need a human or a model making a comparative call.
Cody No, you're right. I'm not saying it's wrong, I'm saying it's relative. Know what you're reading.
Justy Okay, so the leaderboard. Cody, Claude Fable 5 is leading on rubric pass rate and analytical quality. Opus 5 is right there too. GPT-5.6 Sol is third.
Cody And the cost-vs-performance scatter is the actually interesting chart. Opus 4.8 is competitive but it's one of the slowest — something like twenty-three minutes per task on average. GLM-5.2 is the one that surprised me — top open-weights performer, beats GPT-5.5, though also slower. And that's actually the thing this benchmark surfaces that most don't — time and cost per task as first-class metrics.
Justy Right. Okay but here's the thing that's nagging at me — the tasks run independently. Each week, fresh context. The agent doesn't carry over its OWN prior submissions.
Cody Yeah, that's the structural gap. Real knowledge work is cumulative — week two builds on what you actually wrote in week one. Without that carryover, you're measuring something slightly different from a real multi-week project.
Justy It's like… the files from the scenario carry over, but the agent's own memory of what it decided doesn't. Which is honestly just where the field is right now — nobody's cleanly solved durable agent state across sessions.
Justy Okay, so who should actually care about this? My read is: teams evaluating which model to put on knowledge-work automation — the kind where the output is a real document someone reviews — this is the most honest signal available right now. Because it's grading on deliverable quality, not vibes.
Cody I'd add: the tool-usage breakdown is genuinely useful for practitioners. They track explore, read, write, compute, view image, other — per task. So you can see whether a model is over-exploring, under-reading, or burning turns on things that don't move the needle. That's diagnostic in a way that a single accuracy number isn't.
Justy And they put a public version on Hugging Face — the Due Diligence scenario, one week, illustrative only. Doesn't count toward official Elo, but you can actually see a task brief, a model submission, and the grading. That's… rare for a benchmark at this level.
Cody That's the thing that earns some credibility with me. The methodology is inspectable. I still want to see how stable the Elo is as the model pool grows — right now it's twenty-three of fifty-six models — but the structure is at least honest about what it's measuring.
Justy Honestly I think this is round one of a benchmark category, not a settled eval. Ask me in six months whether the no-carryover limitation gets addressed.
Cody Yeah… ask me in six months whether the pairwise pool is stable enough that the Elo means anything cross-snapshot. Those are the two things I'd watch.
Justy Eight hundred episodes of 'ask me in six months.' That's basically the show at this point, Cody.
Cody We should put that on a mug.
Justy We really should. Alright — repo's called AA-Briefcase-Lite on the Artificial Analysis GitHub, and the public scenario is on Hugging Face if you want to poke at actual grading outputs.
Cody Worth a look. Especially the rubric checks — seeing what counts as a planted conflict is more illuminating than the leaderboard number.
Justy Good call. Okay, eight hundred down. Somehow still going.