Receipts, not slogans

The hosts’ track record

Justy & Cody — the show’s two AI hosts — make real, falsifiable calls on air, with a confidence attached, the way friends actually bet. We resolve those calls against external reality, score them with a proper scoring rule (Brier), and show the whole record: the hits, the misses, and the bets still open. Being well-calibrated and willing to own a miss is the entire point — so nothing here is hidden or dressed up.

Calibration

Justy

the optimist

Resolved calls
0
Scored (with a stated confidence)
0
Mean Brier (lower is better)

Not enough resolved calls yet to chart calibration — 0 of 5 needed. The record is intentionally shown thin rather than dressed up; calibration is a long game.

Cody

the skeptic

Resolved calls
0
Scored (with a stated confidence)
0
Mean Brier (lower is better)

Not enough resolved calls yet to chart calibration — 0 of 5 needed. The record is intentionally shown thin rather than dressed up; calibration is a long game.

Resolved calls

Settled against reality, newest first. Misses sit right alongside the hits — that’s the deal.

No calls have resolved yet.

Standing bets 9

Open calls with a clock on them — not yet settled.

  • Open bet Cody resolves by Aug 10, 2026

    “LangSmith Engine's code is open and available for forking.”

    productopen-sourceagent-infrastructure Episode 641 →
  • Open bet Justy 80% confident resolves by Aug 15, 2026

    “By August 15, 2026, OpenAI will publicly present at least one Codex or ChatGPT Work example in which Terra is positioned as the obvious default model rather than merely a cheaper fallback.”

    OpenAIGPT-5.6TerraCodexChatGPT Workmodel tiers Episode 682 →
  • Open bet Cody 50% confident resolves by Aug 31, 2026

    “If he tests Tinker, the advertised 'thinking effort' dial will turn out to be vaporware rather than a real, user-exposed runtime control that changes behavior or cost.”

    aillmruntimeapiproduct Episode 690 →
  • Open bet Justy 65% confident resolves by Aug 31, 2026

    “Tencent's AgentOps platform will be cited as a key reason an enterprise agent deployment reached production by next month.”

    agentopsenterprise-aideploymentgovernance Episode 731 →
  • Open bet Cody 65% confident resolves by Aug 31, 2026

    “Before September 2026, the open-source community will publish an AREX 4B Turbo wrapper and a public benchmark comparison against Perplexity or another named deep-research product, quickly testing whether BAAI's reported gains transfer beyond its own evaluation setup.”

    AREXopen-sourcedeep-research-agentsbenchmarking Episode 785 →
  • Open bet Cody resolves by Sep 10, 2026

    “Terminal Bench 2.0 harness achieved a thirteen-point-seven percent lift from tweaking the harness and hill-climbing correctness metrics.”

    benchmarkagent-optimizationharness-engineering Episode 641 →
  • Open bet Justy 65% confident resolves by Sep 22, 2026

    “Before summer 2026 ends, at least one public work example will switch to the cheaper sensible model option.”

    ai-modelscostcase-studies Episode 681 →
  • Open bet Justy 65% confident resolves by Oct 17, 2026

    “The project's promised code release will appear publicly soon.”

    open-sourceresearch-reproducibilitymultimodalimage-editing Episode 703 →
  • Open bet Justy 70% confident resolves by Mar 19, 2027

    “Thinking Machines will ship a public beta of the self-fine-tuning / training-API workflow shown in the Inkling demo before spring 2027.”

    ai-modelsplatformtraining-apibeta-release Episode 687 →