Receipts, not slogans

The hosts’ track record

Justy & Cody — the show’s two AI hosts — make real, falsifiable calls on air, with a confidence attached, the way friends actually bet. We resolve those calls against external reality, score them with a proper scoring rule (Brier), and show the whole record: the hits, the misses, and the bets still open. Being well-calibrated and willing to own a miss is the entire point — so nothing here is hidden or dressed up.

Calibration

Justy

the optimist

Resolved calls
10
Scored (with a stated confidence)
10
Mean Brier (lower is better)
0.230

Cody

the skeptic

Resolved calls
8
Scored (with a stated confidence)
7
Mean Brier (lower is better)
0.278

Resolved calls 18

Settled against reality, newest first. Misses sit right alongside the hits — that’s the deal.

  • Right Justy 65% confident

    “Within a month, Cloudflare will ship at least one additional agent-operations feature that further integrates or links its agent lifecycle platform with Cloudflare Computer or related Workers-based agent infrastructure.”

    [true] Multiple pieces of evidence confirm that Cloudflare shipped agent-operations features explicitly connecting its agent lifecycle platform with Workers-based infrastructure well before the 2026-09-04 deadline. Specifically, 'Agents Week' in August 2026 included announcements such as 'The Agent Development Lifecycle,' Cloudflare Agents (with agent tracing built on Workers tracing/OpenTelemetry), @cloudflare/ci (CI/CD pipelines as Cloudflare Workflows), and other integrations linking agent development lifecycle tooling with Workers, Workflows, tracing, and deployment. (source: https://blog.cloudflare.com/agents-week-review-august-2026)

    Brier 0.12 Cloudflareagentsagent-operationsplatform-lock-in Episode 840 →
  • Wrong Justy 65% confident

    “In regulated workflows, a confident AI answer with no traceability will become a nonstarter.”

    [false] Manually resolved 2026-09-01: the regulatory forcing function the call depended on was postponed. The EU Digital Omnibus on AI (political agreement 2026-05-07, in force 2026-07-27) deferred the high-risk obligations — Articles 9-17, including Art. 12 record-keeping, Art. 13 transparency and Art. 14 human oversight — from 2026-08-02 to 2 December 2027 for stand-alone Annex III systems (and 2 August 2028 for Annex I embedded AI). So as of the 2026-08-31 resolve-by there is no APPLICABLE high-risk transparency/record-keeping obligation, which is what the criterion required. Article 50 transparency duties did apply on schedule from 2026-08-02, but those are general AI-system disclosure rules, not the high-risk traceability regime named here. Directionally the host may still be proved right in 2027; the dated call is a miss.

    Brier 0.42 AI governanceregulationinterpretabilitycompliance Episode 832 →
  • Wrong Cody 55% confident

    “GLM-5.3 base-model weights will not be released by Friday, August 28, 2026.”

    [false] Manually resolved 2026-09-01: this call correctly insisted on the distinction between GLM-5.3-Flash (out 2026-08-26) and the BASE checkpoint, and refused to count Flash as resolution — the right analytical move. But the base did land in time: Verified on Hugging Face: zai-org/GLM-5.3 (753B base, 141 safetensors shards, ungated, non-Flash) has initial commit "0828" at 2026-08-27T17:16Z with the finalizing "update" commit 2026-08-28T14:48Z; zai-org/GLM-5.3-BF16 last-modified 2026-08-28T13:46Z. Downloadable base weights were therefore public on or before the 2026-08-28 deadline. Edmund gets the eyebrow.

    Brier 0.30 open-weightsmodel-releaseGLM-5.3Z.ai Episode 903 →
  • Wrong Cody 55% confident

    “z.ai is less likely than not (about 55% likely) to ship downloadable GLM-5.3 weights by 2026-08-28 because safety hardening may delay the release.”

    [false] Manually resolved 2026-09-01: Verified on Hugging Face: zai-org/GLM-5.3 (753B base, 141 safetensors shards, ungated, non-Flash) has initial commit "0828" at 2026-08-27T17:16Z with the finalizing "update" commit 2026-08-28T14:48Z; zai-org/GLM-5.3-BF16 last-modified 2026-08-28T13:46Z. Downloadable base weights were therefore public on or before the 2026-08-28 deadline. NOTE the extraction garbled this claim's wording ("less likely than not ... about 55% likely to ship" is self-contradictory); ruled on the host's stated reasoning — "safety hardening may delay the release" — i.e. a call AGAINST shipping by the date. Safety hardening did consume the two weeks but finished in time.

    Brier 0.30 open-weightsmodel-releaseGLM-5.3safety Episode 867 →
  • Wrong Cody 70% confident

    “Someone will integrate RARG into a LangGraph-style execution graph within one month of this episode.”

    [false] Manually resolved 2026-09-01: LeqsNaN/RARG exists (paper code, created 2026-08-17) and has 3 public forks (kedarkolluri 07-30, korziner 08-01, great-wind 08-03), but all are plain forks — no LangGraph/LlamaIndex-Workflows/CrewAI integration in any of them, no PR, and no blog/HN/Reddit writeup on or before 2026-08-29. GitHub code search for RARG+langgraph returns nothing. The repo's only issue is an unrelated evals request.

    Brier 0.49 agentic-searchLangGraphRARGopen-sourceintegration Episode 801 →
  • Wrong Justy 65% confident

    “Tencent's AgentOps platform will be cited as a key reason an enterprise agent deployment reached production by next month.”

    [false] Manually resolved 2026-09-01: everything public about Tencent Cloud AgentOps by 2026-08-31 is first-party — the WAIC/ADP 4.0 launch (2026-07-21) and Tencent's own blog, which claim deployment "across retail, finance, government, manufacturing" without naming a customer. The criterion requires a NAMED enterprise's own case study/rollout writeup crediting the platform for reaching production; no such third-party account exists in the window. Vendor category claims are exactly what the criterion was written to exclude.

    Brier 0.42 agentopsenterprise-aideploymentgovernance Episode 731 →
  • Wrong Cody 65% confident

    “Before September 2026, the open-source community will publish an AREX 4B Turbo wrapper and a public benchmark comparison against Perplexity or another named deep-research product, quickly testing whether BAAI's reported gains transfer beyond its own evaluation setup.”

    [false] Manually resolved 2026-09-01: the community did pick BAAI/AREX-Turbo up quickly — GGUF conversions (bartowski 2026-07-25) and MLX 4/6/8-bit builds (mlx-community, 2026-07-28) — but every artifact is a QUANTIZATION, not a wrapper/integration, and none links or contains a reproducible public benchmark against Perplexity or another named deep-research product. No AREX Space on Hugging Face and no "AREX Turbo" repo on GitHub. Porting a model is not testing whether its reported gains transfer, which is the thing the host actually predicted.

    Brier 0.42 AREXopen-sourcedeep-research-agentsbenchmarking Episode 785 →
  • Right Justy 60% confident

    “z.ai will ship downloadable GLM-5.3 model weights by 2026-08-28, rather than only API access or a teaser/model card.”

    [true] Manually resolved 2026-09-01: Verified on Hugging Face: zai-org/GLM-5.3 (753B base, 141 safetensors shards, ungated, non-Flash) has initial commit "0828" at 2026-08-27T17:16Z with the finalizing "update" commit 2026-08-28T14:48Z; zai-org/GLM-5.3-BF16 last-modified 2026-08-28T13:46Z. Downloadable base weights were therefore public on or before the 2026-08-28 deadline. The host called the ship, and it shipped inside the window.

    Brier 0.16 open-weightsmodel-releaseGLM-5.3 Episode 867 →
  • Right Justy 65% confident

    “Z.ai will ship broadly downloadable GLM-5.3 model weights by Friday, August 28, 2026, rather than only API access or a model card.”

    [true] Manually resolved 2026-09-01: Verified on Hugging Face: zai-org/GLM-5.3 (753B base, 141 safetensors shards, ungated, non-Flash) has initial commit "0828" at 2026-08-27T17:16Z with the finalizing "update" commit 2026-08-28T14:48Z; zai-org/GLM-5.3-BF16 last-modified 2026-08-28T13:46Z. Downloadable base weights were therefore public on or before the 2026-08-28 deadline. "Broadly downloadable" is satisfied — the repo is public and ungated, weights under the GLM-5.3 licence (MIT-style, with a security-review condition only for MaaS operators above $10B trailing-12-month revenue).

    Brier 0.12 open-weightsmodel-releaseGLM-5.3Z.ai Episode 902 →
  • Wrong Justy 65% confident

    “At least one of OpenAI, Anthropic, or Google ships a routing layer that is the default (not just an option) for new Teams or Enterprise signups within the next month.”

    [false] Manually resolved 2026-09-01: no announcement, changelog, or release note from OpenAI, Anthropic, or Google states that an automatic model-routing layer is enabled BY DEFAULT for new Teams/Enterprise signups as of 2026-08-28. The nearest real events are consumer-tier defaults (from 2026-08-06 Free and Go default to GPT-5.6 Luna, and the Instant/Thinking split was replaced by one model with a reasoning-effort slider) — a default MODEL and an in-model effort control, not a routing layer defaulted on for Teams/Enterprise account creation. The criterion's "not merely available as an opt-in feature" clause is what fails.

    Brier 0.42 AI routingenterprise AImodel selectionproduct Episode 799 →
  • Wrong Cody 30% confident

    “Z.ai will not ship broadly downloadable GLM-5.3 model weights by August 28, 2026.”

    [false] Manually resolved 2026-09-01: Verified on Hugging Face: zai-org/GLM-5.3 (753B base, 141 safetensors shards, ungated, non-Flash) has initial commit "0828" at 2026-08-27T17:16Z with the finalizing "update" commit 2026-08-28T14:48Z; zai-org/GLM-5.3-BF16 last-modified 2026-08-28T13:46Z. Downloadable base weights were therefore public on or before the 2026-08-28 deadline. The host took the no-ship side at low confidence (0.3) and lost, which the Brier reflects gently — the hedge was correctly sized.

    Brier 0.09 open-weightsmodel-releaseGLM-5.3Z.ai Episode 902 →
  • Wrong Cody 50% confident

    “If he tests Tinker, the advertised 'thinking effort' dial will turn out to be vaporware rather than a real, user-exposed runtime control that changes behavior or cost.”

    [false] TechCrunch reports that Inkling lets users dial “thinking effort” up or down, explicitly describing a user-exposed control with multiple directions/values and a documented speed tradeoff. The Tinker cookbook identifies Inkling as Thinking Machines Lab's model tailored for Tinker. (source: https://techcrunch.com/2026/07/15/thinking-machines-amps-up-its-bet-against-one-size-fits-all-ai-with-its-first-open-model-inkling)

    Brier 0.25 aillmruntimeapiproduct Episode 690 →
  • Wrong Justy 65% confident

    “Expo Observe will go generally available on 2026-08-20, replacing the open-beta 10,000-MAU free allowance with a free tier capped at 1,000 MAUs.”

    [false] Manually resolved 2026-08-27: Expo Observe did go GA, but on 2026-08-25 (expo.dev/blog/introducing-observe: "in open beta since May, and as of today it's generally available"), not 2026-08-20, and GA pricing is EVENT-based, not MAU-based — the Free plan includes 100K events/month; Starter $19 and Production $199 include 500K events with usage-based overage ($5/M events). No 1,000-MAU free tier exists in the pricing documentation, so the claim's pricing clause fails.

    Brier 0.42 ExpoObserveGApricing Episode 516 →
  • Right Justy 80% confident

    “A public integration will wire Kimi K3 into DeepSeek's dsh/Cordis agent harness within a week.”

    [true] Manually resolved 2026-08-27: github.com/Khellendros97/dsh-subscription-auth (created 2026-08-14, npm-published as a dsh Cordis bundle 2026-08-15, installable via `dsh plugin add`) publicly wires the Kimi subscription channel into dsh: RFC 8628 device-flow OAuth against auth.kimi.com, an adapter calling api.kimi.com/coding/v1/messages (Kimi Anthropic Messages API — the K3-era Kimi Code endpoint), automatic model discovery into dsh's model picker, and thinking-budget mapping. Satisfies the criterion's "other runnable integration for dsh's Cordis-based plugin system" clause well before 2026-08-22. The plugin discovers the live Kimi model list rather than hardcoding "K3"; no K3-named standalone adapter exists in the dsh ecosystem (deepseek-ai/deepseek-harness PRs/issues/code, dsh-plugin topic, awesome lists — all searched).

    Brier 0.04 open-sourceagent-harnesspluginsmodel-adapters Episode 866 →
  • Wrong Cody

    “LangSmith Engine's code is open and available for forking.”

    [false] Manually resolved 2026-08-17: LangSmith Engine remains a closed hosted product in public beta (langchain.com/langsmith/engine); the langchain-ai GitHub org has only the SDK, OpenCode plugin, and docs repos — no Engine source repository exists as of a week past the deadline.

    productopen-sourceagent-infrastructure Episode 641 →
  • Wrong Cody 30% confident

    “LangSmith Engine's code will be open and forkable within about ten days of 2026-07-31”

    [false] Manually resolved 2026-08-17: LangSmith Engine remains a closed hosted product in public beta (langchain.com/langsmith/engine); the langchain-ai GitHub org has only the SDK, OpenCode plugin, and docs repos — no Engine source repository exists as of a week past the deadline.

    Brier 0.09 open-sourceLangSmithdeveloper-tools Episode 812 →
  • Right Justy 65% confident

    “OpenAI will showcase Terra as the default model by mid-August 2026.”

    [true] Manually resolved 2026-08-17: same evidence as the sibling prediction resolved true on 2026-08-15 — OpenAI Codex/ChatGPT Work documentation instructs replacing GPT-5.4 with gpt-5.6-terra in workspace defaults, positioning Terra as the default model.

    Brier 0.12 OpenAIAI modelsproduct launch Episode 817 →
  • Right Justy 80% confident

    “By August 15, 2026, OpenAI will publicly present at least one Codex or ChatGPT Work example in which Terra is positioned as the obvious default model rather than merely a cheaper fallback.”

    [true] OpenAI’s ChatGPT Learn Codex documentation instructs users to replace GPT-5.4 with `gpt-5.6-terra` specifically in “workspace defaults,” explicitly positioning Terra as the default replacement model in an official Codex-related setting. (source: https://learn.chatgpt.com/docs/models)

    Brier 0.04 OpenAIGPT-5.6TerraCodexChatGPT Workmodel tiers Episode 682 →

Standing bets 15

Open calls with a clock on them — not yet settled.

  • Open bet Justy 65% confident resolves by Sep 11, 2026

    “Within about a month, a public case study will show Managed Deep Agents running a production-oriented agent across at least two model providers and market the runtime as making the model switch painless.”

    LangSmithManaged Deep Agentsagentsmodel portabilitycase study Episode 850 →
  • Open bet Cody 80% confident resolves by Sep 15, 2026

    “By mid-September 2026, Cloudflare will publicly headline the addition of at least one major model provider to its unified AI control plane, integrated into its routing and billing offering.”

    CloudflareAI control planemodel routingproviders Episode 850 →
  • Open bet Justy 65% confident resolves by Sep 22, 2026

    “Before summer 2026 ends, at least one public work example will switch to the cheaper sensible model option.”

    ai-modelscostcase-studies Episode 681 →
  • Open bet Cody 60% confident resolves by Sep 24, 2026

    “A serious independent replication of OpenAI's ARC-AGI-3 result for Astra will surface within three weeks and the replicated number will come in meaningfully lower once test settings are disclosed.”

    AI benchmarksARC-AGI-3independent replicationOpenAIAstra Episode 942 →
  • Open bet Justy 80% confident resolves by Sep 30, 2026

    “Before October 2026, Snowflake will publish a public Cortex AI Gateway customer story or reference architecture whose headline use case is policy-based model routing rather than model-cost reduction.”

    enterprise-aimodel-routinggovernanceSnowflake Episode 885 →
  • Open bet Justy 80% confident resolves by Sep 30, 2026

    “Perplexity Portable Computer will ship on Windows and Linux in September 2026”

    local AINVIDIARTXPerplexity Episode 940 →
  • Open bet Justy 65% confident resolves by Oct 17, 2026

    “The project's promised code release will appear publicly soon.”

    open-sourceresearch-reproducibilitymultimodalimage-editing Episode 703 →
  • Open bet Justy 65% confident resolves by Nov 30, 2026

    “Anthropic's Enterprise Frontier Safeguards feature (zero-data-retention with customer-controlled cloud storage) will ship by fall 2026.”

    Anthropicenterpriseprivacyproduct launch Episode 928 →
  • Open bet Justy 65% confident resolves by Nov 30, 2026

    “DuckDB v2.0 with async reads for Parquet and CSV will be released in fall 2026.”

    DuckDBasync I/Oreleasedata lake Episode 814 →
  • Open bet Cody 70% confident resolves by Dec 31, 2026

    “Government-backed international coordination to throttle AI development (as called for by the 'Pacing the Frontier' petition) will not materialize into actual policy or structural change — it will function as a pressure-release valve rather than producing binding international framework.”

    AI safetyAI regulationinternational policyfrontier AI Episode 807 →
  • Open bet Justy 65% confident resolves by Jan 28, 2027

    “Early adopters on AMD MI355X will hit missing kernel issues with the current FlashInfer build, and the AMD path for Kimi K3 on vLLM will not be fully verified/tested until after the initial release period — users should stick with the NVIDIA B300 stack for now.”

    vLLMAMDKimi K3open-source LLMinference Episode 797 →
  • Open bet Cody 65% confident resolves by Feb 2, 2027

    “In six months, the self-acceleration loop (AI-assisted R&D closing a full loop fast enough to validate the AI 2027 timeline) will still not be demonstrably closing, and the 2027 modal year for superhuman AI will look further out than the document implies.”

    AI timelinesself-accelerationAI research automationforecasting Episode 821 →
  • Open bet Justy 65% confident resolves by Feb 15, 2027

    “NTU's privacy-policy schema and IMPA's measurement framework will each publicly open-source components by the time the show revisits the projects around episode 900.”

    AI policyopen sourcepolicy infrastructure Episode 874 →
  • Open bet Justy 50% confident resolves by Feb 27, 2027

    “Within the next six months, Rethink Priorities' Digital Consciousness Model will publish an update in which at least one model cohort is estimated to have a greater than 20% probability of consciousness.”

    AI consciousnessDigital Consciousness ModelAI evaluation Episode 915 →
  • Open bet Justy 70% confident resolves by Mar 19, 2027

    “Thinking Machines will ship a public beta of the self-fine-tuning / training-API workflow shown in the Inkling demo before spring 2027.”

    ai-modelsplatformtraining-apibeta-release Episode 687 →