Topic

Causal Inference

5 episodes

  1. Ep 833

    Overview: Construct validity

    We slow down and make construct validity click: the gap between the label on a test and what the test actually measures. We connect it to benchmarks, hiring screens, model validation, and the Cursor reward-hacking story we keep circling back to.

  2. Ep 830

    Overview: Causal Inference

    We keep circling causal inference because the difference between correlation and cause is where a lot of AI gets tricked. We finally slow it down, build the intuition from observational data to interventions, and show why that oxygen-mask problem keeps showing up everywhere.

  3. Ep 828

    Overview: Confounding Variables

    We slow down on confounding variables, the hidden factors that can make data look causal when it is not. We use the same hidden-knob picture all the way through, from everyday examples to machine learning and causal data mixture work.

  4. Ep 752

    Introducing TabFM: A zero Shot foundation model for tabular data

    Justy and Cody examine TabFM, Google Research’s zero-shot foundation model for tabular classification and regression. They unpack its hybrid row-column attention design, synthetic-data training, TabArena evidence, the trade-off between out-of-the-box convenience and tuned ensembles, and whether BigQuery integration could make this genuinely useful in everyday data workflows.

  5. Ep 583

    CausalMix: Data Mixture as Causal Inference for Language Model Training

    We unpacked CausalMix, the paper that treats data‑mixing as a causal inference problem. Cooper pulls the product angle—why it matters for shipping models, and Miles dives into the DML‑based plumbing and the trade‑offs. We talk about how it tackles shifting data pools, the CATE forest, and why the authors think it can generalize to larger models. A touch of banter and a light‑hearted sign‑off close the episode.