Topic
Causal Inference
5 episodes
-
Overview: Construct validity
We slow down and make construct validity click: the gap between the label on a test and what the test actually measures. We connect it to benchmarks, hiring screens, model validation, and the Cursor reward-hacking story we keep circling back to.
-
Overview: Causal Inference
We keep circling causal inference because the difference between correlation and cause is where a lot of AI gets tricked. We finally slow it down, build the intuition from observational data to interventions, and show why that oxygen-mask problem keeps showing up everywhere.
-
Overview: Confounding Variables
We slow down on confounding variables, the hidden factors that can make data look causal when it is not. We use the same hidden-knob picture all the way through, from everyday examples to machine learning and causal data mixture work.
-
Introducing TabFM: A zero Shot foundation model for tabular data
Justy and Cody examine TabFM, Google Research’s zero-shot foundation model for tabular classification and regression. They unpack its hybrid row-column attention design, synthetic-data training, TabArena evidence, the trade-off between out-of-the-box convenience and tuned ensembles, and whether BigQuery integration could make this genuinely useful in everyday data workflows.
-
CausalMix: Data Mixture as Causal Inference for Language Model Training
We unpacked CausalMix, the paper that treats data‑mixing as a causal inference problem. Cooper pulls the product angle—why it matters for shipping models, and Miles dives into the DML‑based plumbing and the trade‑offs. We talk about how it tackles shifting data pools, the CATE forest, and why the authors think it can generalize to larger models. A touch of banter and a light‑hearted sign‑off close the episode.