Topic
Confounding Variables
3 episodes
-
Overview: Causal Inference
We keep circling causal inference because the difference between correlation and cause is where a lot of AI gets tricked. We finally slow it down, build the intuition from observational data to interventions, and show why that oxygen-mask problem keeps showing up everywhere.
-
Overview: Confounding Variables
We slow down on confounding variables, the hidden factors that can make data look causal when it is not. We use the same hidden-knob picture all the way through, from everyday examples to machine learning and causal data mixture work.
-
CausalMix: Data Mixture as Causal Inference for Language Model Training
We unpacked CausalMix, the paper that treats data‑mixing as a causal inference problem. Cooper pulls the product angle—why it matters for shipping models, and Miles dives into the DML‑based plumbing and the trade‑offs. We talk about how it tackles shifting data pools, the CATE forest, and why the authors think it can generalize to larger models. A touch of banter and a light‑hearted sign‑off close the episode.