Exploring Next / Topics / On Policy Learning Topic On Policy Learning 1 episode Ep 709 Jul 18, 2026 Seed: Self Evolving On Policy Distillation for Agentic Reinforcement Learning Seed tackles the credit-assignment problem in long-horizon agent reinforcement learning by turning completed trajectories into evolving natural-language hindsight skills, then distilling their effect into dense token-level training signals. Vince sees a potentially shippable training pattern for teams already running agentic RL; Ava likes the on-policy design but wants stronger evidence that self-generated skills do not amplify the model’s own blind spots. AgentsTrainingReinforcement Learning From Human FeedbackCredit Assignment