影片高階EN2,844 次觀看
Build A Reasoning Model From Scratch 7: Reinforcement Learning 2 (Tracking, KL Terms, And More)
看影片(在新分頁開啟原站)連到 YouTube・Sebastian Raschka
摘要
深入解析從零訓練推理模型的第七集,重點探討強化學習中的 GRPO 演算法。內容涵蓋如何解讀訓練指標(如損失、獎勵平均值、優勢統計與熵值),並示範如何實現截斷策略比率與 KL 散度懲罰項來穩定訓練。觀眾能學習如何監控訓練過程、診斷不穩定問題,以及參考最新研究最佳化訓練設定。
This video analyzes GRPO training metrics, implements clipped policy ratios and KL loss, and offers tips for stabilizing reinforcement learning-based reasoning models.
摘要、重點與章節標題由語言模型整理,細節(誰說的、數字、先後)可能有誤;要引用請以原始內容為準。
重點
- 詳細解讀 GRPO 訓練中的損失、獎勵與熵值等關鍵指標。
- 實作截斷策略比率與 KL 散度項以穩定模型訓練。
- 提供參考論文技巧,幫助最佳化推理模型的訓練設定。
章節
依話題轉折切分,標題由 AI 產生
- 00:00Introduction and recap
- 01:52Interpreting basic GRPO training metrics
- 06:34Planned improvements to GRPO
- 08:58Running longer training jobs with Python scripts
- 13:39Running the baseline GRPO training script
- 17:29Loading and plotting training logs
- 19:29Diagnosing unstable training
- 23:55Evaluating checkpoints on MATH-500
- 26:26Downloading existing checkpoints
- 30:09Tracking advantage statistics
- 34:53Understanding entropy
- 40:32Computing entropy in PyTorch
- 44:17Interpreting entropy values
- 48:58Adding entropy tracking to GRPO
- 53:36Analyzing advantage and entropy metrics
- 56:18Stabilizing GRPO with clipped policy ratios
- 1:03:27Implementing the clipped policy loss
- 1:09:39Analyzing clipped policy training results
- 1:11:25KL divergence and reward hacking
- 1:15:12Adding a KL loss term
- 1:20:34Limitations of the simplified KL loss
- 1:23:04Format rewards and think tags
- 1:25:47Adding special tokens to the tokenizer
- 1:30:29Implementing the format reward
- 1:35:56Analyzing format reward training
- 1:38:25Rewarding format only for correct answers
- 1:40:48Further GRPO improvements from research
- 1:45:43Next steps and distillation
提到的工具與公司
適合誰看
適合具備程式基礎、正在學習從零訓練大型語言模型或強化學習的開發者與研究者。
摘要依據
- 講者
- Sebastian Raschka
- 依據
- 人工字幕
為什麼排在這裡
- 人氣
- 0.50
- 新鮮
- 1.00
在主題頁與搜尋結果裡,名次由相關、人氣、新鮮三個分數決定;這一頁沒有搜尋的關鍵字,所以沒有相關分數。排序怎麼算
相關內容
- Build A Reasoning Model From Scratch 6: Reinforcement Learning 1 (Implementing GRPO for RLVR)影片 ・ Sebastian Raschka ・ 1 小時 27 分 ・
- Training Agents 3: Reinforcement Learning影片 ・ Hugging Face ・ 1 小時 17 分 ・
- GRPO++: Tricks for Making RL Actually Work文章 ・ Deep (Learning) Focus(Cameron R. Wolfe) ・
- Group Relative Policy Optimization (GRPO)文章 ・ Deep (Learning) Focus(Cameron R. Wolfe) ・
- GRPO without the maths: how an RL trainer nudges the weights文章 ・ Alex Strick van Linschoten ・
- I trained (one of) the smallest Reasoning Language Models EVER from scratch影片 ・ Neural Breakdown with AVB ・
這個來源最近的內容
- Build A Reasoning Model From Scratch 5: Inference Scaling 2 (Logprob Scoring, Self-Refinement) ・ 影片 ・ 1 小時 9 分
- Build A Reasoning Model From Scratch 4: Inference Scaling 1 (Temperature, Top-p, Self-Consistency) ・ 影片 ・ 1 小時 37 分
- Build A Reasoning Model From Scratch 3: The Verifier for Evaluation and RL with Verifiable Rewards ・ 影片 ・ 1 小時 27 分
摘要由 AI 根據原文產生,可能有誤;完整內容請看原站。看影片(在新分頁開啟原站)
