跳到主要內容
AI 武林
影片高階EN2,844 次觀看

Build A Reasoning Model From Scratch 7: Reinforcement Learning 2 (Tracking, KL Terms, And More)

來源 Sebastian Raschka人物 Sebastian Raschka

看影片(在新分頁開啟原站)連到 YouTube・Sebastian Raschka

摘要

深入解析從零訓練推理模型的第七集,重點探討強化學習中的 GRPO 演算法。內容涵蓋如何解讀訓練指標(如損失、獎勵平均值、優勢統計與熵值),並示範如何實現截斷策略比率與 KL 散度懲罰項來穩定訓練。觀眾能學習如何監控訓練過程、診斷不穩定問題,以及參考最新研究最佳化訓練設定。

This video analyzes GRPO training metrics, implements clipped policy ratios and KL loss, and offers tips for stabilizing reinforcement learning-based reasoning models.

摘要、重點與章節標題由語言模型整理,細節(誰說的、數字、先後)可能有誤;要引用請以原始內容為準。

重點

  • 詳細解讀 GRPO 訓練中的損失、獎勵與熵值等關鍵指標。
  • 實作截斷策略比率與 KL 散度項以穩定模型訓練。
  • 提供參考論文技巧,幫助最佳化推理模型的訓練設定。

章節

依話題轉折切分,標題由 AI 產生

  1. 00:00Introduction and recap
  2. 01:52Interpreting basic GRPO training metrics
  3. 06:34Planned improvements to GRPO
  4. 08:58Running longer training jobs with Python scripts
  5. 13:39Running the baseline GRPO training script
  6. 17:29Loading and plotting training logs
  7. 19:29Diagnosing unstable training
  8. 23:55Evaluating checkpoints on MATH-500
  9. 26:26Downloading existing checkpoints
  10. 30:09Tracking advantage statistics
  11. 34:53Understanding entropy
  12. 40:32Computing entropy in PyTorch
  13. 44:17Interpreting entropy values
  14. 48:58Adding entropy tracking to GRPO
  15. 53:36Analyzing advantage and entropy metrics
  16. 56:18Stabilizing GRPO with clipped policy ratios
  17. 1:03:27Implementing the clipped policy loss
  18. 1:09:39Analyzing clipped policy training results
  19. 1:11:25KL divergence and reward hacking
  20. 1:15:12Adding a KL loss term
  21. 1:20:34Limitations of the simplified KL loss
  22. 1:23:04Format rewards and think tags
  23. 1:25:47Adding special tokens to the tokenizer
  24. 1:30:29Implementing the format reward
  25. 1:35:56Analyzing format reward training
  26. 1:38:25Rewarding format only for correct answers
  27. 1:40:48Further GRPO improvements from research
  28. 1:45:43Next steps and distillation

提到的工具與公司

適合誰看

適合具備程式基礎、正在學習從零訓練大型語言模型或強化學習的開發者與研究者。

摘要依據

講者
Sebastian Raschka
依據
人工字幕

為什麼排在這裡

人氣
0.50
新鮮
1.00

在主題頁與搜尋結果裡,名次由相關、人氣、新鮮三個分數決定;這一頁沒有搜尋的關鍵字,所以沒有相關分數。排序怎麼算

這個來源最近的內容

Sebastian Raschka 的所有內容

摘要由 AI 根據原文產生,可能有誤;完整內容請看原站。看影片(在新分頁開啟原站)