影片進階EN7.7 萬 次觀看
Reinforcement Learning from Human Feedback explained with math derivations and the PyTorch code.
來源 Umar Jamil
看影片(在新分頁開啟原站)連到 Umar Jamil
摘要
詳細解釋人類回饋強化學習(RLHF)的數學原理與 PyTorch 實作,涵蓋從語言模型到 PPO 演算法的推導與程式碼。觀眾能透過視覺直覺理解梯度最佳化、獎勵模型與優勢估計等核心概念。
This video explains the mathematical principles and PyTorch implementation of Reinforcement Learning from Human Feedback, covering RLHF, PPO, and reward models with visual intuition.
這筆內容還沒有取得字幕或內文,這段摘要只根據標題與說明欄產生,可能不夠準確;實際內容請以原站為準。
提到的工具與公司
- PyTorch
- PPO
適合誰看
具備程式基礎或機器學習背景,想深入理解 RLHF 原理與實作的開發者。
摘要依據
- 依據
- 標題與說明欄(還沒有取得字幕或內文)
為什麼排在這裡
- 人氣
- 0.39
- 新鮮
- 0.03
在主題頁與搜尋結果裡,名次由相關、人氣、新鮮三個分數決定;這一頁沒有搜尋的關鍵字,所以沒有相關分數。排序怎麼算
相關內容
- Reinforcement Learning with Human Feedback (RLHF) in 4 minutes影片 ・ Sebastian Raschka ・ 4 分鐘
- Reinforcement Learning with Human Feedback (RLHF), Clearly Explained!!!影片 ・ StatQuest with Josh Starmer ・ 18 分鐘
- 【生成式AI導論 2024】第8講:大型語言模型修練史 — 第三階段: 參與實戰,打磨技巧 (Reinforcement Learning from Human Feedback, RLHF)影片 ・ Hung-yi Lee ・ 37 分鐘
- Build A Reasoning Model From Scratch 6: Reinforcement Learning 1 (Implementing GRPO for RLVR)影片 ・ Sebastian Raschka
- Reinforcement Learning with Neural Networks: Mathematical Details影片 ・ StatQuest with Josh Starmer ・ 25 分鐘
- 5 useful things you'll learn in my new post-training textbook (shipping now!)文章 ・ Nathan Lambert(部落格)
摘要由 AI 根據標題與說明欄產生(還沒有取得原文),可能有誤;完整內容請看原站。看影片(在新分頁開啟原站)
