看影片(在新分頁開啟原站)連到 Alexander Amini
摘要
深入介紹強化學習與深度學習的結合,解釋狀態、行動與獎勵等核心概念,並探討 Q 學習與策略梯度等演算法。學員將了解如何利用 AI 在無監督環境中自我演進,達到超越人類水準的表現。
This lecture explains reinforcement learning fundamentals and compares Q-learning with policy gradient methods.
摘要、重點與章節標題由語言模型整理,細節(誰說的、數字、先後)可能有誤;要引用請以原始內容為準。
重點
- 釐清強化學習中狀態、行動與獎勵的基礎定義。
- 比較 Q 學習與策略梯度演算法的優缺點與應用。
- 了解 AlphaGo 等模型如何實現超越人類的策略。
章節
依話題轉折切分,標題由 AI 產生
- 00:00Introduction
- 02:20Classes of learning problems
- 06:33Definitions
- 12:30The Q function
- 17:29Deeper into the Q function
- 23:12Deep Q Networks
- 30:36Atari results and limitations
- 34:24Policy learning algorithms
- 39:31Discrete vs continuous actions
- 43:21Training policy gradients
- 49:10RL in real life
- 51:33VISTA simulator
- 53:24AlphaGo and AlphaZero and MuZero
- 58:58Summary
適合誰看
正在學習深度學習或機器學習理論,希望理解強化學習原理的學生或研究者。
摘要依據
- 依據
- 自動字幕
為什麼排在這裡
- 人氣
- 0.53
- 新鮮
- 0.04
在主題頁與搜尋結果裡,名次由相關、人氣、新鮮三個分數決定;這一頁沒有搜尋的關鍵字,所以沒有相關分數。排序怎麼算
相關內容
- MIT 6.S191 (2025): Reinforcement Learning影片 ・ Alexander Amini ・ 1 小時 2 分(在新分頁開啟原站)
- MIT 6.S191: Reinforcement Learning影片 ・ Alexander Amini ・ 59 分鐘(在新分頁開啟原站)
- Stanford AA203 Optimal and Learning-Based Control | Spring 2026 | Lecture 17: RL Value-Based Methods影片 ・ Stanford Online ・ 1 小時 18 分
- Stanford CS229 Machine Learning | Spring 2026 | Lecture 18: GMM (EM), PCA影片 ・ Stanford Online ・ 1 小時 16 分
- Reinforcement Learning: Essential Concepts影片 ・ StatQuest with Josh Starmer ・ 18 分鐘(在新分頁開啟原站)
- 人工智慧:機器學習與理論基礎 05. Reinforcement learning影片 ・ NTU OpenCourseWare ・ 2 小時 16 分(在新分頁開啟原站)
摘要由 AI 根據原文產生,可能有誤;完整內容請看原站。看影片(在新分頁開啟原站)
