影片進階EN6.7 萬 次觀看
Coding LLaMA 2 from scratch in PyTorch - KV Cache, Grouped Query Attention, Rotary PE, RMSNorm
來源 Umar Jamil
看影片(在新分頁開啟原站)連到 Umar Jamil
摘要
詳細示範如何從零開始用 PyTorch 編寫 LLaMA 2 模型,涵蓋旋轉位置嵌入、RMS 歸一化、群組查詢注意力等核心機制與推理策略。觀眾可學習 Transformer 架構的實作細節與自訂模型的開發流程。
This video teaches how to code the LLaMA 2 model from scratch in PyTorch, covering key components like rotary embeddings and inference strategies.
這筆內容還沒有取得字幕或內文,這段摘要只根據標題與說明欄產生,可能不夠準確;實際內容請以原站為準。
提到的工具與公司
- PyTorch
- LLaMA 2
- SwiGLU
- RMSNorm
- Rotary PE
- GQA
適合誰看
具備 Transformer 基礎知識並想深入理解與實作大型語言模型的開發者。
摘要依據
- 依據
- 標題與說明欄(還沒有取得字幕或內文)
為什麼排在這裡
- 人氣
- 0.36
- 新鮮
- 0.01
在主題頁與搜尋結果裡,名次由相關、人氣、新鮮三個分數決定;這一頁沒有搜尋的關鍵字,所以沒有相關分數。排序怎麼算
相關內容
- LLaMA explained: KV-Cache, Rotary Positional Embedding, RMS Norm, Grouped Query Attention, SwiGLU影片 ・ Umar Jamil ・ 1 小時 11 分
- Coding a Transformer from scratch on PyTorch, with full explanation, training and inference.影片 ・ Umar Jamil ・ 2 小時 59 分
- What I Learned From Implementing LLM Architectures From Scratch (And How to Get Started)影片 ・ Sebastian Raschka ・ 53 分鐘
- Coding a ChatGPT Like Transformer From Scratch in PyTorch影片 ・ StatQuest with Josh Starmer ・ 31 分鐘
- Build an LLM from Scratch 4: Implementing a GPT model from Scratch To Generate Text影片 ・ Sebastian Raschka ・ 1 小時 46 分
- Implementing GPT-2 From Scratch (Transformer Walkthrough Part 2/2)影片 ・ Neel Nanda ・ 1 小時 19 分
摘要由 AI 根據標題與說明欄產生(還沒有取得原文),可能有誤;完整內容請看原站。看影片(在新分頁開啟原站)
