看影片(在新分頁開啟原站)連到 YouTube・Julia Turc
摘要
介紹讓擴散式大型語言模型變快的技術,包含自我蒸餾、課程學習、KV 快取與區塊擴散等方法。觀眾能了解如何最佳化推理速度與降低計算成本。
This video explains techniques to accelerate diffusion LLMs, including self-distillation, guided diffusion, and approximate KV caching.
這筆內容還沒有取得字幕或內文,這段摘要只根據標題與說明欄產生,可能不夠準確;實際內容請以原站為準。
提到的工具與公司
- FlashDLM
- dLLM-Cache
- dKV-Cache
- LLaDA
- Mercury
適合誰看
正在研究或開發擴散式大型語言模型並希望提升效能的開發者。
摘要依據
- 講者
- Julia Turc
- 依據
- 標題與說明欄(還沒有取得字幕或內文)
為什麼排在這裡
- 人氣
- 0.42
- 新鮮
- 0.39
在主題頁與搜尋結果裡,名次由相關、人氣、新鮮三個分數決定;這一頁沒有搜尋的關鍵字,所以沒有相關分數。排序怎麼算
相關內容
- 加快語言模型生成速度 (2/2):KV Cache影片 ・ Hung-yi Lee ・ 39 分鐘
- The Most Absurd Way To Train LLMs... With 3x Less Memory!?影片 ・ bycloud ・ 16 分鐘
- Why Diffusion Will Win AI Inference with Inception Co-Founder and CEO Stefano Ermon影片 ・ No Priors
- Build A Reasoning Model From Scratch 2: Loading a Base Model, Text Generation, and KV Caching影片 ・ Sebastian Raschka ・ 1 小時 37 分
- How LLMs Get Faster Without Changing Their Outputs | Speculative Decoding影片 ・ Jia-Bin Huang
- 理解 KV Cache 与 Prompt Caching:LLM 推理加速的核心机制文章 ・ chaofa用代码打点酱油(袁朝发)
摘要由 AI 根據標題與說明欄產生(還沒有取得原文),可能有誤;完整內容請看原站。看影片(在新分頁開啟原站)
