看影片(在新分頁開啟原站)連到 AI Engineer
摘要
Philip Kiely 回顧自 2026 年 2 月著作出版後,推論工程的新進展,涵蓋本地與資料中心兩種場景。他討論了 TurboQuant 因計算成本過高而不適合生產環境,並介紹了 KV 壓縮技術與基於訓練的推斷解碼新方法 DFlash 及 DSpark。
Philip Kiely reviews recent advancements in inference engineering, covering TurboQuant limitations, KV compaction, and new speculative decoding methods like DFlash and DSpark.
摘要、重點與章節標題由語言模型整理,細節(誰說的、數字、先後)可能有誤;要引用請以原始內容為準。
重點
- 推論工程分為本地壓縮與資料中心降低延遲兩種策略。
- TurboQuant 節省記憶體但計算成本過高,不適合生產。
- DFlash 與 DSpark 利用擴散模型大幅提升推斷解碼效率。
章節
依話題轉折切分,標題由 AI 產生
- 00:00Three years at the World's Fair, and a book
- 02:30Two kinds of inference engineering: local and data center
- 03:52Training for inference blurs the handoff
- 05:20What happened to TurboQuant
- 07:37Where four bit quantization actually lands
- 09:28KV compaction and a learned cache
- 12:11Speculation, from small models to trained drafters
- 13:32DFlash: diffusion for drafting
- 15:11DSpark, and continuous speculator retraining
- 17:04What comes next
提到的工具與公司
- TurboQuant
- NVFP4
- DFlash
- DSpark
- Eagle 3
- Medusa
- B200
- Nvidia Dynamo
適合誰看
正在開發或最佳化大型語言模型推斷系統的工程師與架構師。
摘要依據
- 講者
- Philip Kiely
- 依據
- 自動字幕
為什麼排在這裡
- 人氣
- 0.75
- 新鮮
- 0.94
在主題頁與搜尋結果裡,名次由相關、人氣、新鮮三個分數決定;這一頁沒有搜尋的關鍵字,所以沒有相關分數。排序怎麼算
相關內容
- The Inference Engineering Masterclass — Philip Kiely & Ali Taha, BasetenPodcast ・ Latent Space ・ 1 小時 41 分
- How to Engineer AI Inference Systems with Philip Kiely - #766Podcast ・ The TWIML AI Podcast ・ 55 分鐘
- NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light" — Nader Khalil (Brev), Kyle Kranen (Dynamo)Podcast ・ Latent Space ・ 1 小時 24 分
- DSpark: DeepSeek-V4's Insane Compute Optimization Explained影片 ・ bycloud ・ 16 分鐘(在新分頁開啟原站)
- Large Transformer Model Inference Optimization文章 ・ Lilian Weng(部落格)
- EP131 - Google 最新 TurboQuant 技術血洗記憶體股票!華爾街反應過度了嗎?深入解析、PAMO車禍線上律師影片 ・ 科技浪 ・ 1 小時 25 分
摘要由 AI 根據原文產生,可能有誤;完整內容請看原站。看影片(在新分頁開啟原站)
