跳到主要內容
AI 武林
影片高階EN1.4 萬 次觀看

What's New in Inference Engineering — Philip Kiely, Baseten

來源 AI Engineer

看影片(在新分頁開啟原站)連到 AI Engineer

其他版本:AIE Talks 摘要頁(在新分頁開啟)

摘要

Philip Kiely 回顧自 2026 年 2 月著作出版後,推論工程的新進展,涵蓋本地與資料中心兩種場景。他討論了 TurboQuant 因計算成本過高而不適合生產環境,並介紹了 KV 壓縮技術與基於訓練的推斷解碼新方法 DFlash 及 DSpark。

Philip Kiely reviews recent advancements in inference engineering, covering TurboQuant limitations, KV compaction, and new speculative decoding methods like DFlash and DSpark.

摘要、重點與章節標題由語言模型整理,細節(誰說的、數字、先後)可能有誤;要引用請以原始內容為準。

重點

  • 推論工程分為本地壓縮與資料中心降低延遲兩種策略。
  • TurboQuant 節省記憶體但計算成本過高,不適合生產。
  • DFlash 與 DSpark 利用擴散模型大幅提升推斷解碼效率。

章節

依話題轉折切分,標題由 AI 產生

  1. 00:00Three years at the World's Fair, and a book
  2. 02:30Two kinds of inference engineering: local and data center
  3. 03:52Training for inference blurs the handoff
  4. 05:20What happened to TurboQuant
  5. 07:37Where four bit quantization actually lands
  6. 09:28KV compaction and a learned cache
  7. 12:11Speculation, from small models to trained drafters
  8. 13:32DFlash: diffusion for drafting
  9. 15:11DSpark, and continuous speculator retraining
  10. 17:04What comes next

提到的工具與公司

  • TurboQuant
  • NVFP4
  • DFlash
  • DSpark
  • Eagle 3
  • Medusa
  • B200
  • Nvidia Dynamo

適合誰看

正在開發或最佳化大型語言模型推斷系統的工程師與架構師。

摘要依據

講者
Philip Kiely
依據
自動字幕

為什麼排在這裡

人氣
0.75
新鮮
0.94

在主題頁與搜尋結果裡,名次由相關、人氣、新鮮三個分數決定;這一頁沒有搜尋的關鍵字,所以沒有相關分數。排序怎麼算

摘要由 AI 根據原文產生,可能有誤;完整內容請看原站。看影片(在新分頁開啟原站)