跳到主要內容
AI 武林
影片高階EN8,989 次觀看

Two Bugs That Hid in Plain Sight: A vLLM Debugging Detective Story — Asaf Gardin & Yuval Belfer

來源 AI Engineer

看影片(在新分頁開啟原站)連到 AI Engineer

其他版本:AIE Talks 摘要頁(在新分頁開啟)

摘要

Asaf Gardin 與 Yuval Belfer 分享在 vLLM 部署 AI21 Jamba 模型時發現的兩個隱性錯誤。第一個導致偶爾輸出無意義內容,根源在於解碼階段早於預填充階段執行;第二個則由 32 位索引溢位引發。透過降低記憶體使用率與對照基準模型,團隊成功定位並修復了這些問題。

Two engineers explain how they debugged silent inference bugs in vLLM serving the Jamba model by manipulating memory constraints and comparing log probabilities against a baseline.

摘要、重點與章節標題由語言模型整理,細節(誰說的、數字、先後)可能有誤;要引用請以原始內容為準。

重點

  • vLLM 中解碼階段早於預填充導致 Mamba 模型讀取陳舊狀態。
  • 32 位索引在超過 40 億時溢位,需改用 64 位大小型別修正。
  • 透過降低記憶體壓力與對照基準模型可復現並定位隱性錯誤。

章節

依話題轉折切分,標題由 AI 產生

  1. 00:00No crash, no warning, high confidence
  2. 02:48The one in a thousand gibberish
  3. 03:43Reproducing it fast by starving the GPU
  4. 05:18A logprob comparison against a plain baseline
  5. 08:48Threading a request ID through the forward pass
  6. 09:30Decode before prefill, and why only Mamba noticed
  7. 11:47Case two: logprob spikes every twelve steps
  8. 12:29A lever that changes the shape of failure
  9. 14:33A 32 bit index that wrapped around
  10. 15:42Two scenes, one criminal

提到的工具與公司

  • vLLM
  • Jamba
  • Mamba
  • Hugging Face Transformers
  • CUDA

適合誰看

負責大語言模型推理系統開發、除錯或效能最佳化的工程師。

摘要依據

講者
Asaf Gardin、Yuval Belfer
依據
自動字幕

為什麼排在這裡

人氣
0.71
新鮮
0.94

在主題頁與搜尋結果裡,名次由相關、人氣、新鮮三個分數決定;這一頁沒有搜尋的關鍵字,所以沒有相關分數。排序怎麼算

摘要由 AI 根據原文產生,可能有誤;完整內容請看原站。看影片(在新分頁開啟原站)