影片高階EN8,989 次觀看
Two Bugs That Hid in Plain Sight: A vLLM Debugging Detective Story — Asaf Gardin & Yuval Belfer
來源 AI Engineer
看影片(在新分頁開啟原站)連到 AI Engineer
摘要
Asaf Gardin 與 Yuval Belfer 分享在 vLLM 部署 AI21 Jamba 模型時發現的兩個隱性錯誤。第一個導致偶爾輸出無意義內容,根源在於解碼階段早於預填充階段執行;第二個則由 32 位索引溢位引發。透過降低記憶體使用率與對照基準模型,團隊成功定位並修復了這些問題。
Two engineers explain how they debugged silent inference bugs in vLLM serving the Jamba model by manipulating memory constraints and comparing log probabilities against a baseline.
摘要、重點與章節標題由語言模型整理,細節(誰說的、數字、先後)可能有誤;要引用請以原始內容為準。
重點
- vLLM 中解碼階段早於預填充導致 Mamba 模型讀取陳舊狀態。
- 32 位索引在超過 40 億時溢位,需改用 64 位大小型別修正。
- 透過降低記憶體壓力與對照基準模型可復現並定位隱性錯誤。
章節
依話題轉折切分,標題由 AI 產生
- 00:00No crash, no warning, high confidence
- 02:48The one in a thousand gibberish
- 03:43Reproducing it fast by starving the GPU
- 05:18A logprob comparison against a plain baseline
- 08:48Threading a request ID through the forward pass
- 09:30Decode before prefill, and why only Mamba noticed
- 11:47Case two: logprob spikes every twelve steps
- 12:29A lever that changes the shape of failure
- 14:33A 32 bit index that wrapped around
- 15:42Two scenes, one criminal
提到的工具與公司
- vLLM
- Jamba
- Mamba
- Hugging Face Transformers
- CUDA
適合誰看
負責大語言模型推理系統開發、除錯或效能最佳化的工程師。
摘要依據
- 講者
- Asaf Gardin、Yuval Belfer
- 依據
- 自動字幕
為什麼排在這裡
- 人氣
- 0.71
- 新鮮
- 0.94
在主題頁與搜尋結果裡,名次由相關、人氣、新鮮三個分數決定;這一頁沒有搜尋的關鍵字,所以沒有相關分數。排序怎麼算
相關內容
- Never Trust An LLM影片 ・ Matt Pocock ・ 14 分鐘
- Stealing Reasoning Traces from Proprietary LLM APIs — Ilia Shumailov & Alexander PanfilovPodcast ・ Machine Learning Street Talk ・ 49 分鐘
- Inside vLLM: The Engine Powering Open-Source AIPodcast ・ AI + a16z ・ 47 分鐘
- Detecting Hidden AI Failures影片 ・ Braintrust ・ 12 分鐘(在新分頁開啟原站)
- 如何从专有LLM API中窃取推理痕迹 | LLM安全 | 推理痕迹 | 大语言模型 | AI安全 | 越狱攻击 | 模型蒸馏 | 推理模型 | 加密推理 | 隐私泄露 | 提示注入影片 ・ Best Partners TV ・ 16 分鐘
- Fast inference changes what you can build影片 ・ DeepLearningAI ・ 2 分鐘(在新分頁開啟原站)
摘要由 AI 根據原文產生,可能有誤;完整內容請看原站。看影片(在新分頁開啟原站)
