跳到主要內容
AI 武林
影片高階EN1.6 萬 次觀看

Weight Folding, CUDA Streams, and the Bug That Made My Model Speak Backwards — Filip Makraduli

來源 AI Engineer

看影片(在新分頁開啟原站)連到 AI Engineer

其他版本:AIE Talks 摘要頁(在新分頁開啟)

摘要

Filip Makraduli 分享 FlashNorm 技術,透過將 RMS 規範係數合併至權重並延遲除法,讓矩陣乘法與向量運算並行執行,提升 Transformer 推理速度。他還揭露並修復了因 CUDA 串流隱式匯總而導致的模型輸出重複與延遲的 Bug,並說明為何需要開放式推理引擎來部署自定義核碼。

Filip Makraduli explains FlashNorm optimizations for RMS norm in transformers and shares a CUDA stream bug fix that caused model output errors.

摘要、重點與章節標題由語言模型整理,細節(誰說的、數字、先後)可能有誤;要引用請以原始內容為準。

重點

  • FlashNorm 透過權重合併與延遲除法,讓 RMS 規範與矩陣運算並行。
  • CUDA 串流隱式匯總導致讀取未完成緩區,引發模型輸出重複與延遲。
  • 開放式推理引擎能支援自定義核碼部署,避免租端限制。

章節

依話題轉折切分,標題由 AI 產生

  1. 00:00Two lines of algebra for RMS norm
  2. 02:07Why a layer with no math costs so much time
  3. 03:44Weight folding, deferred division, dropped pre norm
  4. 06:33The output that came from the past
  5. 07:25Tensor cores and CUDA cores in parallel
  6. 08:45An implicit join and a race condition
  7. 09:58Making the post scale wait on both streams
  8. 10:42Results on open models, and what works out of the box
  9. 13:54Why you cannot do kernel work on a rented endpoint
  10. 16:00Paper, repo, and where to find him

提到的工具與公司

  • FlashNorm
  • RMS norm
  • CUDA
  • CUDA cores
  • Hugging Face
  • Site

適合誰看

正在最佳化 Transformer 推理速度、熟悉 CUDA 程式設計或部署自定義模型開發者的工程師。

摘要依據

講者
Filip Makraduli
依據
自動字幕

為什麼排在這裡

人氣
0.76
新鮮
0.94

在主題頁與搜尋結果裡,名次由相關、人氣、新鮮三個分數決定;這一頁沒有搜尋的關鍵字,所以沒有相關分數。排序怎麼算

摘要由 AI 根據原文產生,可能有誤;完整內容請看原站。看影片(在新分頁開啟原站)