跳到主要內容
AI 武林
文章進階EN

Diffusion Models for Video Generation

來源 Lilian Weng(部落格)人物 Lilian Weng

讀原文(在新分頁開啟原站)連到 Lilian Weng(部落格)

摘要

介紹了從頭開始構建自定義的 Diffusion Video Models (VDM) 的方法。文章詳述了 DDIM 更新規則、VDM 的 3D U-Net 架構設計,以及 Imagen Video、Make-A-Video 等現有模型的關鍵技術細節,如空間與時間注意力機制、超解析度處理及跨時序注意力。

This article introduces the approach of building custom Diffusion Video Models (VDM) from scratch. It details DDIM update rules, the 3D U-Net architecture, and key techniques from existing models like Imagen Video and Make-A-Video, including spatial-temporal attention, super-resolution, and distill.

重點

  • DDIM 更新規則採用 v 預測引數化,有效避免色彩偏移。
  • VDM 架構將 2D U-Net 擴充套件為 3D 結構,並加入時間注意力塊捕捉時序一致性。
  • Imagen Video 使用 7 個分層模型,透過超解析度與分層蒸餾加速生成。

提到的工具與公司

  • DDIM
  • 3D U-Net
  • CLIP
  • T5

適合誰看

學習 AI 模型架構、程式開發或研究生成式 AI 的技術人員。

摘要依據

講者
Lilian Weng
依據
文章全文

為什麼排在這裡

人氣
0.75
新鮮
0.03

在主題頁與搜尋結果裡,名次由相關、人氣、新鮮三個分數決定;這一頁沒有搜尋的關鍵字,所以沒有相關分數。排序怎麼算

摘要由 AI 根據原文產生,可能有誤;完整內容請看原站。讀原文(在新分頁開啟原站)