影片進階EN1.2 萬 次觀看
Stanford CME296 Diffusion & Large Vision Models | Spring 2026 | Lecture 7 - Evaluation
看影片(在新分頁開啟原站)連到 Stanford Online
摘要
第七講介紹如何評估文字生成影像模型的輸出品質,涵蓋人類評分、參考無指標(如 FID、CLIPScore)與參考有指標(如 MSE、PSNR、SSIM)等評估方法。學員將學習如何從美學與提示遵循度兩個維度判斷生成結果,並了解 GenEval、LongTextBench 等主流基準測試的運作原理。
This lecture covers methods for evaluating text-to-image generation quality, including human ratings, reference-free metrics like FID, and reference-based metrics like SSIM, alongside benchmark discussions.
摘要、重點與章節標題由語言模型整理,細節(誰說的、數字、先後)可能有誤;要引用請以原始內容為準。
重點
- 學習人類評分與二值化評分的優缺點與噪音來源。
- 掌握畫素級指標(MSE、PSNR)與結構相似性指標(SSIM)的計算邏輯。
- 了解 GenEval、LongTextBench 等基準測試如何驗證提示遵循度與文字渲染能力。
章節
依話題轉折切分,標題由 AI 產生
- 00:00Introduction
- 05:19Motivation
- 10:48Human ratings
- 19:43Elo rating system
- 26:37Reference-free metrics
- 29:15Fréchet inception distance (FID)
- 42:30CLIPScore
- 44:51PickScore
- 45:41Reference-based metrics
- 48:07Mean squared error (MSE)
- 49:36Peak signal-to-noise ratio (PSNR)
- 51:54Structural similarity (SSIM)
- 1:01:09Perceptual similarity (LPIPS)
- 1:05:03Multimodal LLMs
- 1:13:10Faithfulness evaluation (TIFA)
- 1:17:29Visual question answering score (VQA)
- 1:24:40MLLM-as-a-Judge
- 1:34:17Benchmarks
提到的工具與公司
- GenEval
- DPG bench
- TIFA
- CLIPScore
- PickScore
- MLLM-as-a-Judge
適合誰看
正在修讀或研究擴增生成模型評估方法的研究生或工程師。
摘要依據
- 依據
- 人工字幕
為什麼排在這裡
- 人氣
- 0.56
- 新鮮
- 0.61
在主題頁與搜尋結果裡,名次由相關、人氣、新鮮三個分數決定;這一頁沒有搜尋的關鍵字,所以沒有相關分數。排序怎麼算
相關內容
- Stanford CME296 Diffusion & Large Vision Models | Spring 2026 | Lecture 8 - Trending Topics影片 ・ Stanford Online ・ 1 小時 50 分
- 【生成式人工智慧與機器學習導論2025】第 4 講:評估生成式人工智慧能力時可能遇到的各種坑影片 ・ Hung-yi Lee ・ 2 小時 1 分
- Stop Trusting MTEB Rankings (Kelly Hong, Chroma)文章 ・ Jason Liu(部落格)
- Evaluating the Effectiveness of LLM-Evaluators (aka LLM-as-Judge)文章 ・ Eugene Yan(部落格)
- MIT 6.S191 (2023): Text-to-Image Generation影片 ・ Alexander Amini ・ 45 分鐘(在新分頁開啟原站)
- The Race to Production-Grade Diffusion LLMs with Stefano Ermon - #764Podcast ・ The TWIML AI Podcast ・ 1 小時 3 分
摘要由 AI 根據原文產生,可能有誤;完整內容請看原站。看影片(在新分頁開啟原站)
