影片進階EN6.6 萬 次觀看
Stanford CS329A Self-Improving AI Agents | Part 2 | Test-Time Compute Scaling
看影片(在新分頁開啟原站)連到 Stanford Online
摘要
探討如何在推理階段透過增加計算量來提升大型語言模型效能,涵蓋並行取樣、驗證機制及架構搜尋等技術。觀眾可學習如何利用測試時計算超越現有模型,並掌握相關Scaling Law與實務架構設計。
This lecture explores test-time compute scaling techniques to enhance LLM performance without retraining, covering parallel sampling, reward models, and architecture search.
摘要、重點與章節標題由語言模型整理,細節(誰說的、數字、先後)可能有誤;要引用請以原始內容為準。
重點
- 透過並行取樣與驗證,可讓小模型在特定任務上超越大模型。
- 利用過程獎勵模型與結果獎勵模型引導搜尋與架構最佳化。
- Arkon架構透過多層推理技術堆疊,顯著提升任務解決率。
章節
依話題轉折切分,標題由 AI 產生
提到的工具與公司
- Llama 3
- GPT-4o
- Claude 3.5
- DeepSeek-V3
- Palm
- Arkon
- CUDA
適合誰看
適合正在研究或實踐大語言模型推理最佳化、希望提升模型效能的開發者與研究者。
摘要依據
- 講者
- Test-Time Compute Scaling
- 依據
- 人工字幕
為什麼排在這裡
- 人氣
- 0.79
- 新鮮
- 0.79
在主題頁與搜尋結果裡,名次由相關、人氣、新鮮三個分數決定;這一頁沒有搜尋的關鍵字,所以沒有相關分數。排序怎麼算
相關內容
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters (Paper)影片 ・ Yannic Kilcher ・ 53 分鐘(在新分頁開啟原站)
- Why We Think文章 ・ Lilian Weng(部落格)
- Build A Reasoning Model From Scratch 4: Inference Scaling 1 (Temperature, Top-p, Self-Consistency)影片 ・ Sebastian Raschka ・ 1 小時 37 分(在新分頁開啟原站)
- Stanford CS329A Self-Improving AI Agents | Part 1 | Course Overview影片 ・ Stanford Online ・ 1 小時 10 分
- #490 – State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGIPodcast ・ Lex Fridman Podcast
- The End of GPU Scaling? Compute & The Agent Era — Tim Dettmers (Ai2) & Dan Fu (Together AI)Podcast ・ The MAD Podcast ・ 1 小時 4 分
摘要由 AI 根據原文產生,可能有誤;完整內容請看原站。看影片(在新分頁開啟原站)
