影片高階EN9,639 次觀看
Vertical Mobility: Inference from MVP to Trillion-Parameter Workloads — Sitanshu Gupta, CoreWeave
來源 AI Engineer
看影片(在新分頁開啟原站)連到 AI Engineer
摘要
Sitanshu Gupta 介紹 CoreWeave 正在構建的推理平台,支援無伺服器、預留通量與專屬部署。平台透過 KV 快取感知的路由、快取外掛與量化技術,最佳化不同工作負載的價格與效能表現。
Sitanshu Gupta explains CoreWeave's unified inference platform for serverless, provisioned, and dedicated deployments, highlighting KV cache optimization and workload scheduling strategies.
摘要、重點與章節標題由語言模型整理,細節(誰說的、數字、先後)可能有誤;要引用請以原始內容為準。
重點
- 支援無伺服器、預留通量與專屬部署三種消費模式。
- 透過 KV 快取感知的路由與快取外掛降低重計算成本。
- 將批處理工作負載安排在閒暇時間以最大化資源利用率。
章節
依話題轉折切分,標題由 AI 產生
- 00:00Four months in, leading inference at CoreWeave
- 01:22Serverless, dedicated, and provisioned throughput
- 03:10Four workload shapes and a game of Tetris
- 05:15The request flow, from gateway to GPU
- 07:17Why the router is KV cache aware
- 10:21Scheduling batch into idle overnight capacity
- 11:05Offloading KV cache between turns
- 12:31Quantization and custom trained speculators
提到的工具與公司
- CoreWeave
- VLM
- NVIDIA GPUs
- LM Cache
- Weights and Biases
適合誰看
正在規劃或建構大規模 AI 推理基礎設施的工程師與架構師。
摘要依據
- 講者
- Sitanshu Gupta
- 依據
- 自動字幕
為什麼排在這裡
- 人氣
- 0.71
- 新鮮
- 0.94
在主題頁與搜尋結果裡,名次由相關、人氣、新鮮三個分數決定;這一頁沒有搜尋的關鍵字,所以沒有相關分數。排序怎麼算
相關內容
- How to Get Started with Managed Inference in CoreWeave ARENA影片 ・ CoreWeave(在新分頁開啟原站)
- How to get started with managed inference in CoreWeave ARENA影片 ・ CoreWeave ・ 3 分鐘(在新分頁開啟原站)
- Inside the $41B AI Cloud Challenging Big Tech | CoreWeave SVPPodcast ・ Gradient Dissent ・ 53 分鐘
- The Inference Engineering Masterclass — Philip Kiely & Ali Taha, BasetenPodcast ・ Latent Space ・ 1 小時 41 分
- The Future of AI Infrastructure with CoreWeavePodcast ・ Practical AI ・ 50 分鐘
- Stanford CS336 Language Modeling from Scratch | Spring 2026 | Guest Lecture: Dan Fu影片 ・ Stanford Online ・ 1 小時 12 分
摘要由 AI 根據原文產生,可能有誤;完整內容請看原站。看影片(在新分頁開啟原站)
