看影片(在新分頁開啟原站)連到 AI Engineer
摘要
John Ousterhout 指出 AI 推理與代理工作負載中,小訊息的延遲比吞吐量更重要,傳統 TCP 與 RDMA 因無訊息邊界且延遲控制機制不適合,導致短訊息尾延遲高。Homa 是全新設計的協議,透過訊息本位、接收端延遲控制與交換機優先佇列,大幅降低短訊息延遲並提升整體效能。
John Ousterhout presents Homa, a new network protocol designed to drastically reduce tail latency for short messages in AI inference and agentic workloads by using message-based transport and receiver-side congestion control.
摘要、重點與章節標題由語言模型整理,細節(誰說的、數字、先後)可能有誤;要引用請以原始內容為準。
重點
- AI 推理與代理工作負載中,短訊息的尾延遲會限制整體系統吞吐量。
- TCP 與 RDMA 因無訊息邊界且延遲控制位於傳送端,無法有效處理短訊息。
- Homa 採用訊息本位、接收端控制與優先佇列,將短訊息延遲降低至微秒級。
章節
依話題轉折切分,標題由 AI 產生
- 00:00Why latency is becoming the metric that matters
- 02:19The old workload: gigabytes and throughput
- 03:09The new workload: metadata and coordination
- 04:24How one slow exchange stalls every GPU
- 06:07Incast, and where the queue actually builds
- 07:10Why congestion control lives on the wrong end
- 10:05A byte stream has no message boundaries
- 11:08Homa, and a clean slate redesign
- 12:12Messages, not streams
- 13:30Controlling congestion from the receiver
- 14:59Using the priority queues already in the switch
- 15:52The benchmark against TCP
提到的工具與公司
- Homa
- TCP
- RDMA
- SRPT
適合誰看
負責 AI 系統架構、網路最佳化或正在使用大規模 GPU 叢集的開發人員。
摘要依據
- 講者
- John Ousterhout
- 依據
- 自動字幕
為什麼排在這裡
- 人氣
- 0.92
- 新鮮
- 0.94
在主題頁與搜尋結果裡,名次由相關、人氣、新鮮三個分數決定;這一頁沒有搜尋的關鍵字,所以沒有相關分數。排序怎麼算
相關內容
- Why AI needs a new kind of supercomputer network - Episode 18Podcast ・ OpenAI Podcast ・ 38 分鐘
- Daniel Homola - From Multi Agent Patterns to Reliable Orchestration影片 ・ Berkeley RDI ・ 6 分鐘(在新分頁開啟原站)
- Scaling Agentic Inference Across Heterogeneous Compute with Zain Asgar - #757Podcast ・ The TWIML AI Podcast ・ 49 分鐘
- Introduction to Mission Control agent影片 ・ CoreWeave ・ 3 分鐘(在新分頁開啟原站)
- Routing LLM Inference in Production: From Engine Signals to Policy — Qianru Lao & Lu Zhang, OpenAI影片 ・ AI Engineer ・ 18 分鐘(在新分頁開啟原站)
- 高效电商 AI 智能体解剖指南 | 选自 Anthropic 的 Claude 博客文章 ・ 寶玉
摘要由 AI 根據原文產生,可能有誤;完整內容請看原站。看影片(在新分頁開啟原站)
