跳到主要內容
AI 武林
影片高階EN22.3 萬 次觀看

Homa: The End of TCP for AI Clusters — John Ousterhout, Stanford

來源 AI Engineer

看影片(在新分頁開啟原站)連到 AI Engineer

其他版本:AIE Talks 摘要頁(在新分頁開啟)

摘要

John Ousterhout 指出 AI 推理與代理工作負載中,小訊息的延遲比吞吐量更重要,傳統 TCP 與 RDMA 因無訊息邊界且延遲控制機制不適合,導致短訊息尾延遲高。Homa 是全新設計的協議,透過訊息本位、接收端延遲控制與交換機優先佇列,大幅降低短訊息延遲並提升整體效能。

John Ousterhout presents Homa, a new network protocol designed to drastically reduce tail latency for short messages in AI inference and agentic workloads by using message-based transport and receiver-side congestion control.

摘要、重點與章節標題由語言模型整理,細節(誰說的、數字、先後)可能有誤;要引用請以原始內容為準。

重點

  • AI 推理與代理工作負載中,短訊息的尾延遲會限制整體系統吞吐量。
  • TCP 與 RDMA 因無訊息邊界且延遲控制位於傳送端,無法有效處理短訊息。
  • Homa 採用訊息本位、接收端控制與優先佇列,將短訊息延遲降低至微秒級。

章節

依話題轉折切分,標題由 AI 產生

  1. 00:00Why latency is becoming the metric that matters
  2. 02:19The old workload: gigabytes and throughput
  3. 03:09The new workload: metadata and coordination
  4. 04:24How one slow exchange stalls every GPU
  5. 06:07Incast, and where the queue actually builds
  6. 07:10Why congestion control lives on the wrong end
  7. 10:05A byte stream has no message boundaries
  8. 11:08Homa, and a clean slate redesign
  9. 12:12Messages, not streams
  10. 13:30Controlling congestion from the receiver
  11. 14:59Using the priority queues already in the switch
  12. 15:52The benchmark against TCP

提到的工具與公司

  • Homa
  • TCP
  • RDMA
  • SRPT

適合誰看

負責 AI 系統架構、網路最佳化或正在使用大規模 GPU 叢集的開發人員。

摘要依據

講者
John Ousterhout
依據
自動字幕

為什麼排在這裡

人氣
0.92
新鮮
0.94

在主題頁與搜尋結果裡,名次由相關、人氣、新鮮三個分數決定;這一頁沒有搜尋的關鍵字,所以沒有相關分數。排序怎麼算

摘要由 AI 根據原文產生,可能有誤;完整內容請看原站。看影片(在新分頁開啟原站)