跳到主要內容
AI 武林
影片進階EN9,995 次觀看

Large clusters for small models — Daniel Svonava, Superlinked

來源 AI Engineer

看影片(在新分頁開啟原站)連到 AI Engineer

其他版本:AIE Talks 摘要頁(在新分頁開啟)

摘要

Daniel Svonava 分享如何讓小型開源模型在生產環境中發揮 frontier 級效能,透過任務分割與專用模型組合降低成本。他介紹 Superlinked 的自定義架構,利用共享佇列與自組批次解決傳統路由器效能不足的問題,並透過自動調優迴圈讓團隊無需頻繁溝通即可快速部署新模型。

A talk on deploying fleets of small open-source models efficiently using custom infrastructure to overcome routing bottlenecks and reduce costs.

摘要、重點與章節標題由語言模型整理,細節(誰說的、數字、先後)可能有誤;要引用請以原始內容為準。

重點

  • 小型開源模型在特定任務上可達 frontier 級效能且成本更低。
  • 傳統頂層路由器無法有效處理大量小型請求,導致 GPU 利用率低。
  • Superlinked 架構透過共享佇列與自組批次,讓吞吐量翻倍並簡化部署。

章節

依話題轉折切分,標題由 AI 產生

  1. 00:00Small open source models, do it yourself
  2. 02:30Small models are catching the frontier
  3. 03:42One task, one model: a nine model contract agent
  4. 06:37Serving tools are do it yourself research projects
  5. 07:31Why top down routing chokes on small requests
  6. 08:42LoRAs, fine tunes, and the conversation that kills velocity
  7. 11:12A gateway, a shared queue, and workers that pull
  8. 14:37Workers form their own batches, and throughput doubles
  9. 16:28Three runtimes and a Rust sidecar
  10. 19:29Half a million embedding tokens a second on one GPU
  11. 22:38Pack models on the same GPU
  12. 23:33Autoresearch that ships tuned configs
  13. 24:17An 80 cent LoRA

提到的工具與公司

  • Superlinked
  • Apache 2.0
  • vLLM
  • SGLang
  • PyTorch
  • Kendall
  • NATS JetStream
  • Rust

適合誰看

正在構建或最佳化 AI 代理、需要部署多模型叢集的軟體工程師與基礎設施工程師。

摘要依據

講者
Daniel Svonava
依據
自動字幕

為什麼排在這裡

人氣
0.72
新鮮
0.94

在主題頁與搜尋結果裡,名次由相關、人氣、新鮮三個分數決定;這一頁沒有搜尋的關鍵字,所以沒有相關分數。排序怎麼算

摘要由 AI 根據原文產生,可能有誤;完整內容請看原站。看影片(在新分頁開啟原站)