看影片(在新分頁開啟原站)連到 AI Engineer
摘要
Daniel Svonava 分享如何讓小型開源模型在生產環境中發揮 frontier 級效能,透過任務分割與專用模型組合降低成本。他介紹 Superlinked 的自定義架構,利用共享佇列與自組批次解決傳統路由器效能不足的問題,並透過自動調優迴圈讓團隊無需頻繁溝通即可快速部署新模型。
A talk on deploying fleets of small open-source models efficiently using custom infrastructure to overcome routing bottlenecks and reduce costs.
摘要、重點與章節標題由語言模型整理,細節(誰說的、數字、先後)可能有誤;要引用請以原始內容為準。
重點
- 小型開源模型在特定任務上可達 frontier 級效能且成本更低。
- 傳統頂層路由器無法有效處理大量小型請求,導致 GPU 利用率低。
- Superlinked 架構透過共享佇列與自組批次,讓吞吐量翻倍並簡化部署。
章節
依話題轉折切分,標題由 AI 產生
- 00:00Small open source models, do it yourself
- 02:30Small models are catching the frontier
- 03:42One task, one model: a nine model contract agent
- 06:37Serving tools are do it yourself research projects
- 07:31Why top down routing chokes on small requests
- 08:42LoRAs, fine tunes, and the conversation that kills velocity
- 11:12A gateway, a shared queue, and workers that pull
- 14:37Workers form their own batches, and throughput doubles
- 16:28Three runtimes and a Rust sidecar
- 19:29Half a million embedding tokens a second on one GPU
- 22:38Pack models on the same GPU
- 23:33Autoresearch that ships tuned configs
- 24:17An 80 cent LoRA
提到的工具與公司
- Superlinked
- Apache 2.0
- vLLM
- SGLang
- PyTorch
- Kendall
- NATS JetStream
- Rust
適合誰看
正在構建或最佳化 AI 代理、需要部署多模型叢集的軟體工程師與基礎設施工程師。
摘要依據
- 講者
- Daniel Svonava
- 依據
- 自動字幕
為什麼排在這裡
- 人氣
- 0.72
- 新鮮
- 0.94
在主題頁與搜尋結果裡,名次由相關、人氣、新鮮三個分數決定;這一頁沒有搜尋的關鍵字,所以沒有相關分數。排序怎麼算
相關內容
- What Makes Open Models Fast in Production — Sujee Maniyam, Nebius影片 ・ AI Engineer(在新分頁開啟原站)
- Stanford MS&E435 Economics of the AI Supercycle | Spring 2026 | Applications, Applied AI影片 ・ Stanford Online ・ 49 分鐘
- Why Specialized AI Could Beat The God Model影片 ・ a16z(在新分頁開啟原站)
- How Prisma Built An AI SRE影片 ・ Mastra ・ 57 分鐘(在新分頁開啟原站)
- What comes next with open modelsPodcast ・ Interconnects ・ 18 分鐘
- Jev 爆紅後 Strands Decider、Clef 接連登場,決策模型正走出哪些不同路線?文章 ・ TechOrange 科技報橘
摘要由 AI 根據原文產生,可能有誤;完整內容請看原站。看影片(在新分頁開啟原站)
