影片高階EN1.7 萬 次觀看
The Frontier AI Inference Cloud for Agents — Byung-Gon (Gon) Chun, FriendliAI
來源 AI Engineer
看影片(在新分頁開啟原站)連到 AI Engineer
摘要
探討代理推理與聊天不同的架構需求,強調以任務完成時間為指標。透過開源模型與閉源模型比較,展示開源模型能以極低成本達成可用品質。FriendliAI 透過字首快取、分層 KV 快取管理、感知快取的路由及感知代理的排程,最佳化任務端到端延遲。
A talk on optimizing agent inference for end-to-end task completion using open-weight models and specialized caching infrastructure.
摘要、重點與章節標題由語言模型整理,細節(誰說的、數字、先後)可能有誤;要引用請以原始內容為準。
重點
- 代理任務應以完成時間衡量,而非單一請求延遲。
- 開源模型已達可用品質,能以極低成本執行代理。
- FriendliAI 透過四項技術最佳化代理推理的端到端速度。
章節
依話題轉折切分,標題由 AI 產生
- 00:00The team that invented continuous batching
- 02:06One task, two models, one bill
- 03:18The unit is the task, not the request
- 05:25Consecutive steps share a huge prefix
- 07:02Four pillars of an agentic inference cloud
- 09:19Cache aware routing versus a naive load balancer
- 10:04Agent aware optimization
- 11:00The same task, end to end, on two providers
提到的工具與公司
- GLM 5.2
- Claude Opus 4.8
- FriendliAI
- vLLM
- LG
- Hilo
- Orca
適合誰看
正在構建或部署 AI 代理、需要最佳化推理成本與延遲的開發者或工程師。
摘要依據
- 依據
- 自動字幕
為什麼排在這裡
- 人氣
- 0.77
- 新鮮
- 0.95
在主題頁與搜尋結果裡,名次由相關、人氣、新鮮三個分數決定;這一頁沒有搜尋的關鍵字,所以沒有相關分數。排序怎麼算
相關內容
- 實戰 AI Agents 應用開發: TTFT 和 Prompt Caching文章 ・ ihower(張文鈿)
- Fireworks CTO: Why AI Is About To Get 1000x Better影片 ・ David Ondrej ・ 51 分鐘(在新分頁開啟原站)
- GLM-5.2 is the step change for open agentsPodcast ・ Interconnects ・ 9 分鐘
- GLM 5.2 total victory: the week open source won and nobody panickedPodcast ・ ThursdAI ・ 1 小時 30 分
- The Age of Async Agents — Cognition's Walden Yan & OpenInspect's Cole MurrayPodcast ・ Latent Space ・ 1 小時 8 分
- 愛好 AI Engineer 電子報 🚀 新型態代理人 OpenClaw 正夯,電子報改版 #35文章 ・ ihower(張文鈿)
摘要由 AI 根據原文產生,可能有誤;完整內容請看原站。看影片(在新分頁開啟原站)
