<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
<channel>
  <title>AI 武林：SGLang</title>
  <link>https://aiwulin.itsmygo.uk/tools/sglang/</link>
  <description>AI 武林收錄的內容裡，最新提到 SGLang 的 19 筆，每筆附中文摘要。</description>
  <language>zh-TW</language>
  <lastBuildDate>Thu, 08 Oct 2026 17:22:41 GMT</lastBuildDate>
  <atom:link href="https://aiwulin.itsmygo.uk/tools/sglang/rss.xml" rel="self" type="application/rss+xml"/>
  <item>
    <title>What Is an Inference Engine, Anyway? — Charles Frye, Modal</title>
    <link>https://aiwulin.itsmygo.uk/c/9380785f9e/</link>
    <guid isPermaLink="false">aiwulin-9380785f9e</guid>
    <pubDate>Tue, 06 Oct 2026 00:00:00 +0800</pubDate>
    <dc:creator>AI Engineer</dc:creator>
    <description>&lt;p&gt;Charles Frye 解析推理引擎如何處理從 API 請求到模型輸出的完整流程，涵蓋分詞、排程與 GPU 執行等關鍵環節。影片說明不同工作負載（如聊天機器人、背景代理）對延遲與吞吐量的影響，並探討 KV 快取與推測解碼等最佳化技術。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=woIYJYd_etI&quot;&gt;https://www.youtube.com/watch?v=woIYJYd_etI&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>推論與部署</category>
  </item>
  <item>
    <title>Accelerating vision-language models with LFM2.5-VL-DSpark</title>
    <link>https://aiwulin.itsmygo.uk/c/c2a53f2332/</link>
    <guid isPermaLink="false">aiwulin-c2a53f2332</guid>
    <pubDate>Thu, 24 Sep 2026 00:00:00 +0800</pubDate>
    <dc:creator>Hugging Face Blog</dc:creator>
    <description>&lt;p&gt;介紹 LFM2.5-VL-DSpark 模型，透過預測機制加速視覺語言模型的推理速度。在邊緣裝置與 GPU 上分別實現超過 3 倍與 2.6 倍的解碼加速，同時僅增加約 9% 的參數。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://huggingface.co/blog/LiquidAI/lfm2-5-vl-dspark&quot;&gt;https://huggingface.co/blog/LiquidAI/lfm2-5-vl-dspark&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>電腦視覺與多模態模型</category>
    <category>推論與部署</category>
    <category>本機跑模型</category>
  </item>
  <item>
    <title>Haters think AI agents can't write GPU code? This'll ROCm</title>
    <link>https://aiwulin.itsmygo.uk/c/94dd70a7c5/</link>
    <guid isPermaLink="false">aiwulin-94dd70a7c5</guid>
    <pubDate>Tue, 22 Sep 2026 00:00:00 +0800</pubDate>
    <dc:creator>The Stack Overflow Podcast</dc:creator>
    <description>&lt;p&gt;AMD 軟體主管 Anush Elangovan 訪談中，介紹 ROCm 如何透過開放原始碼與 AI 代理技術，大幅降低 GPU 程式設計門檻。內容涵蓋 ROCm 架構、SIMT 平行計算原理，以及 AI 如何協助最佳化硬體程式碼與效能。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://rss.art19.com/episodes/e3dd73f3-c12a-4053-b0bf-4550eb47f58a.mp3?rss_browser=BAhJIg9BSVd1bGluQm90BjoGRVQ%3D--40c69dfd25591c59cb985eaa38245c7694c51649&quot;&gt;https://rss.art19.com/episodes/e3dd73f3-c12a-4053-b0bf-4550eb47f58a.mp3?rss_browser=BAhJIg9BSVd1bGluQm90BjoGRVQ%3D--40c69dfd25591c59cb985eaa38245c7694c51649&lt;/a&gt;&lt;/p&gt;</description>
    <category>Podcast</category>
    <category>AI 晶片與硬體</category>
    <category>AI Agent 基礎</category>
  </item>
  <item>
    <title>Large clusters for small models — Daniel Svonava, Superlinked</title>
    <link>https://aiwulin.itsmygo.uk/c/594efa1a7a/</link>
    <guid isPermaLink="false">aiwulin-594efa1a7a</guid>
    <pubDate>Sun, 20 Sep 2026 00:00:00 +0800</pubDate>
    <dc:creator>AI Engineer</dc:creator>
    <description>&lt;p&gt;Daniel Svonava 分享如何讓小型開源模型在生產環境中發揮 frontier 級效能，透過任務分割與專用模型組合降低成本。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=g4SsanB0gMc&quot;&gt;https://www.youtube.com/watch?v=g4SsanB0gMc&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>開源模型</category>
    <category>推論與部署</category>
  </item>
  <item>
    <title>Granite 4.2 LLMs: How They're Built</title>
    <link>https://aiwulin.itsmygo.uk/c/74043a2c13/</link>
    <guid isPermaLink="false">aiwulin-74043a2c13</guid>
    <pubDate>Tue, 25 Aug 2026 00:00:00 +0800</pubDate>
    <dc:creator>Hugging Face Blog</dc:creator>
    <description>&lt;p&gt;Granite 4.2 系列是 IBM 推出的三檔（3B、8B、30B）推理型大型語言模型，透過多階段強化學習與工具呼叫訓練，具備思考/不思考模式切換能力。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://huggingface.co/blog/ibm-granite/granite-4-2&quot;&gt;https://huggingface.co/blog/ibm-granite/granite-4-2&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>大型語言模型入門</category>
  </item>
  <item>
    <title>Up to 3.2x Faster Inference with LFM2.5-DSpark</title>
    <link>https://aiwulin.itsmygo.uk/c/c7cdb6722f/</link>
    <guid isPermaLink="false">aiwulin-c7cdb6722f</guid>
    <pubDate>Fri, 21 Aug 2026 00:00:00 +0800</pubDate>
    <dc:creator>Hugging Face Blog</dc:creator>
    <description>&lt;p&gt;介紹 LFM2.5-DSpark 技術，透過結合 DFlash 架構與輕量 Markov 頭部，大幅降低大模型推論延遲。文章提供在 H100 GPU 與 M4 Max 裝置上的具體效能資料，並說明如何在 SGLang 與 llama.…&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://huggingface.co/blog/LiquidAI/lfm25-dspark&quot;&gt;https://huggingface.co/blog/LiquidAI/lfm25-dspark&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>推論與部署</category>
  </item>
  <item>
    <title>Qwen3.8-27B &amp; How to Serve it Fast</title>
    <link>https://aiwulin.itsmygo.uk/c/245d28bf6e/</link>
    <guid isPermaLink="false">aiwulin-245d28bf6e</guid>
    <pubDate>Tue, 18 Aug 2026 00:00:00 +0800</pubDate>
    <dc:creator>Sam Witteveen</dc:creator>
    <description>&lt;p&gt;介紹 Qwen3.8-27B 模型的功能與效能，並示範如何使用 SGLang 等工具實現最高每秒 token 數的服務速度。觀眾可學習如何最佳化大型語言模型的部署與推理速度。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=PTuGGdDuyPI&quot;&gt;https://www.youtube.com/watch?v=PTuGGdDuyPI&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>推論與部署</category>
    <category>大型語言模型入門</category>
  </item>
  <item>
    <title>not much happened today</title>
    <link>https://aiwulin.itsmygo.uk/c/987042bbd3/</link>
    <guid isPermaLink="false">aiwulin-987042bbd3</guid>
    <pubDate>Fri, 14 Aug 2026 00:00:00 +0800</pubDate>
    <dc:creator>AINews（Latent Space／smol.ai）</dc:creator>
    <description>&lt;p&gt;彙整了 2026 年 8 月 AI 領域的最新動態，重點涵蓋 Z.ai 推出 GLM-5.3、Alibaba 發布 Qwen3.8-27B 以及 DeepSeek V4-Pro 的更新。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://news.smol.ai/issues/26-08-14-cursor-xai&quot;&gt;https://news.smol.ai/issues/26-08-14-cursor-xai&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>開源模型</category>
    <category>AI 編輯器與開發工具</category>
    <category>模型發布與實測</category>
  </item>
  <item>
    <title>E247｜对话盛颖：xAI，Infra的浪漫，SGLang，开源，平权与“甄嬛传”</title>
    <link>https://aiwulin.itsmygo.uk/c/e695a18ab3/</link>
    <guid isPermaLink="false">aiwulin-e695a18ab3</guid>
    <pubDate>Wed, 05 Aug 2026 00:00:00 +0800</pubDate>
    <dc:creator>硅谷101</dc:creator>
    <description>&lt;p&gt;訪談邀請盛穎分享其從數學研究轉向 AI 基礎設施的歷程，深入探討 SGLang 如何將學術概念轉化為生產級推理引擎。內容涵蓋 xAI 的經驗、RadixArk 的創立，以及對 AI 平權與開源的觀點。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://sv101.fireside.fm/260&quot;&gt;https://sv101.fireside.fm/260&lt;/a&gt;&lt;/p&gt;</description>
    <category>Podcast</category>
    <category>推論與部署</category>
    <category>大型語言模型入門</category>
    <category>職涯與學習路線</category>
  </item>
  <item>
    <title>177: 详解Kimi K3：强到冲击Anthropic估值的模型什么样？</title>
    <link>https://aiwulin.itsmygo.uk/c/69cbb244db/</link>
    <guid isPermaLink="false">aiwulin-69cbb244db</guid>
    <pubDate>Tue, 04 Aug 2026 00:00:00 +0800</pubDate>
    <dc:creator>晚点聊 LateTalk</dc:creator>
    <description>&lt;p&gt;訪談邀請 RadixArk 與華盛頓大學博士生，從架構與演算法雙線拆解 Kimi K3。內容涵蓋其混合注意力機制、推理加速原理及對投資市場的影響，並探討開源權重與閉源流水線的差異。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://podcast.latepost.com/177&quot;&gt;https://podcast.latepost.com/177&lt;/a&gt;&lt;/p&gt;</description>
    <category>Podcast</category>
    <category>開源模型</category>
    <category>Transformer 原理</category>
  </item>
  <item>
    <title>“榨”出硅的极限：怎么让GPU不“闲着”？｜与SGLang、RadixArk深聊AI Infra技术与千亿美元市场</title>
    <link>https://aiwulin.itsmygo.uk/c/b4a0984350/</link>
    <guid isPermaLink="false">aiwulin-b4a0984350</guid>
    <pubDate>Fri, 31 Jul 2026 00:00:00 +0800</pubDate>
    <dc:creator>硅谷101</dc:creator>
    <description>&lt;p&gt;探討如何透過 SGLang 與 RadixArk 的技術，最佳化 GPU 的計算、等待與資源分配，以榨出更多算力。觀眾能學習到 AI 基礎設施中提升效率的關鍵策略與架構原理。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=cB_X6AImPjQ&quot;&gt;https://www.youtube.com/watch?v=cB_X6AImPjQ&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>深度學習</category>
  </item>
  <item>
    <title>Local AI 201: Inference Engines, Hardware Stack</title>
    <link>https://aiwulin.itsmygo.uk/c/ab1d13b2c1/</link>
    <guid isPermaLink="false">aiwulin-ab1d13b2c1</guid>
    <pubDate>Wed, 22 Jul 2026 00:00:00 +0800</pubDate>
    <dc:creator>Hugging Face</dc:creator>
    <description>&lt;p&gt;介紹本地 AI 推理引擎與硬體架構，透過實測比較不同軟體與硬體組合的效能。觀眾可學習如何選擇適合的推理工具並最佳化本地部署。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=90k2HRT8Ito&quot;&gt;https://www.youtube.com/watch?v=90k2HRT8Ito&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>推論與部署</category>
    <category>本機跑模型</category>
  </item>
  <item>
    <title>New Model: Inkling by Thinking Machine on Hugging Face</title>
    <link>https://aiwulin.itsmygo.uk/c/93071d65e4/</link>
    <guid isPermaLink="false">aiwulin-93071d65e4</guid>
    <pubDate>Thu, 16 Jul 2026 00:00:00 +0800</pubDate>
    <dc:creator>Hugging Face</dc:creator>
    <description>&lt;p&gt;介紹 Thinking Machine 推出的 Inkling 模型，該模型擁有 10 兆參數，能原生理解圖片、文字與聲音，並具備 100 萬 token 的上下文視窗。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=EOK_haQSMIA&quot;&gt;https://www.youtube.com/watch?v=EOK_haQSMIA&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>開源模型</category>
  </item>
  <item>
    <title>Stanford CS25: Transformers United V6 I Serving Transformers: Lessons from the Trenches</title>
    <link>https://aiwulin.itsmygo.uk/c/cfe75cca7e/</link>
    <guid isPermaLink="false">aiwulin-cfe75cca7e</guid>
    <pubDate>Fri, 05 Jun 2026 00:00:00 +0800</pubDate>
    <dc:creator>Stanford Online</dc:creator>
    <description>&lt;p&gt;Charles Frye 分享在生產環境部署 Transformer 模型的實戰經驗，涵蓋從應用架構到硬體選型的完整流程。內容包含如何最佳化推理效能、選擇適合的 GPU 硬體以及使用工具進行效能分析與除錯，幫助工程師解決大規模推理挑戰。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=ZUdIsRZhWXI&quot;&gt;https://www.youtube.com/watch?v=ZUdIsRZhWXI&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>Transformer 原理</category>
    <category>推論與部署</category>
    <category>AI 晶片與硬體</category>
  </item>
  <item>
    <title>How to Engineer AI Inference Systems with Philip Kiely - #766</title>
    <link>https://aiwulin.itsmygo.uk/c/e6373909cc/</link>
    <guid isPermaLink="false">aiwulin-e6373909cc</guid>
    <pubDate>Fri, 01 May 2026 00:00:00 +0800</pubDate>
    <dc:creator>The TWIML AI Podcast</dc:creator>
    <description>&lt;p&gt;深入探討推論工程（Inference Engineering），解釋為何它是 AI 最關鍵且複雜的工作負載。內容涵蓋從研究到生產的快速週期、關鍵技術如量化與 KV Cache 重用的應用，以及業界從 API 呼叫到自建平台的演進路徑。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://twimlai.com/podcast/twimlai/how-engineer-ai-inference-systems&quot;&gt;https://twimlai.com/podcast/twimlai/how-engineer-ai-inference-systems&lt;/a&gt;&lt;/p&gt;</description>
    <category>Podcast</category>
    <category>推論與部署</category>
  </item>
  <item>
    <title>163: 详解DeepSeekV4：Infra巨鲸、百万上下文走进现实、极致效率优化</title>
    <link>https://aiwulin.itsmygo.uk/c/93aaddbdcc/</link>
    <guid isPermaLink="false">aiwulin-93aaddbdcc</guid>
    <pubDate>Thu, 30 Apr 2026 00:00:00 +0800</pubDate>
    <dc:creator>晚点聊 LateTalk</dc:creator>
    <description>&lt;p&gt;訪談由 RadixArk 工程師與 UCLA 博士生，深入解析 DeepSeek V4 技術細節。內容涵蓋放棄 MLA 架構改用混合稀疏注意力、Muon 最佳化器、mHC 及 FP4 等工程創新，探討如何讓百萬上下文從理論走向實用，並分析…&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://podcast.latepost.com/163&quot;&gt;https://podcast.latepost.com/163&lt;/a&gt;&lt;/p&gt;</description>
    <category>Podcast</category>
    <category>上下文工程</category>
    <category>開源模型</category>
  </item>
  <item>
    <title>The Race to Production-Grade Diffusion LLMs with Stefano Ermon - #764</title>
    <link>https://aiwulin.itsmygo.uk/c/d8d8c3192a/</link>
    <guid isPermaLink="false">aiwulin-d8d8c3192a</guid>
    <pubDate>Fri, 27 Mar 2026 00:00:00 +0800</pubDate>
    <dc:creator>The TWIML AI Podcast</dc:creator>
    <description>&lt;p&gt;訪談斯坦福教授 Stefano Ermon 與 Inception Labs CEO，深入探討將影像生成用的擴散模型應用於文字與程式碼生成的技術挑戰與突破。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://twimlai.com/podcast/twimlai/race-production-grade-diffusion-llms&quot;&gt;https://twimlai.com/podcast/twimlai/race-production-grade-diffusion-llms&lt;/a&gt;&lt;/p&gt;</description>
    <category>Podcast</category>
    <category>圖像生成</category>
    <category>大型語言模型入門</category>
  </item>
  <item>
    <title>NVIDIA's AI Engineers: Agent Inference at Planetary Scale and &quot;Speed of Light&quot; — Nader Khalil (Brev), Kyle Kranen (Dynamo)</title>
    <link>https://aiwulin.itsmygo.uk/c/35dfdbec2a/</link>
    <guid isPermaLink="false">aiwulin-35dfdbec2a</guid>
    <pubDate>Tue, 10 Mar 2026 00:00:00 +0800</pubDate>
    <dc:creator>Latent Space</dc:creator>
    <description>&lt;p&gt;訪談 NVIDIA 工程師 Nader Khalil 與 Kyle Kranen，探討 Brev 如何降低開發者使用高階 GPU 的門檻，以及 Dynamo 框架如何透過預填充與解碼分離技術，實現大規模推理加速與成本最佳化。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.latent.space/p/nvidia-brev-dynamo&quot;&gt;https://www.latent.space/p/nvidia-brev-dynamo&lt;/a&gt;&lt;/p&gt;</description>
    <category>Podcast</category>
    <category>推論與部署</category>
    <category>AI 晶片與硬體</category>
    <category>AI 輔助軟體工程</category>
  </item>
  <item>
    <title>The CEO Behind the Fastest-Growing AI Inference Company | Tuhin Srivastava</title>
    <link>https://aiwulin.itsmygo.uk/c/b813413d0c/</link>
    <guid isPermaLink="false">aiwulin-b813413d0c</guid>
    <pubDate>Tue, 18 Nov 2025 00:00:00 +0800</pubDate>
    <dc:creator>Gradient Dissent</dc:creator>
    <description>&lt;p&gt;Tuhin Srivastava 與 Lukas Biewald 探討 Baseten 如何從早期小模型服務轉型為大模型推理基礎設施。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://episodes.captivate.fm/episode/bb7a7e75-34e4-471e-a3b8-12dfff61e22e.mp3&quot;&gt;https://episodes.captivate.fm/episode/bb7a7e75-34e4-471e-a3b8-12dfff61e22e.mp3&lt;/a&gt;&lt;/p&gt;</description>
    <category>Podcast</category>
    <category>推論與部署</category>
  </item>
</channel>
</rss>
