<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
<channel>
  <title>AI 武林：H100</title>
  <link>https://aiwulin.itsmygo.uk/tools/h100/</link>
  <description>AI 武林收錄的內容裡，最新提到 H100 的 30 筆，每筆附中文摘要。</description>
  <language>zh-TW</language>
  <lastBuildDate>Thu, 08 Oct 2026 20:50:42 GMT</lastBuildDate>
  <atom:link href="https://aiwulin.itsmygo.uk/tools/h100/rss.xml" rel="self" type="application/rss+xml"/>
  <item>
    <title>【曲博股市科技EP308】資料中心的未來在於太空：Starcloud公司網站產品大解密</title>
    <link>https://aiwulin.itsmygo.uk/c/7494cdbcab/</link>
    <guid isPermaLink="false">aiwulin-7494cdbcab</guid>
    <pubDate>Tue, 06 Oct 2026 00:00:00 +0800</pubDate>
    <dc:creator>曲博科技教室 Dr. J Class</dc:creator>
    <description>&lt;p&gt;介紹 Starcloud 公司的太空資料中心產品，解析其如何利用軌道設施部署 NVIDIA H100 GPU 以推動人工智慧革命。內容涵蓋從設計原則到星雲系列產品的具體規格與應用前景。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=tCf3B-pLP5M&quot;&gt;https://www.youtube.com/watch?v=tCf3B-pLP5M&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>AI 晶片與硬體</category>
  </item>
  <item>
    <title>Accelerating No-Code Enterprise AI: Inside Simplismart’s NVIDIA Inception Journey</title>
    <link>https://aiwulin.itsmygo.uk/c/d3094a82e3/</link>
    <guid isPermaLink="false">aiwulin-d3094a82e3</guid>
    <pubDate>Sat, 03 Oct 2026 00:00:00 +0800</pubDate>
    <dc:creator>NVIDIA</dc:creator>
    <description>&lt;p&gt;Simplismart 透過無程式碼平台讓企業快速部署自訂 AI 模型，專注於最佳化生成式 AI 的推理效能。內容介紹如何利用 NVIDIA GPU、TensorRT-LLM 及 NIM 微服務，針對不同場景（如語音代理、檔案解析）調整延…&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=UxOQ2qyY6Qo&quot;&gt;https://www.youtube.com/watch?v=UxOQ2qyY6Qo&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>推論與部署</category>
    <category>AI 晶片與硬體</category>
  </item>
  <item>
    <title>In-Context Retrieval with Siddharth Gollapudi - Weaviate Podcast #146!</title>
    <link>https://aiwulin.itsmygo.uk/c/a18cc40cee/</link>
    <guid isPermaLink="false">aiwulin-a18cc40cee</guid>
    <pubDate>Thu, 01 Oct 2026 00:00:00 +0800</pubDate>
    <dc:creator>Weaviate Podcast</dc:creator>
    <description>&lt;p&gt;探討將整個文庫置入大型語言模型上下文，利用注意力機制進行檢索的「情境檢索」技術。與需重新訓練的參數化檢索不同，此方法能更精確地還原文字，並解決長上下文中的注意力稀釋問題。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://podcasters.spotify.com/pod/show/weaviate/episodes/In-Context-Retrieval-with-Siddharth-Gollapudi---Weaviate-Podcast-146-e3pn04l&quot;&gt;https://podcasters.spotify.com/pod/show/weaviate/episodes/In-Context-Retrieval-with-Siddharth-Gollapudi---Weaviate-Podcast-146-e3pn04l&lt;/a&gt;&lt;/p&gt;</description>
    <category>Podcast</category>
    <category>Transformer 原理</category>
    <category>大型語言模型入門</category>
    <category>向量搜尋與 Embedding</category>
  </item>
  <item>
    <title>What exactly happens when you press send: Journey from Chat to a GPU</title>
    <link>https://aiwulin.itsmygo.uk/c/cae56ef28a/</link>
    <guid isPermaLink="false">aiwulin-cae56ef28a</guid>
    <pubDate>Wed, 30 Sep 2026 00:00:00 +0800</pubDate>
    <dc:creator>Vizuara</dc:creator>
    <description>&lt;p&gt;追蹤從手機按下傳送鍵到資料中心 GPU 的完整過程，解析大型語言模型如何透過矩陣運算與並行處理完成回答。觀眾能了解 Qwen3-8B 模型在 NVIDIA H100 上的實際運作機制與效能資料。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=sjLFTUp-1T0&quot;&gt;https://www.youtube.com/watch?v=sjLFTUp-1T0&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>AI 晶片與硬體</category>
    <category>大型語言模型入門</category>
  </item>
  <item>
    <title>Inside Rubin: NVIDIA's most powerful AI chip</title>
    <link>https://aiwulin.itsmygo.uk/c/e62166908e/</link>
    <guid isPermaLink="false">aiwulin-e62166908e</guid>
    <pubDate>Tue, 29 Sep 2026 00:00:00 +0800</pubDate>
    <dc:creator>Vizuara</dc:creator>
    <description>&lt;p&gt;拆解 NVIDIA Rubin 晶片內部結構，從電路層級分析其運算核心與記憶體設計，並對比 H100 探討限制 AI 速度的關鍵因素。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=EfYWEXnfXA8&quot;&gt;https://www.youtube.com/watch?v=EfYWEXnfXA8&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>AI 晶片與硬體</category>
  </item>
  <item>
    <title>How far behind Nvidia is Huawei?</title>
    <link>https://aiwulin.itsmygo.uk/c/910ee61b95/</link>
    <guid isPermaLink="false">aiwulin-910ee61b95</guid>
    <pubDate>Fri, 25 Sep 2026 00:00:00 +0800</pubDate>
    <dc:creator>Epoch AI（Gradient Updates）</dc:creator>
    <description>&lt;p&gt;分析華為在 AI 晶片與算力上的發展路徑，指出受出口管制限制，其晶片效能與產量難以在 2030 年前追上輝達。即使達成所有技術目標，華為仍將落後約四年，主要瓶頸在於製程裝置、記憶體供應及軟體生態。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://epochai.substack.com/p/how-far-behind-nvidia-is-huawei&quot;&gt;https://epochai.substack.com/p/how-far-behind-nvidia-is-huawei&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>AI 晶片與硬體</category>
  </item>
  <item>
    <title>#151 AI蛋糕的第三层｜Neoclouds新云的诞生</title>
    <link>https://aiwulin.itsmygo.uk/c/44f19a4410/</link>
    <guid isPermaLink="false">aiwulin-44f19a4410</guid>
    <pubDate>Mon, 21 Sep 2026 00:00:00 +0800</pubDate>
    <dc:creator>科技慢半拍</dc:creator>
    <description>&lt;p&gt;探討 AI 時代如何從「搜尋」轉向「生成」，導致基礎設施投入激增，催生以 CoreWeave 等為代表的「新雲」產業。內容解析新雲如何透過整合電力、冷卻與 GPU 叢集，將算力轉化為可交易資產，並分析其與傳統雲及輝達的生態關係。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.xiaoyuzhoufm.com/episode/6aa365a29d32647781675d04&quot;&gt;https://www.xiaoyuzhoufm.com/episode/6aa365a29d32647781675d04&lt;/a&gt;&lt;/p&gt;</description>
    <category>Podcast</category>
    <category>AI 晶片與硬體</category>
  </item>
  <item>
    <title>Two techniques for working with System One models</title>
    <link>https://aiwulin.itsmygo.uk/c/aa617abcb5/</link>
    <guid isPermaLink="false">aiwulin-aa617abcb5</guid>
    <pubDate>Fri, 18 Sep 2026 00:00:00 +0800</pubDate>
    <dc:creator>Sean Goedecke</dc:creator>
    <description>&lt;p&gt;介紹將大型語言模型轉化為只輸出決策的「System One」模型技術，並透過實作 Qwen3-8B 玩 Doom 與 Wikiracing 遊戲，分享設定層級目標與賽局抽樣等實作技巧。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://seangoedecke.com/two-techniques-for-working-with-system-one-models&quot;&gt;https://seangoedecke.com/two-techniques-for-working-with-system-one-models&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>大型語言模型入門</category>
    <category>本機跑模型</category>
  </item>
  <item>
    <title>#149 AI蛋糕的最底层｜重塑电力基础设施</title>
    <link>https://aiwulin.itsmygo.uk/c/0dd69d38a3/</link>
    <guid isPermaLink="false">aiwulin-0dd69d38a3</guid>
    <pubDate>Mon, 07 Sep 2026 00:00:00 +0800</pubDate>
    <dc:creator>科技慢半拍</dc:creator>
    <description>&lt;p&gt;探討電力如何成為 AI 發展的瓶頸，從黃仁勳的「五層蛋糕」模型分析能源競爭。內容涵蓋從傳統交流電到 800V 高壓直流電的架構演進，以及垂直供電與晶片背面供電等技術革新，旨在說明縮短電力路徑、降低損耗的重要性。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.xiaoyuzhoufm.com/episode/6a94e541f03e74ee6b026129&quot;&gt;https://www.xiaoyuzhoufm.com/episode/6a94e541f03e74ee6b026129&lt;/a&gt;&lt;/p&gt;</description>
    <category>Podcast</category>
    <category>AI 晶片與硬體</category>
  </item>
  <item>
    <title>20VC: How to Build Your Own Data Center &amp; Why Every Startup Should Do It | How ElevenLabs Leapfrogged Us: What I Learned | The AI Talent War: How Your Hiring Process Needs to Change with Cliff Weitzman, Speechify</title>
    <link>https://aiwulin.itsmygo.uk/c/5ee744f249/</link>
    <guid isPermaLink="false">aiwulin-5ee744f249</guid>
    <pubDate>Sat, 05 Sep 2026 00:00:00 +0800</pubDate>
    <dc:creator>The Twenty Minute VC（20VC）</dc:creator>
    <description>&lt;p&gt;Cliff Weitzman 分享 Speechify 自購 Nvidia GPU 建立資料中心的策略，解釋為何自建比租機更划算且能加速模型訓練。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://thetwentyminutevc.libsyn.com/20vc-how-to-build-your-own-data-center-why-every-startup-should-do-it-how-elevenlabs-leapfrogged-us-what-i-learned-the-ai-talent-war-how-your-hiring-process-needs-to-change-with-cliff-weitzman-speechify&quot;&gt;https://thetwentyminutevc.libsyn.com/20vc-how-to-build-your-own-data-center-why-every-startup-should-do-it-how-elevenlabs-leapfrogged-us-what-i-learned-the-ai-talent-war-how-your-hiring-process-needs-to-change-with-cliff-weitzman-speechify&lt;/a&gt;&lt;/p&gt;</description>
    <category>Podcast</category>
    <category>深度學習</category>
  </item>
  <item>
    <title>Up to 3.2x Faster Inference with LFM2.5-DSpark</title>
    <link>https://aiwulin.itsmygo.uk/c/c7cdb6722f/</link>
    <guid isPermaLink="false">aiwulin-c7cdb6722f</guid>
    <pubDate>Fri, 21 Aug 2026 00:00:00 +0800</pubDate>
    <dc:creator>Hugging Face Blog</dc:creator>
    <description>&lt;p&gt;介紹 LFM2.5-DSpark 技術，透過結合 DFlash 架構與輕量 Markov 頭部，大幅降低大模型推論延遲。文章提供在 H100 GPU 與 M4 Max 裝置上的具體效能資料，並說明如何在 SGLang 與 llama.…&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://huggingface.co/blog/LiquidAI/lfm25-dspark&quot;&gt;https://huggingface.co/blog/LiquidAI/lfm25-dspark&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>推論與部署</category>
  </item>
  <item>
    <title>IsoExec: Unified Execution to Eliminate Trainer-Inference Mismatch in SkyRL</title>
    <link>https://aiwulin.itsmygo.uk/c/f3884b8465/</link>
    <guid isPermaLink="false">aiwulin-f3884b8465</guid>
    <pubDate>Fri, 21 Aug 2026 00:00:00 +0800</pubDate>
    <dc:creator>vLLM Blog</dc:creator>
    <description>&lt;p&gt;介紹 IsoExec，一種統一執行抽象，用於消除訓練與推論引擎間的浮點數不匹配問題。透過執行合約與位元一致的核心，確保 RL 訓練與推論使用相同策略時結果一致，降低除錯成本並提升系統穩定性。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://vllm.ai/blog/2026-08-21-isoexec&quot;&gt;https://vllm.ai/blog/2026-08-21-isoexec&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>推論與部署</category>
  </item>
  <item>
    <title>The Bloomberg Terminal for AI Compute</title>
    <link>https://aiwulin.itsmygo.uk/c/7b4da98209/</link>
    <guid isPermaLink="false">aiwulin-7b4da98209</guid>
    <pubDate>Thu, 13 Aug 2026 00:00:00 +0800</pubDate>
    <dc:creator>The Data Exchange with Ben Lorica</dc:creator>
    <description>&lt;p&gt;Steve Hou 與 Ben Lorica 探討 AI 運算市場資料，包括 GPU 租賃指數、前向曲線及 Token 支出趨勢。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://dts.podtrac.com/redirect.mp3/www.buzzsprout.com/682433/episodes/19604199-the-bloomberg-terminal-for-ai-compute.mp3&quot;&gt;https://dts.podtrac.com/redirect.mp3/www.buzzsprout.com/682433/episodes/19604199-the-bloomberg-terminal-for-ai-compute.mp3&lt;/a&gt;&lt;/p&gt;</description>
    <category>Podcast</category>
    <category>AI 晶片與硬體</category>
  </item>
  <item>
    <title>Building the First Data Centers in Space</title>
    <link>https://aiwulin.itsmygo.uk/c/36af638b2f/</link>
    <guid isPermaLink="false">aiwulin-36af638b2f</guid>
    <pubDate>Thu, 06 Aug 2026 00:00:00 +0800</pubDate>
    <dc:creator>Y Combinator Startup Podcast</dc:creator>
    <description>&lt;p&gt;Starcloud 執行長 Philip Johnston 分享公司如何先預訂 SpaceX 火箭再開發產品，並成功將 Nvidia H100 晶片送入太空執行。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://podcasters.spotify.com/pod/show/ycombinator/episodes/Building-the-First-Data-Centers-in-Space-e3n13ek&quot;&gt;https://podcasters.spotify.com/pod/show/ycombinator/episodes/Building-the-First-Data-Centers-in-Space-e3n13ek&lt;/a&gt;&lt;/p&gt;</description>
    <category>Podcast</category>
    <category>AI 晶片與硬體</category>
  </item>
  <item>
    <title>The Inference Engineering Masterclass — Philip Kiely &amp; Ali Taha, Baseten</title>
    <link>https://aiwulin.itsmygo.uk/c/a8f46de4b5/</link>
    <guid isPermaLink="false">aiwulin-a8f46de4b5</guid>
    <pubDate>Tue, 04 Aug 2026 00:00:00 +0800</pubDate>
    <dc:creator>Latent Space</dc:creator>
    <description>&lt;p&gt;深入探討推論工程如何將訓練好的模型轉化為高效、可靠且具規模的產品。內容涵蓋長問答處理中的 KV 快取路由、預解碼加速、量化最佳化、工具呼叫輸出約束等技術細節，並討論了持續學習與 KV 快取壓縮的未來方向。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.latent.space/p/inference-eng&quot;&gt;https://www.latent.space/p/inference-eng&lt;/a&gt;&lt;/p&gt;</description>
    <category>Podcast</category>
    <category>推論與部署</category>
  </item>
  <item>
    <title>Inside the Model Factory — Eiso Kant, Poolside AI</title>
    <link>https://aiwulin.itsmygo.uk/c/332d880765/</link>
    <guid isPermaLink="false">aiwulin-332d880765</guid>
    <pubDate>Thu, 23 Jul 2026 00:00:00 +0800</pubDate>
    <dc:creator>Latent Space</dc:creator>
    <description>&lt;p&gt;訪談 Eiso Kant 分享 Poolside AI 如何建立 Model Factory，在八週內將模型從預訓練到發布。內容涵蓋其十年研發歷程、為何轉向開放權重、Laguna S 2.1 的效能表現及團隊運作模式。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.latent.space/p/poolside&quot;&gt;https://www.latent.space/p/poolside&lt;/a&gt;&lt;/p&gt;</description>
    <category>Podcast</category>
    <category>開源模型</category>
    <category>從零打造語言模型</category>
  </item>
  <item>
    <title>Expenditure Horizon: Measuring Optimization Ability, with an Application to NanoGPT</title>
    <link>https://aiwulin.itsmygo.uk/c/ca678a16dd/</link>
    <guid isPermaLink="false">aiwulin-ca678a16dd</guid>
    <pubDate>Tue, 21 Jul 2026 00:00:00 +0800</pubDate>
    <dc:creator>METR</dc:creator>
    <description>&lt;p&gt;提出「支出地平線」指標，用於量化 AI 代理在最佳化任務上的能力，並透過 NanoGPT 速度挑戰的資料進行實證分析。文章估算人類最佳化每提升 1% 需約 2,500 美元人力成本，並比較 AI 代理的投入曲線與人類曲線交點。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://metr.org/blog/2026-07-21-expenditure-horizon&quot;&gt;https://metr.org/blog/2026-07-21-expenditure-horizon&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>從零打造語言模型</category>
    <category>AI Agent 基礎</category>
    <category>AI 工作術</category>
  </item>
  <item>
    <title>Devin Outposts on Modal</title>
    <link>https://aiwulin.itsmygo.uk/c/7f0a59b170/</link>
    <guid isPermaLink="false">aiwulin-7f0a59b170</guid>
    <pubDate>Tue, 21 Jul 2026 00:00:00 +0800</pubDate>
    <dc:creator>Modal Blog</dc:creator>
    <description>&lt;p&gt;介紹 Devin 與 Modal 合作推出的 Devin Outposts，讓 AI 工程師能在使用者自訂的 GPU 環境中執行程式碼。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://modal.com/blog/devin-outposts-run-devin-in-modal-sandoxes&quot;&gt;https://modal.com/blog/devin-outposts-run-devin-in-modal-sandoxes&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>Coding Agent</category>
    <category>AI 輔助軟體工程</category>
    <category>AI 晶片與硬體</category>
  </item>
  <item>
    <title>Building custom code models for Ericsson proprietary silicon | AI Now Summit 2026</title>
    <link>https://aiwulin.itsmygo.uk/c/3d08c15e9f/</link>
    <guid isPermaLink="false">aiwulin-3d08c15e9f</guid>
    <pubDate>Thu, 02 Jul 2026 00:00:00 +0800</pubDate>
    <dc:creator>Mistral</dc:creator>
    <description>&lt;p&gt;介紹 Ericsson 與 Mistral 合作開發專為其 ASIC 晶片量身打造的自訂大型語言模型，解決資料稀缺問題並達成高準確度。看完能了解如何將模型導入生產環境以實現自主訓練。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=ArWG4pmTXPQ&quot;&gt;https://www.youtube.com/watch?v=ArWG4pmTXPQ&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>大型語言模型入門</category>
    <category>AI 晶片與硬體</category>
    <category>開源模型</category>
  </item>
  <item>
    <title>🔬 The Coolest Diffusion Research Isn't in LLMs — Evan Feinberg &amp; Sergey Edunov, Genesis Molecular AI</title>
    <link>https://aiwulin.itsmygo.uk/c/455620d3c8/</link>
    <guid isPermaLink="false">aiwulin-455620d3c8</guid>
    <pubDate>Wed, 01 Jul 2026 00:00:00 +0800</pubDate>
    <dc:creator>Latent Space</dc:creator>
    <description>&lt;p&gt;訪談 Genesis Molecular AI 創辦人 Evan Feinberg 與 CTO Sergey Edunov，探討他們如何透過專注於小分子藥物發現中的蛋白質 - 配體相互作用，利用擴散模型達到亞埃級精度。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.latent.space/p/the-coolest-diffusion-research-isnt&quot;&gt;https://www.latent.space/p/the-coolest-diffusion-research-isnt&lt;/a&gt;&lt;/p&gt;</description>
    <category>Podcast</category>
    <category>大型語言模型入門</category>
    <category>AI 與科學研究</category>
    <category>圖像生成</category>
  </item>
  <item>
    <title>Most of the Web Will Never Get APIs for AI Agents | Dhruv Batra, Yutori</title>
    <link>https://aiwulin.itsmygo.uk/c/9610a7e17d/</link>
    <guid isPermaLink="false">aiwulin-9610a7e17d</guid>
    <pubDate>Thu, 18 Jun 2026 00:00:00 +0800</pubDate>
    <dc:creator>Chain of Thought</dc:creator>
    <description>&lt;p&gt;訪談 Dhruv Batra 與 Yutori 公司，探討為何多數網頁不會為 AI 代理提供 API，並介紹 Yutori Navigator 如何透過感知畫素與即時編寫 JavaScript 來執行任務。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://chainofthought.show/podcast/63-most-of-the-web-will-never-get-apis-for-ai-agents-dhruv-batra&quot;&gt;https://chainofthought.show/podcast/63-most-of-the-web-will-never-get-apis-for-ai-agents-dhruv-batra&lt;/a&gt;&lt;/p&gt;</description>
    <category>Podcast</category>
    <category>AI Agent 基礎</category>
    <category>瀏覽器與電腦操作</category>
    <category>Coding Agent</category>
  </item>
  <item>
    <title>Stanford CS336 Language Modeling from Scratch | Spring 2026 | Guest Lecture: Dan Fu</title>
    <link>https://aiwulin.itsmygo.uk/c/71d9953c8f/</link>
    <guid isPermaLink="false">aiwulin-71d9953c8f</guid>
    <pubDate>Sat, 06 Jun 2026 00:00:00 +0800</pubDate>
    <dc:creator>Stanford Online</dc:creator>
    <description>&lt;p&gt;Dan Fu 演講分享語言模型從訓練到推論的完整流程，重點在於如何將 GPU 轉化為實際的「智慧引擎」。內容涵蓋推論服務架構、最佳化 KV Cache 的策略，以及透過 Mega Kernels 技術減少 GPU 等待時間以提升效率。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=9EEm4iMAF5s&quot;&gt;https://www.youtube.com/watch?v=9EEm4iMAF5s&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>推論與部署</category>
    <category>從零打造語言模型</category>
    <category>大型語言模型入門</category>
  </item>
  <item>
    <title>Stanford CS25: Transformers United V6 I Serving Transformers: Lessons from the Trenches</title>
    <link>https://aiwulin.itsmygo.uk/c/cfe75cca7e/</link>
    <guid isPermaLink="false">aiwulin-cfe75cca7e</guid>
    <pubDate>Fri, 05 Jun 2026 00:00:00 +0800</pubDate>
    <dc:creator>Stanford Online</dc:creator>
    <description>&lt;p&gt;Charles Frye 分享在生產環境部署 Transformer 模型的實戰經驗，涵蓋從應用架構到硬體選型的完整流程。內容包含如何最佳化推理效能、選擇適合的 GPU 硬體以及使用工具進行效能分析與除錯，幫助工程師解決大規模推理挑戰。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=ZUdIsRZhWXI&quot;&gt;https://www.youtube.com/watch?v=ZUdIsRZhWXI&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>Transformer 原理</category>
    <category>推論與部署</category>
    <category>AI 晶片與硬體</category>
  </item>
  <item>
    <title>Introducing Claude Managed Agents with Modal Sandboxes</title>
    <link>https://aiwulin.itsmygo.uk/c/a6e5735bb5/</link>
    <guid isPermaLink="false">aiwulin-a6e5735bb5</guid>
    <pubDate>Tue, 19 May 2026 00:00:00 +0800</pubDate>
    <dc:creator>Modal Blog</dc:creator>
    <description>&lt;p&gt;Modal 與 Anthropic 合作推出整合，讓 Claude Managed Agents 能在自託管的 Modal Sandbox 中執行工具呼叫。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://modal.com/blog/introducing-claude-managed-agents-with-modal-sandboxes&quot;&gt;https://modal.com/blog/introducing-claude-managed-agents-with-modal-sandboxes&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>Agent Harness</category>
    <category>AI Agent 基礎</category>
  </item>
  <item>
    <title>Boosting multimodal inference performance by &gt;10% with a single Python dictionary</title>
    <link>https://aiwulin.itsmygo.uk/c/f551621157/</link>
    <guid isPermaLink="false">aiwulin-f551621157</guid>
    <pubDate>Mon, 04 May 2026 00:00:00 +0800</pubDate>
    <dc:creator>Modal Blog</dc:creator>
    <description>&lt;p&gt;分析 SGLang 在處理多模態模型時，因重複執行 GPU 記憶體共享書寫作業導致效能瓶頸。透過簡單的 Python 字典快取機制替代昂貴的重複呼叫，使吞吐量提升 16% 且延遲降低 10%。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://modal.com/blog/boosting-multimodal-inference-performance-by-greater-than-10-with-a-single-python-dictionary&quot;&gt;https://modal.com/blog/boosting-multimodal-inference-performance-by-greater-than-10-with-a-single-python-dictionary&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>電腦視覺與多模態模型</category>
    <category>推論與部署</category>
  </item>
  <item>
    <title>Evidence on AI R&amp;D Progress from NanoGPT</title>
    <link>https://aiwulin.itsmygo.uk/c/6af4a563a9/</link>
    <guid isPermaLink="false">aiwulin-6af4a563a9</guid>
    <pubDate>Tue, 21 Apr 2026 00:00:00 +0800</pubDate>
    <dc:creator>METR</dc:creator>
    <description>&lt;p&gt;透過 NanoGPT 速度挑戰賽的公開資料，分析 AI 代理在加速 AI 研發方面的實際貢獻與進展。作者分類了人類與 AI 提出的最佳化方案，指出雖然 AI 已參與部分最佳化，但深度創新仍多來自人類，並探討了如何利用此類挑戰作為評估基準。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://metr.org/notes/2026-04-21-ai-rd-nanogpt-progress&quot;&gt;https://metr.org/notes/2026-04-21-ai-rd-nanogpt-progress&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>從零打造語言模型</category>
    <category>AI Agent 基礎</category>
  </item>
  <item>
    <title>Autoscaling Autoresearch: Give your agents elastic GPUs on Modal</title>
    <link>https://aiwulin.itsmygo.uk/c/fa55094f14/</link>
    <guid isPermaLink="false">aiwulin-fa55094f14</guid>
    <pubDate>Tue, 14 Apr 2026 00:00:00 +0800</pubDate>
    <dc:creator>Modal Blog</dc:creator>
    <description>&lt;p&gt;介紹如何使用 Modal 平台配合 Autoresearch 工具，讓 AI 代理自動彈性調配 GPU 資源進行研究。透過簡單的 API 呼叫，代理能根據任務需求即時啟動或釋放多張 GPU，既避免資源浪費，又大幅提升實驗速度與效率。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://modal.com/blog/autoscaling-autoresearch&quot;&gt;https://modal.com/blog/autoscaling-autoresearch&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>AI Agent 基礎</category>
    <category>AI 晶片與硬體</category>
  </item>
  <item>
    <title>Greetings, Earthlings: Philip Johnston of Starcloud on Data Centers in Space</title>
    <link>https://aiwulin.itsmygo.uk/c/8e5ee7e57b/</link>
    <guid isPermaLink="false">aiwulin-8e5ee7e57b</guid>
    <pubDate>Tue, 17 Mar 2026 00:00:00 +0800</pubDate>
    <dc:creator>Training Data</dc:creator>
    <description>&lt;p&gt;Philip Johnston 解釋為何太空資料中心將成為未來十年 AI 運算的主要地點，分析地緣限制與太空能源成本優勢。內容涵蓋熱耗散物理原理、射線防護測試、以及太空運算如何重塑雲端架構與經濟模型。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://pscrb.fm/rss/p/traffic.megaphone.fm/CPUAI3959632960.mp3&quot;&gt;https://pscrb.fm/rss/p/traffic.megaphone.fm/CPUAI3959632960.mp3&lt;/a&gt;&lt;/p&gt;</description>
    <category>Podcast</category>
    <category>AI 晶片與硬體</category>
    <category>推論與部署</category>
  </item>
  <item>
    <title>Dylan Patel — Deep dive on the 3 big bottlenecks to scaling AI compute</title>
    <link>https://aiwulin.itsmygo.uk/c/a243926cf6/</link>
    <guid isPermaLink="false">aiwulin-a243926cf6</guid>
    <pubDate>Sat, 14 Mar 2026 00:00:00 +0800</pubDate>
    <dc:creator>Dwarkesh Podcast</dc:creator>
    <description>&lt;p&gt;Dylan Patel 深入剖析 AI 運算擴展的三大瓶頸：邏輯、記憶體與電力，並探討實驗室、超大型雲端廠商與晶圓廠的經濟動態。內容涵蓋 H100 價格走勢、記憶體供需壓力、機器人集中化運算趨勢，以及台灣在地化供應鏈風險。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.dwarkesh.com/p/dylan-patel&quot;&gt;https://www.dwarkesh.com/p/dylan-patel&lt;/a&gt;&lt;/p&gt;</description>
    <category>Podcast</category>
    <category>AI 晶片與硬體</category>
  </item>
  <item>
    <title>How is hardware reshaping LLM design?</title>
    <link>https://aiwulin.itsmygo.uk/c/a8d7264f5e/</link>
    <guid isPermaLink="false">aiwulin-a8d7264f5e</guid>
    <pubDate>Tue, 03 Mar 2026 00:00:00 +0800</pubDate>
    <dc:creator>Julia Turc</dc:creator>
    <description>&lt;p&gt;解析為何傳統大型語言模型推理受記憶體牆限制，並探討 HBM、屋頂線模型及 vLLM 等最佳化技術。觀眾可了解記憶體與運算密集型任務的差異，以及擴散式模型如何改變工作負載。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=BSzhrZOp2x8&quot;&gt;https://www.youtube.com/watch?v=BSzhrZOp2x8&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>大型語言模型入門</category>
    <category>推論與部署</category>
  </item>
</channel>
</rss>
