<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
<channel>
  <title>AI 武林：GRPO</title>
  <link>https://aiwulin.itsmygo.uk/tools/grpo/</link>
  <description>AI 武林收錄的內容裡，最新提到 GRPO 的 16 筆，每筆附中文摘要。</description>
  <language>zh-TW</language>
  <lastBuildDate>Thu, 08 Oct 2026 17:22:41 GMT</lastBuildDate>
  <atom:link href="https://aiwulin.itsmygo.uk/tools/grpo/rss.xml" rel="self" type="application/rss+xml"/>
  <item>
    <title>Build &amp; Train a GLM-5.3-Flash Model From Scratch with Python</title>
    <link>https://aiwulin.itsmygo.uk/c/f6477d5af6/</link>
    <guid isPermaLink="false">aiwulin-f6477d5af6</guid>
    <pubDate>Wed, 07 Oct 2026 00:00:00 +0800</pubDate>
    <dc:creator>freeCodeCamp.org</dc:creator>
    <description>&lt;p&gt;示範如何從零開始用 Python 訓練一個 2500 萬參數的 GLM-5.3 Flash 多模態語言模型，涵蓋分詞、預訓練與強化學習流程。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=-gfgQfw2g_E&quot;&gt;https://www.youtube.com/watch?v=-gfgQfw2g_E&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>從零打造語言模型</category>
    <category>強化學習</category>
    <category>大型語言模型入門</category>
  </item>
  <item>
    <title>PewDiePie is setting AI free... and OpenAI is furious</title>
    <link>https://aiwulin.itsmygo.uk/c/f16ac80d2c/</link>
    <guid isPermaLink="false">aiwulin-f16ac80d2c</guid>
    <pubDate>Tue, 06 Oct 2026 00:00:00 +0800</pubDate>
    <dc:creator>Fireship</dc:creator>
    <description>&lt;p&gt;介紹 PewDiePie 訓練的 Ajax 模型，這是基於 Qwen 3.5 並移除安全限制後，透過 GPT Soul 輸出進行蒸餾訓練而成的。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=_5p1_TNSWqQ&quot;&gt;https://www.youtube.com/watch?v=_5p1_TNSWqQ&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>開源模型</category>
    <category>強化學習</category>
    <category>量化與蒸餾</category>
  </item>
  <item>
    <title>Build A Reasoning Model From Scratch 6: Reinforcement Learning 1 (Implementing GRPO for RLVR)</title>
    <link>https://aiwulin.itsmygo.uk/c/1fbb7ea673/</link>
    <guid isPermaLink="false">aiwulin-1fbb7ea673</guid>
    <pubDate>Sat, 03 Oct 2026 00:00:00 +0800</pubDate>
    <dc:creator>Sebastian Raschka</dc:creator>
    <description>&lt;p&gt;由 Sebastian Raschka 講解如何從零開始用 Python 實現 GRPO 演算法，訓練小型推理模型。觀眾能學習 RLVR 的具體實作細節與訓練流程。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=237Hf7Q3lgg&quot;&gt;https://www.youtube.com/watch?v=237Hf7Q3lgg&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>強化學習</category>
    <category>研究前沿</category>
    <category>從零打造語言模型</category>
  </item>
  <item>
    <title>Beating RL With Reflection: GEPA and Optimize Anything — Lakshya A. Agrawal, GEPA</title>
    <link>https://aiwulin.itsmygo.uk/c/1e091cd2dc/</link>
    <guid isPermaLink="false">aiwulin-1e091cd2dc</guid>
    <pubDate>Sun, 27 Sep 2026 00:00:00 +0800</pubDate>
    <dc:creator>AI Engineer</dc:creator>
    <description>&lt;p&gt;Lakshya A. Agrawal 介紹 GEPA，一種透過完整執行軌跡與文字反思來最佳化提示詞的高效方法，比傳統強化學習更省資源。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=OA-Mc60Rboo&quot;&gt;https://www.youtube.com/watch?v=OA-Mc60Rboo&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>強化學習</category>
    <category>AGI 與長期趨勢</category>
    <category>提示工程</category>
  </item>
  <item>
    <title>Recursive-in-Recursive AI for Scientific AI Agents</title>
    <link>https://aiwulin.itsmygo.uk/c/467b03174b/</link>
    <guid isPermaLink="false">aiwulin-467b03174b</guid>
    <pubDate>Thu, 17 Sep 2026 00:00:00 +0800</pubDate>
    <dc:creator>Discover AI</dc:creator>
    <description>&lt;p&gt;介紹 ScienceBuddy 如何透過模型與工具協同進化，提升科學 AI 代理的準確率。內容說明將人類專家修正轉化為可執行任務，並透過 GWAS 與文獻推理實現自我改進。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=aMR_54K-KxE&quot;&gt;https://www.youtube.com/watch?v=aMR_54K-KxE&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>AI Agent 基礎</category>
  </item>
  <item>
    <title>Import AI 470: No rights for machines; automating environment generation with SPADE; and building better GPU kernels with Hawkeye</title>
    <link>https://aiwulin.itsmygo.uk/c/46c5642234/</link>
    <guid isPermaLink="false">aiwulin-46c5642234</guid>
    <pubDate>Mon, 24 Aug 2026 00:00:00 +0800</pubDate>
    <dc:creator>Import AI（Jack Clark）</dc:creator>
    <description>&lt;p&gt;探討 AI 在網路安全、數學與 AI 研究本身的加速差異，並介紹 SPADE 與 Hawkeye 等工具如何自動生成訓練環境與最佳化 GPU 核心程式碼。最後也討論了關於 AI 是否應有權利、人類意識危機等哲學議題。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://importai.substack.com/p/import-ai-470-no-rights-for-machines&quot;&gt;https://importai.substack.com/p/import-ai-470-no-rights-for-machines&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>AI 晶片與硬體</category>
  </item>
  <item>
    <title>Why Image Generation Needs More Than Bigger Models with Fatih Porikli - #773</title>
    <link>https://aiwulin.itsmygo.uk/c/2d6af6b27c/</link>
    <guid isPermaLink="false">aiwulin-2d6af6b27c</guid>
    <pubDate>Thu, 13 Aug 2026 00:00:00 +0800</pubDate>
    <dc:creator>The TWIML AI Podcast</dc:creator>
    <description>&lt;p&gt;探討文字轉影像模型雖能生成逼真畫面，卻在人物身份一致性、高解析度與邊緣裝置運算上仍有挑戰。Fatih Porikli 介紹了利用強化學習最佳化訓練目標，以提升畫面可控性、消除偽影並實現高效本地生成。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://twimlai.com/podcast/twimlai/why-image-generation-needs-more-than-bigger-models&quot;&gt;https://twimlai.com/podcast/twimlai/why-image-generation-needs-more-than-bigger-models&lt;/a&gt;&lt;/p&gt;</description>
    <category>Podcast</category>
    <category>圖像生成</category>
  </item>
  <item>
    <title>Chelsea Finn: This is the State of the Art in Robotics</title>
    <link>https://aiwulin.itsmygo.uk/c/19e26ce482/</link>
    <guid isPermaLink="false">aiwulin-19e26ce482</guid>
    <pubDate>Wed, 12 Aug 2026 00:00:00 +0800</pubDate>
    <dc:creator>Y Combinator</dc:creator>
    <description>&lt;p&gt;Chelsea Finn 分享 Physical Intelligence 如何透過強化學習與大模型訓練，開發能自主長時間運作的一般目的機器人。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=cRZNwgvcWUg&quot;&gt;https://www.youtube.com/watch?v=cRZNwgvcWUg&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>機器人與具身智慧</category>
    <category>強化學習</category>
  </item>
  <item>
    <title>GLM-5.2: DeepSeek Was Wrong About RL?</title>
    <link>https://aiwulin.itsmygo.uk/c/8cb27560b9/</link>
    <guid isPermaLink="false">aiwulin-8cb27560b9</guid>
    <pubDate>Fri, 24 Jul 2026 00:00:00 +0800</pubDate>
    <dc:creator>bycloud</dc:creator>
    <description>&lt;p&gt;討論 GLM-5.2 模型訓練策略，指出其放棄 GRPO 改用 PPO 的決定值得關注，並分享作者對模型表現之外的洞見。適合對大模型訓練技術感興趣的開發者與研究者。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=3KwpmSpEplY&quot;&gt;https://www.youtube.com/watch?v=3KwpmSpEplY&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>強化學習</category>
    <category>開源模型</category>
  </item>
  <item>
    <title>Why a Nation Can't Outsource Its Frontier AI - Alistair Pullen (Cosine AI)</title>
    <link>https://aiwulin.itsmygo.uk/c/5793875ca6/</link>
    <guid isPermaLink="false">aiwulin-5793875ca6</guid>
    <pubDate>Tue, 14 Jul 2026 00:00:00 +0800</pubDate>
    <dc:creator>Machine Learning Street Talk</dc:creator>
    <description>&lt;p&gt;Cosine 創辦人 Alistair Pullen 與 Tim Scarfe 訪談，探討因美國出口管制導致 Fable 模型受限後，英國如何透過 Isambard 超級電腦與企業協作，以低成本訓練國有大型語言模型。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://podcasters.spotify.com/pod/show/machinelearningstreettalk/episodes/Why-a-Nation-Cant-Outsource-Its-Frontier-AI---Alistair-Pullen-Cosine-AI-e3m1r3m&quot;&gt;https://podcasters.spotify.com/pod/show/machinelearningstreettalk/episodes/Why-a-Nation-Cant-Outsource-Its-Frontier-AI---Alistair-Pullen-Cosine-AI-e3m1r3m&lt;/a&gt;&lt;/p&gt;</description>
    <category>Podcast</category>
    <category>大型語言模型入門</category>
    <category>開源模型</category>
    <category>推論與部署</category>
  </item>
  <item>
    <title>Ornith 1.0: 自我優化代理式程式設計模型 [正體中文字幕]</title>
    <link>https://aiwulin.itsmygo.uk/c/6577c389cc/</link>
    <guid isPermaLink="false">aiwulin-6577c389cc</guid>
    <pubDate>Mon, 29 Jun 2026 00:00:00 +0800</pubDate>
    <dc:creator>Will 保哥</dc:creator>
    <description>&lt;p&gt;解析 Ornith-1 模型如何透過專業化訓練與動態控制架構，讓小型模型在效能上超越龐大對手。觀眾可了解其防止獎勵駭客機制與真實測試結果，掌握 AI 程式設計新趨勢。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=ruPMUk9xNzI&quot;&gt;https://www.youtube.com/watch?v=ruPMUk9xNzI&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>本機跑模型</category>
    <category>大型語言模型入門</category>
  </item>
  <item>
    <title>Introducing Ornith 1.0 - Agentic Coding LLMs</title>
    <link>https://aiwulin.itsmygo.uk/c/c64b281756/</link>
    <guid isPermaLink="false">aiwulin-c64b281756</guid>
    <pubDate>Fri, 26 Jun 2026 00:00:00 +0800</pubDate>
    <dc:creator>Sam Witteveen</dc:creator>
    <description>&lt;p&gt;介紹 Ornith 1.0 系列自架構編碼模型，能自動設計並執行任務指令碼。觀眾可學習模型如何利用強化學習訓練「自支撐」能力，以及如何在本地環境執行不同規模模型。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=uD4-uy0GmHE&quot;&gt;https://www.youtube.com/watch?v=uD4-uy0GmHE&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>Coding Agent</category>
    <category>大型語言模型入門</category>
  </item>
  <item>
    <title>AI Trends 2026: OpenClaw Agents, Reasoning LLMs, and More with Sebastian Raschka - #762</title>
    <link>https://aiwulin.itsmygo.uk/c/ef5550a5b0/</link>
    <guid isPermaLink="false">aiwulin-ef5550a5b0</guid>
    <pubDate>Fri, 27 Feb 2026 00:00:00 +0800</pubDate>
    <dc:creator>The TWIML AI Podcast</dc:creator>
    <description>&lt;p&gt;由 Sebastian Raschka 主講，分析 2026 年 AI 趨勢，重點在於從原始模型擴展轉向後訓練與推理技術。內容涵蓋自一致性、自最佳化等提升數學與編碼能力的推理方法，以及工具整合與代理工作流的實際應用。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://twimlai.com/podcast/twimlai/ai-trends-2026-openclaw-agents-reasoning-llms&quot;&gt;https://twimlai.com/podcast/twimlai/ai-trends-2026-openclaw-agents-reasoning-llms&lt;/a&gt;&lt;/p&gt;</description>
    <category>Podcast</category>
    <category>AI Agent 基礎</category>
    <category>大型語言模型入門</category>
    <category>Coding Agent</category>
  </item>
  <item>
    <title>State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka</title>
    <link>https://aiwulin.itsmygo.uk/c/bfc6d2e316/</link>
    <guid isPermaLink="false">aiwulin-bfc6d2e316</guid>
    <pubDate>Thu, 29 Jan 2026 00:00:00 +0800</pubDate>
    <dc:creator>The MAD Podcast</dc:creator>
    <description>&lt;p&gt;Sebastian Raschka 與 Matt Turck 深入探討 2025 至 2026 年大型語言模型的演進，涵蓋架構選擇、RLVR 與 GRPO 強化學習技術、推理縮放策略及評估方法。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://podcasters.spotify.com/pod/show/firstmark/episodes/State-of-LLMs-2026-RLVR--GRPO--Inference-Scaling--Sebastian-Raschka-e3eaha2&quot;&gt;https://podcasters.spotify.com/pod/show/firstmark/episodes/State-of-LLMs-2026-RLVR--GRPO--Inference-Scaling--Sebastian-Raschka-e3eaha2&lt;/a&gt;&lt;/p&gt;</description>
    <category>Podcast</category>
    <category>推論與部署</category>
    <category>大型語言模型入門</category>
    <category>研究前沿</category>
  </item>
  <item>
    <title>The State Of LLMs 2025: Progress, Problems, and Predictions</title>
    <link>https://aiwulin.itsmygo.uk/c/b1581a551b/</link>
    <guid isPermaLink="false">aiwulin-b1581a551b</guid>
    <pubDate>Tue, 30 Dec 2025 00:00:00 +0800</pubDate>
    <dc:creator>Sebastian Raschka（部落格）</dc:creator>
    <description>&lt;p&gt;Sebastian Raschka 回顧 2025 年大型語言模型發展，重點分析 DeepSeek R1 帶來的推理模型突破與 RLVR 演算法。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://magazine.sebastianraschka.com/p/state-of-llms-2025&quot;&gt;https://magazine.sebastianraschka.com/p/state-of-llms-2025&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>大型語言模型入門</category>
    <category>推論與部署</category>
    <category>研究前沿</category>
  </item>
  <item>
    <title>From DeepSeek V3 to V3.2: Architecture, Sparse Attention, and RL Updates</title>
    <link>https://aiwulin.itsmygo.uk/c/f02a2aa74e/</link>
    <guid isPermaLink="false">aiwulin-f02a2aa74e</guid>
    <pubDate>Wed, 03 Dec 2025 00:00:00 +0800</pubDate>
    <dc:creator>Sebastian Raschka（部落格）</dc:creator>
    <description>&lt;p&gt;深入解析 DeepSeek V3.2 的架構演進，涵蓋稀疏注意力機制、自驗證訓練方法與 GRPO 最佳化細節，並對比其與 V3、R1 及 Qwen3 等模型的差異。讀者可掌握最新開源大模型的技術亮點與訓練策略。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://magazine.sebastianraschka.com/p/technical-deepseek&quot;&gt;https://magazine.sebastianraschka.com/p/technical-deepseek&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>開源模型</category>
    <category>Transformer 原理</category>
    <category>強化學習</category>
  </item>
</channel>
</rss>
