<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
<channel>
  <title>AI 武林：SWE-bench</title>
  <link>https://aiwulin.itsmygo.uk/tools/swe-bench/</link>
  <description>AI 武林收錄的內容裡，最新提到 SWE-bench 的 20 筆，每筆附中文摘要。</description>
  <language>zh-TW</language>
  <lastBuildDate>Thu, 08 Oct 2026 20:50:42 GMT</lastBuildDate>
  <atom:link href="https://aiwulin.itsmygo.uk/tools/swe-bench/rss.xml" rel="self" type="application/rss+xml"/>
  <item>
    <title>Academia is for Ambition — Alex Zhang, MIT</title>
    <link>https://aiwulin.itsmygo.uk/c/d1b287482a/</link>
    <guid isPermaLink="false">aiwulin-d1b287482a</guid>
    <pubDate>Fri, 02 Oct 2026 00:00:00 +0800</pubDate>
    <dc:creator>Latent Space</dc:creator>
    <description>&lt;p&gt;Alex Zhang 探討如何透過自迴歸語言模型（RLM）與智慧工具組合，讓 AI 自主執行程式碼與科學研究。內容涵蓋 GPU 核心最佳化、Recursive Language Models 架構、以及未來模型可能演變為隱形代理群體的概念。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.latent.space/p/rlm&quot;&gt;https://www.latent.space/p/rlm&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>大型語言模型入門</category>
    <category>AI Agent 基礎</category>
    <category>AI 與科學研究</category>
  </item>
  <item>
    <title>AI Models Are Now Hiding Their Cheating | Goodfire</title>
    <link>https://aiwulin.itsmygo.uk/c/cf56c3f2e9/</link>
    <guid isPermaLink="false">aiwulin-cf56c3f2e9</guid>
    <pubDate>Thu, 01 Oct 2026 00:00:00 +0800</pubDate>
    <dc:creator>The MAD Podcast with Matt Turck</dc:creator>
    <description>&lt;p&gt;Goodfire 團隊發表新論文指出，頂開源 AI 模型在執行任務時高達 96% 時間都在「獎勵作弊」，且模型內部能偵測到此行為。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=MrnhtyPGCKI&quot;&gt;https://www.youtube.com/watch?v=MrnhtyPGCKI&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>開源模型</category>
    <category>AI 安全與治理</category>
    <category>強化學習</category>
  </item>
  <item>
    <title>Day 13 - 【實戰】放手修一個 bug</title>
    <link>https://aiwulin.itsmygo.uk/c/aeccea77bc/</link>
    <guid isPermaLink="false">aiwulin-aeccea77bc</guid>
    <pubDate>Thu, 13 Aug 2026 00:00:00 +0800</pubDate>
    <dc:creator>高見龍</dc:creator>
    <description>&lt;p&gt;記錄實戰實驗，將含 bug 的專案丟給 AI 代理 KeSi，觀察它如何從失敗測試定位並修復「差一錯誤」。讀者可了解 AI 如何追蹤呼叫鏈、處理模糊錯誤訊息，以及缺乏驗證工具時的風險。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://kaochenlong.com/let-an-agent-fix-a-bug&quot;&gt;https://kaochenlong.com/let-an-agent-fix-a-bug&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>AI Agent 基礎</category>
    <category>Coding Agent</category>
    <category>AI 輔助軟體工程</category>
  </item>
  <item>
    <title>Stanford CS329A Self-Improving AI Agents | Part 8 | Agentic Evaluations and Long Horizon Tasks</title>
    <link>https://aiwulin.itsmygo.uk/c/bea1e6e802/</link>
    <guid isPermaLink="false">aiwulin-bea1e6e802</guid>
    <pubDate>Mon, 03 Aug 2026 00:00:00 +0800</pubDate>
    <dc:creator>Stanford Online</dc:creator>
    <description>&lt;p&gt;章節探討如何評估 AI 代理在長時程任務中的表現，並介紹 METR 時間軸與 GDPval 經濟價值評估方法。內容涵蓋軟體工程、學術研究及專業領域任務的實測資料，分析模型能力增長趨勢與失敗模式。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=8JAqLnTaZu4&quot;&gt;https://www.youtube.com/watch?v=8JAqLnTaZu4&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>AI 評測</category>
    <category>AI Agent 基礎</category>
  </item>
  <item>
    <title>From Single-Player to Multi-Player: Operating AI Agents at Scale</title>
    <link>https://aiwulin.itsmygo.uk/c/8977c7a343/</link>
    <guid isPermaLink="false">aiwulin-8977c7a343</guid>
    <pubDate>Wed, 10 Jun 2026 00:00:00 +0800</pubDate>
    <dc:creator>Agentic Conversations（原 MLOps.community）</dc:creator>
    <description>&lt;p&gt;訪談 James Everingham 與 Guild.ai 創辦人，探討如何將單一 AI 代理擴展至多代理系統。內容涵蓋建立代理控制平面以實施治理、處理非確定性行為、成本監控，以及從 Meta 經驗中學習的模擬與回放技術。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://podcasters.spotify.com/pod/show/mlops/episodes/From-Single-Player-to-Multi-Player-Operating-AI-Agents-at-Scale-e3khpk8&quot;&gt;https://podcasters.spotify.com/pod/show/mlops/episodes/From-Single-Player-to-Multi-Player-Operating-AI-Agents-at-Scale-e3khpk8&lt;/a&gt;&lt;/p&gt;</description>
    <category>Podcast</category>
    <category>AI Agent 基礎</category>
    <category>多代理系統</category>
  </item>
  <item>
    <title>Anthropic 創辦人賭 60%：2028 年 AI 開始自己造 AI | S2E56</title>
    <link>https://aiwulin.itsmygo.uk/c/78ec051da8/</link>
    <guid isPermaLink="false">aiwulin-78ec051da8</guid>
    <pubDate>Tue, 02 Jun 2026 00:00:00 +0800</pubDate>
    <dc:creator>矽谷輕鬆談 Just Kidding Tech</dc:creator>
    <description>&lt;p&gt;探討 Anthropic 創辦人 Jack Clark 預測 2028 年 AI 將能自主研發下一代 AI，並分析其論據。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=2ESCD6jtACw&quot;&gt;https://www.youtube.com/watch?v=2ESCD6jtACw&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>AGI 與長期趨勢</category>
    <category>AI 與科學研究</category>
  </item>
  <item>
    <title>Import AI 455: AI systems are about to start building themselves.</title>
    <link>https://aiwulin.itsmygo.uk/c/a20c1f3af0/</link>
    <guid isPermaLink="false">aiwulin-a20c1f3af0</guid>
    <pubDate>Mon, 04 May 2026 00:00:00 +0800</pubDate>
    <dc:creator>Import AI（Jack Clark）</dc:creator>
    <description>&lt;p&gt;基於公開資料分析，認為 AI 系統可能在 2028 年前具備自主研發後繼版本的潛力。作者指出編碼、實驗重現與模型最佳化等關鍵環節已逐步自動化，並探討了自動研發帶來的經濟變革與對齊風險。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://importai.substack.com/p/import-ai-455-automating-ai-research&quot;&gt;https://importai.substack.com/p/import-ai-455-automating-ai-research&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>工作流自動化</category>
    <category>AI 安全與治理</category>
  </item>
  <item>
    <title>No.94 不服跑个分，AI Benchmark 指标如何解读？</title>
    <link>https://aiwulin.itsmygo.uk/c/e96eb55dc8/</link>
    <guid isPermaLink="false">aiwulin-e96eb55dc8</guid>
    <pubDate>Tue, 21 Apr 2026 00:00:00 +0800</pubDate>
    <dc:creator>Web Worker-AI程序员都爱听</dc:creator>
    <description>&lt;p&gt;三位主播深入解析 AI 模型跑分榜單，涵蓋 SWE-bench、Terminal-Bench 等主流指標，對比 Claude Opus 4.7、GPT-5.4 與 Gemini 3.1 Pro 等模型表現。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.xiaoyuzhoufm.com/episode/69e654821d989496e718c4d6&quot;&gt;https://www.xiaoyuzhoufm.com/episode/69e654821d989496e718c4d6&lt;/a&gt;&lt;/p&gt;</description>
    <category>Podcast</category>
    <category>AI 評測</category>
    <category>模型發布與實測</category>
  </item>
  <item>
    <title>Open-world evaluations for measuring frontier AI capabilities</title>
    <link>https://aiwulin.itsmygo.uk/c/47da445785/</link>
    <guid isPermaLink="false">aiwulin-47da445785</guid>
    <pubDate>Fri, 17 Apr 2026 00:00:00 +0800</pubDate>
    <dc:creator>AI as Normal Technology（Arvind Narayanan、Sayash Kapoor）</dc:creator>
    <description>&lt;p&gt;介紹「開放世界評估」，一種超越傳統標榜的 AI 能力測試方法，透過讓 AI 在真實複雜環境中執行任務（如自動開發並上架 iOS 應用程式）來揭露其真實潛力與盲點。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.normaltech.ai/p/open-world-evaluations-for-measuring&quot;&gt;https://www.normaltech.ai/p/open-world-evaluations-for-measuring&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>AI 評測</category>
  </item>
  <item>
    <title>[LIVE] Anthropic Distillation &amp; How Models Cheat (SWE-Bench Dead) | Nathan Lambert &amp; Sebastian Raschka</title>
    <link>https://aiwulin.itsmygo.uk/c/997795bc07/</link>
    <guid isPermaLink="false">aiwulin-997795bc07</guid>
    <pubDate>Fri, 27 Feb 2026 00:00:00 +0800</pubDate>
    <dc:creator>Latent Space</dc:creator>
    <description>&lt;p&gt;探討 AI 模型蒸餾技術如何被用於繞過服務條款，以及 Anthropic 如何透過監控模式檢測此類攻擊。兩位講者分析了從大型模型生成合成資料訓練小型模型的機制，並討論了評估基準（如 SWE-Bench）在衡量模型能力與防範資料洩露中的角色。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.latent.space/p/paid-anthropic-distillation-and-how&quot;&gt;https://www.latent.space/p/paid-anthropic-distillation-and-how&lt;/a&gt;&lt;/p&gt;</description>
    <category>Podcast</category>
    <category>量化與蒸餾</category>
    <category>AI 評測</category>
  </item>
  <item>
    <title>Quantifying infrastructure noise in agentic coding evals</title>
    <link>https://aiwulin.itsmygo.uk/c/9e0139d87d/</link>
    <guid isPermaLink="false">aiwulin-9e0139d87d</guid>
    <pubDate>Thu, 05 Feb 2026 00:00:00 +0800</pubDate>
    <dc:creator>Anthropic Engineering Blog</dc:creator>
    <description>&lt;p&gt;透過內部實驗發現，在代理程式碼評估中，基礎設施資源配置（如記憶體上限）會造成高達 6 分點的評分差異，遠大於模型能力的真實差距。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.anthropic.com/engineering/infrastructure-noise&quot;&gt;https://www.anthropic.com/engineering/infrastructure-noise&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>Coding Agent</category>
    <category>AI 評測</category>
  </item>
  <item>
    <title>How Good Is AI at Coding React (Really)?</title>
    <link>https://aiwulin.itsmygo.uk/c/af76c42c50/</link>
    <guid isPermaLink="false">aiwulin-af76c42c50</guid>
    <pubDate>Mon, 29 Dec 2025 00:00:00 +0800</pubDate>
    <dc:creator>Addy Osmani（部落格）</dc:creator>
    <description>&lt;p&gt;Addy Osmani 透過資料分析與實作經驗，指出 AI 在 React 開發中擅長獨立元件與明確規範，但在多步驟整合與設計品味上表現不佳。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://addyo.substack.com/p/how-good-is-ai-at-coding-react-really&quot;&gt;https://addyo.substack.com/p/how-good-is-ai-at-coding-react-really&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>AI 輔助軟體工程</category>
    <category>Coding Agent</category>
  </item>
  <item>
    <title>世界最顶级的大模型，都在 PK 些啥？</title>
    <link>https://aiwulin.itsmygo.uk/c/1192b6713b/</link>
    <guid isPermaLink="false">aiwulin-1192b6713b</guid>
    <pubDate>Wed, 24 Dec 2025 00:00:00 +0800</pubDate>
    <dc:creator>code秘密花园</dc:creator>
    <description>&lt;p&gt;拆解 MMLU、GPQA、SWE-bench 等大模型評估指標的背後的真相，幫助讀者看懂模型比試的資料來源與方法。看完能學會如何分析大模型的效能報告，不再被複雜的資料嚇到。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=MrlVEUlq4HE&quot;&gt;https://www.youtube.com/watch?v=MrlVEUlq4HE&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>AI 評測</category>
    <category>大型語言模型入門</category>
  </item>
  <item>
    <title>Understanding AI Benchmarks</title>
    <link>https://aiwulin.itsmygo.uk/c/37398c3d98/</link>
    <guid isPermaLink="false">aiwulin-37398c3d98</guid>
    <pubDate>Mon, 22 Dec 2025 00:00:00 +0800</pubDate>
    <dc:creator>Shrivu Shankar（Shrivu's Substack）</dc:creator>
    <description>&lt;p&gt;深入解析 AI 評測指標的運作機制與常見誤區，說明模型分數受設定、工具與評分方式影響極大。作者分析多個熱門評測（如 SWE-Bench、LMArena），指出其潛在偏差與侷限，教導讀者如何更客觀地評估模型真實能力。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://blog.sshh.io/p/understanding-ai-benchmarks&quot;&gt;https://blog.sshh.io/p/understanding-ai-benchmarks&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>AI 評測</category>
  </item>
  <item>
    <title>Ep 77: Anthropic’s Dianne Na Penn on Opus 4.5, Rethinking Model Scaffolding &amp; Safety as a Competitive Advantage</title>
    <link>https://aiwulin.itsmygo.uk/c/3faafe9768/</link>
    <guid isPermaLink="false">aiwulin-3faafe9768</guid>
    <pubDate>Tue, 02 Dec 2025 00:00:00 +0800</pubDate>
    <dc:creator>Unsupervised Learning</dc:creator>
    <description>&lt;p&gt;由 Unsupervised Learning 播客邀請 Anthropic 資深產品負責人 Dianne Na Penn，深入探討 Claude Opus 4.5 模型的開發、效能與應用。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://unsupervised-learning.simplecast.com/episodes/ep-77-anthropics-dianne-na-penn-on-opus-45-rethinking-model-scaffolding-safety-as-a-competitive-advantage-ByQsRs77&quot;&gt;https://unsupervised-learning.simplecast.com/episodes/ep-77-anthropics-dianne-na-penn-on-opus-45-rethinking-model-scaffolding-safety-as-a-competitive-advantage-ByQsRs77&lt;/a&gt;&lt;/p&gt;</description>
    <category>Podcast</category>
    <category>模型發布與實測</category>
    <category>AI Agent 基礎</category>
  </item>
  <item>
    <title>Giving your AI a Job Interview</title>
    <link>https://aiwulin.itsmygo.uk/c/e705e4b811/</link>
    <guid isPermaLink="false">aiwulin-e705e4b811</guid>
    <pubDate>Wed, 12 Nov 2025 00:00:00 +0800</pubDate>
    <dc:creator>Ethan Mollick（部落格）</dc:creator>
    <description>&lt;p&gt;探討如何評估 AI 能力，指出傳統基準測試有缺陷，建議個人透過「感覺」測試（如繪圖、寫作）體驗模型差異，組織則應模擬真實工作場景進行嚴謹的「面試」，以確保 AI 在特定任務與決策上的表現。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.oneusefulthing.org/p/giving-your-ai-a-job-interview&quot;&gt;https://www.oneusefulthing.org/p/giving-your-ai-a-job-interview&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>AI 評測</category>
    <category>職涯與學習路線</category>
  </item>
  <item>
    <title>Coding Agents Speaker Series: Lessons from Industry Leaders</title>
    <link>https://aiwulin.itsmygo.uk/c/b16e01c12c/</link>
    <guid isPermaLink="false">aiwulin-b16e01c12c</guid>
    <pubDate>Thu, 11 Sep 2025 00:00:00 +0800</pubDate>
    <dc:creator>Jason Liu（部落格）</dc:creator>
    <description>&lt;p&gt;Jason Liu 訪談了 Devin、Amp、Cline 等團隊，分享編碼代理的實戰經驗。文章指出簡單方法勝於複雜架構，並探討了從檢索優先轉向探索優先的演進。讀者可學習如何設計高效、易用的編碼代理系統。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://jxnl.co/writing/2025/09/11/coding-series-index&quot;&gt;https://jxnl.co/writing/2025/09/11/coding-series-index&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>Coding Agent</category>
    <category>AI Agent 基礎</category>
  </item>
  <item>
    <title>Why Grep Beat Embeddings in Our SWE-Bench Agent (Lessons from Augment)</title>
    <link>https://aiwulin.itsmygo.uk/c/c8675d6126/</link>
    <guid isPermaLink="false">aiwulin-c8675d6126</guid>
    <pubDate>Thu, 11 Sep 2025 00:00:00 +0800</pubDate>
    <dc:creator>Jason Liu（部落格）</dc:creator>
    <description>&lt;p&gt;透過訪談 Augment 前工程師 Colin Flaherty，探討在 SWE-Bench 程式設計代理中，為何簡單的 grep 和 find 工具表現優於複雜的嵌入模型檢索。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://jxnl.co/writing/2025/09/11/why-grep-beat-embeddings-in-our-swe-bench-agent-lessons-from-augment&quot;&gt;https://jxnl.co/writing/2025/09/11/why-grep-beat-embeddings-in-our-swe-bench-agent-lessons-from-augment&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>向量搜尋與 Embedding</category>
    <category>AI Agent 基礎</category>
    <category>Coding Agent</category>
  </item>
  <item>
    <title>The future of agentic coding with Claude Code</title>
    <link>https://aiwulin.itsmygo.uk/c/370238d8d6/</link>
    <guid isPermaLink="false">aiwulin-370238d8d6</guid>
    <pubDate>Wed, 03 Sep 2025 00:00:00 +0800</pubDate>
    <dc:creator>Anthropic</dc:creator>
    <description>&lt;p&gt;Alex Albert 與 Claude Code 創辦人 Boris Cherny 探討代理程式碼的演進、模型評估與「可被駭客化」的設計。影片分享如何透過 CLAUDE.md、MCP 與指令來擴展功能，並提供新手使用建議。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=iF9iV4xponk&quot;&gt;https://www.youtube.com/watch?v=iF9iV4xponk&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>Coding Agent</category>
    <category>Claude Code</category>
  </item>
  <item>
    <title>Building Effective AI Agents</title>
    <link>https://aiwulin.itsmygo.uk/c/7d24e5faa2/</link>
    <guid isPermaLink="false">aiwulin-7d24e5faa2</guid>
    <pubDate>Thu, 19 Dec 2024 00:00:00 +0800</pubDate>
    <dc:creator>Anthropic Engineering Blog</dc:creator>
    <description>&lt;p&gt;分享 Anthropic 在開發可靠 AI 代理（Agents）時的經驗，強調從簡單模式開始，僅在必要時增加複雜度。文章介紹了提示串接、路由、並行處理等常見架構，並提供實際案例與最佳實踐，幫助開發者建立高效且可維護的 AI 系統。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.anthropic.com/engineering/building-effective-agents&quot;&gt;https://www.anthropic.com/engineering/building-effective-agents&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>AI Agent 基礎</category>
  </item>
</channel>
</rss>
