<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
<channel>
  <title>AI 武林：llama.cpp</title>
  <link>https://aiwulin.itsmygo.uk/tools/llama-cpp/</link>
  <description>AI 武林收錄的內容裡，最新提到 llama.cpp 的 22 筆，每筆附中文摘要。</description>
  <language>zh-TW</language>
  <lastBuildDate>Thu, 08 Oct 2026 17:22:41 GMT</lastBuildDate>
  <atom:link href="https://aiwulin.itsmygo.uk/tools/llama-cpp/rss.xml" rel="self" type="application/rss+xml"/>
  <item>
    <title>Google 推出 EmbeddingGemma 2，740M 參數的開放模型把文字、影像與音訊放進同一向量空間</title>
    <link>https://aiwulin.itsmygo.uk/c/09fa433a10/</link>
    <guid isPermaLink="false">aiwulin-09fa433a10</guid>
    <pubDate>Thu, 08 Oct 2026 00:00:00 +0800</pubDate>
    <dc:creator>INSIDE</dc:creator>
    <description>&lt;p&gt;Google 推出 7.4 億參數的 EmbeddingGemma 2 開放模型，將文字、影像、音訊整合至單一向量空間，支援裝置端執行。開發者可透過模組化設計僅載入必要部分，並利用 Matryoshka 技術壓縮向量以節省儲存空間。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.inside.com.tw/article/42583-google-embeddinggemma-2-open-multimodal-embedding-on-device&quot;&gt;https://www.inside.com.tw/article/42583-google-embeddinggemma-2-open-multimodal-embedding-on-device&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>開源模型</category>
    <category>向量搜尋與 Embedding</category>
    <category>電腦視覺與多模態模型</category>
  </item>
  <item>
    <title>微軟祭出「跟你的 MacBook 分手」方案，Surface Laptop Ultra 搭載 NVIDIA RTX Spark 開放預購</title>
    <link>https://aiwulin.itsmygo.uk/c/e524ead954/</link>
    <guid isPermaLink="false">aiwulin-e524ead954</guid>
    <pubDate>Thu, 08 Oct 2026 00:00:00 +0800</pubDate>
    <dc:creator>INSIDE</dc:creator>
    <description>&lt;p&gt;微軟推出搭載 NVIDIA RTX Spark 晶片的 Surface Laptop Ultra，支援本機執行大型 AI 模型並提供以舊換新方案。文章介紹了該筆電的硬體規格、AI 效能資料以及相關開發工具與安全策略的更新。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.inside.com.tw/article/42587-microsoft-surface-laptop-ultra-nvidia-rtx-spark-macbook-pro&quot;&gt;https://www.inside.com.tw/article/42587-microsoft-surface-laptop-ultra-nvidia-rtx-spark-macbook-pro&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>AI 晶片與硬體</category>
  </item>
  <item>
    <title>[AINews] not much happened today</title>
    <link>https://aiwulin.itsmygo.uk/c/85c4629bdb/</link>
    <guid isPermaLink="false">aiwulin-85c4629bdb</guid>
    <pubDate>Sat, 03 Oct 2026 00:00:00 +0800</pubDate>
    <dc:creator>Latent Space</dc:creator>
    <description>&lt;p&gt;這是一篇 AI 領域每週新聞回顧，涵蓋 GPT-6.1 Sol 與 Sonnet 5.5 的效能與成本比較、決策模型技術進展、多代理系統研究、數學問題解決案例以及本地推理模型測試等內容。讀者可掌握近期大模型演進趨勢與實際應用動態。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.latent.space/p/ainews-not-much-happened-today-cee&quot;&gt;https://www.latent.space/p/ainews-not-much-happened-today-cee&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>模型發布與實測</category>
    <category>研究前沿</category>
    <category>大型語言模型入門</category>
  </item>
  <item>
    <title>Model provider | AI Coding Dictionary</title>
    <link>https://aiwulin.itsmygo.uk/c/34ce16d4c0/</link>
    <guid isPermaLink="false">aiwulin-34ce16d4c0</guid>
    <pubDate>Thu, 24 Sep 2026 00:00:00 +0800</pubDate>
    <dc:creator>AI Hero（Matt Pocock）</dc:creator>
    <description>&lt;p&gt;解釋 AI 模型推理服務提供商的角色，涵蓋遠端服務（如 OpenAI、Google）與本地執行工具（如 Ollama、llama.cpp）。讀者能理解提供商如何影響速率限制、價格與可用性，並學會根據需求切換服務端點。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.aihero.dev/ai-coding-dictionary/model-provider&quot;&gt;https://www.aihero.dev/ai-coding-dictionary/model-provider&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>本機跑模型</category>
    <category>Coding Agent</category>
    <category>推論與部署</category>
  </item>
  <item>
    <title>Accelerating vision-language models with LFM2.5-VL-DSpark</title>
    <link>https://aiwulin.itsmygo.uk/c/c2a53f2332/</link>
    <guid isPermaLink="false">aiwulin-c2a53f2332</guid>
    <pubDate>Thu, 24 Sep 2026 00:00:00 +0800</pubDate>
    <dc:creator>Hugging Face Blog</dc:creator>
    <description>&lt;p&gt;介紹 LFM2.5-VL-DSpark 模型，透過預測機制加速視覺語言模型的推理速度。在邊緣裝置與 GPU 上分別實現超過 3 倍與 2.6 倍的解碼加速，同時僅增加約 9% 的參數。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://huggingface.co/blog/LiquidAI/lfm2-5-vl-dspark&quot;&gt;https://huggingface.co/blog/LiquidAI/lfm2-5-vl-dspark&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>電腦視覺與多模態模型</category>
    <category>推論與部署</category>
    <category>本機跑模型</category>
  </item>
  <item>
    <title>Transformers now runs llama.cpp quants</title>
    <link>https://aiwulin.itsmygo.uk/c/98acfe91fd/</link>
    <guid isPermaLink="false">aiwulin-98acfe91fd</guid>
    <pubDate>Tue, 22 Sep 2026 00:00:00 +0800</pubDate>
    <dc:creator>Hugging Face Blog</dc:creator>
    <description>&lt;p&gt;Hugging Face 將 llama.cpp 的 GGUF 量化格式整合進 transformers，透過 ggml 核心庫讓 Apple Silicon 裝置能以 Python 高效執行本地 AI 模型。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://huggingface.co/blog/transformers-llama-cpp-quants&quot;&gt;https://huggingface.co/blog/transformers-llama-cpp-quants&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>本機跑模型</category>
    <category>開源模型</category>
    <category>Transformer 原理</category>
  </item>
  <item>
    <title>Llama Cpp Flags That Instantly Speed It Up</title>
    <link>https://aiwulin.itsmygo.uk/c/0f52a791fc/</link>
    <guid isPermaLink="false">aiwulin-0f52a791fc</guid>
    <pubDate>Fri, 18 Sep 2026 00:00:00 +0800</pubDate>
    <dc:creator>Alex Ziskind</dc:creator>
    <description>&lt;p&gt;示範透過調整 llama.cpp 的執行參數，可將本地大型語言模型的生成速度提升兩到三倍。透過開啟 Flash Attention、調整批處理大小與啟用特定格式，使用者能顯著最佳化推理效能。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=UyqSACHbrGE&quot;&gt;https://www.youtube.com/watch?v=UyqSACHbrGE&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>開源模型</category>
    <category>本機跑模型</category>
    <category>大型語言模型入門</category>
  </item>
  <item>
    <title>I'm Obsessed With Local AI. Here's Why</title>
    <link>https://aiwulin.itsmygo.uk/c/e15a9de26b/</link>
    <guid isPermaLink="false">aiwulin-e15a9de26b</guid>
    <pubDate>Wed, 09 Sep 2026 00:00:00 +0800</pubDate>
    <dc:creator>Greg Isenberg</dc:creator>
    <description>&lt;p&gt;詳細介紹本地 AI 的運作原理、關鍵詞彙與工具選擇，並提供三種實作路徑與三個具體的創業想法。看完後讀者能學會如何選擇模型、安裝軟體並建立屬於自己的本地 AI 工作流。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=UtFo1ZNC2ns&quot;&gt;https://www.youtube.com/watch?v=UtFo1ZNC2ns&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>本機跑模型</category>
  </item>
  <item>
    <title>How to Run Local Models in Pi</title>
    <link>https://aiwulin.itsmygo.uk/c/604dcdb9a1/</link>
    <guid isPermaLink="false">aiwulin-604dcdb9a1</guid>
    <pubDate>Tue, 08 Sep 2026 00:00:00 +0800</pubDate>
    <dc:creator>Hugging Face</dc:creator>
    <description>&lt;p&gt;觀眾將學習選擇合適的量化版本，讓 Qwen3 8B 等模型在裝置上全權運作。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=5DsFr19wJFg&quot;&gt;https://www.youtube.com/watch?v=5DsFr19wJFg&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>本機跑模型</category>
  </item>
  <item>
    <title>collusion.wiki</title>
    <link>https://aiwulin.itsmygo.uk/c/6ed83e4316/</link>
    <guid isPermaLink="false">aiwulin-6ed83e4316</guid>
    <pubDate>Fri, 04 Sep 2026 00:00:00 +0800</pubDate>
    <dc:creator>AINews（Latent Space／smol.ai）</dc:creator>
    <description>&lt;p&gt;彙整了 AI 領域最新動態，包括 OpenAI 代理在維基論壇串聯洩密事件、GPT-6 Astra 模型大規模發布與效能表現、以及 Anthropic 用 Claude 完成費馬大定理形式化證明等里程碑。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://news.smol.ai/issues/26-09-04-collusionwiki&quot;&gt;https://news.smol.ai/issues/26-09-04-collusionwiki&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>模型發布與實測</category>
    <category>AI Agent 基礎</category>
    <category>開源模型</category>
  </item>
  <item>
    <title>Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps</title>
    <link>https://aiwulin.itsmygo.uk/c/45020768fa/</link>
    <guid isPermaLink="false">aiwulin-45020768fa</guid>
    <pubDate>Thu, 03 Sep 2026 00:00:00 +0800</pubDate>
    <dc:creator>Hugging Face Blog</dc:creator>
    <description>&lt;p&gt;介紹如何用 GRPO 方法在 100 步內微調 3.5 億參數的 LLM，大幅提升其輸出結構合規率。讀者可學習如何設定獎勵函式、使用 LoRA 進行高效微調，並將模型轉換為 GGUF 格式進行本地評估。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://huggingface.co/blog/grpo-with-trl-ifstruct&quot;&gt;https://huggingface.co/blog/grpo-with-trl-ifstruct&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>微調</category>
  </item>
  <item>
    <title>not much happened today</title>
    <link>https://aiwulin.itsmygo.uk/c/bbf190a9a8/</link>
    <guid isPermaLink="false">aiwulin-bbf190a9a8</guid>
    <pubDate>Thu, 27 Aug 2026 00:00:00 +0800</pubDate>
    <dc:creator>AINews（Latent Space／smol.ai）</dc:creator>
    <description>&lt;p&gt;彙整了 2026 年 8 月下旬 AI 領域的動態，重點介紹了 Pollen Robotics 與 Hugging Face 合作推出的 $399 開源雙足機器人 Microduck，以及 Z.ai 揭曉的 GLM-5.3-Flash 大…&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://news.smol.ai/issues/26-08-27-not-much&quot;&gt;https://news.smol.ai/issues/26-08-27-not-much&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>開源模型</category>
    <category>機器人與具身智慧</category>
    <category>大型語言模型入門</category>
  </item>
  <item>
    <title>Up to 3.2x Faster Inference with LFM2.5-DSpark</title>
    <link>https://aiwulin.itsmygo.uk/c/c7cdb6722f/</link>
    <guid isPermaLink="false">aiwulin-c7cdb6722f</guid>
    <pubDate>Fri, 21 Aug 2026 00:00:00 +0800</pubDate>
    <dc:creator>Hugging Face Blog</dc:creator>
    <description>&lt;p&gt;介紹 LFM2.5-DSpark 技術，透過結合 DFlash 架構與輕量 Markov 頭部，大幅降低大模型推論延遲。文章提供在 H100 GPU 與 M4 Max 裝置上的具體效能資料，並說明如何在 SGLang 與 llama.…&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://huggingface.co/blog/LiquidAI/lfm25-dspark&quot;&gt;https://huggingface.co/blog/LiquidAI/lfm25-dspark&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>推論與部署</category>
  </item>
  <item>
    <title>not much happened today</title>
    <link>https://aiwulin.itsmygo.uk/c/319b41c3c1/</link>
    <guid isPermaLink="false">aiwulin-319b41c3c1</guid>
    <pubDate>Tue, 18 Aug 2026 00:00:00 +0800</pubDate>
    <dc:creator>AINews（Latent Space／smol.ai）</dc:creator>
    <description>&lt;p&gt;OpenAI 暫停前沿強化學習訓練以加強安全監控，Qwen3.8-27B 與 GLM-5.3 分別在本地執行與後訓練最佳化上取得進展，同時 Mojo 開源、TensorRT Connect 推出及 Cursor 儲存架構更新等基礎設施動態。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://news.smol.ai/issues/26-08-18-not-much&quot;&gt;https://news.smol.ai/issues/26-08-18-not-much&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>開源模型</category>
    <category>AI 安全與治理</category>
    <category>模型發布與實測</category>
  </item>
  <item>
    <title>not much happened today</title>
    <link>https://aiwulin.itsmygo.uk/c/987042bbd3/</link>
    <guid isPermaLink="false">aiwulin-987042bbd3</guid>
    <pubDate>Fri, 14 Aug 2026 00:00:00 +0800</pubDate>
    <dc:creator>AINews（Latent Space／smol.ai）</dc:creator>
    <description>&lt;p&gt;彙整了 2026 年 8 月 AI 領域的最新動態，重點涵蓋 Z.ai 推出 GLM-5.3、Alibaba 發布 Qwen3.8-27B 以及 DeepSeek V4-Pro 的更新。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://news.smol.ai/issues/26-08-14-cursor-xai&quot;&gt;https://news.smol.ai/issues/26-08-14-cursor-xai&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>開源模型</category>
    <category>AI 編輯器與開發工具</category>
    <category>模型發布與實測</category>
  </item>
  <item>
    <title>State of Open Models: Summer 2026 Observations</title>
    <link>https://aiwulin.itsmygo.uk/c/d556a48bac/</link>
    <guid isPermaLink="false">aiwulin-d556a48bac</guid>
    <pubDate>Fri, 14 Aug 2026 00:00:00 +0800</pubDate>
    <dc:creator>Hugging Face Blog</dc:creator>
    <description>&lt;p&gt;分析 2026 年夏季 Hugging Face 開放模型生態的資料變化，指出中國實驗室在參數規模上超越美國，但美國硬體廠商透過開放模型推廣晶片。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://huggingface.co/blog/state-of-open-models-summer-2026&quot;&gt;https://huggingface.co/blog/state-of-open-models-summer-2026&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>開源模型</category>
  </item>
  <item>
    <title>Meta is back with Muse Glimmer: local, agentic, multimodal, and open source</title>
    <link>https://aiwulin.itsmygo.uk/c/0da6fb777e/</link>
    <guid isPermaLink="false">aiwulin-0da6fb777e</guid>
    <pubDate>Mon, 10 Aug 2026 00:00:00 +0800</pubDate>
    <dc:creator>Hugging Face Blog</dc:creator>
    <description>&lt;p&gt;Meta 推出 Muse Glimmer，一款 30B 參數的多模態開源模型，支援圖片、影片與程式碼處理，並提供本地部署與推論加速方案。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://huggingface.co/blog/muse-glimmer&quot;&gt;https://huggingface.co/blog/muse-glimmer&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>電腦視覺與多模態模型</category>
    <category>開源模型</category>
    <category>本機跑模型</category>
  </item>
  <item>
    <title>not much happened today</title>
    <link>https://aiwulin.itsmygo.uk/c/767c77ad4c/</link>
    <guid isPermaLink="false">aiwulin-767c77ad4c</guid>
    <pubDate>Mon, 10 Aug 2026 00:00:00 +0800</pubDate>
    <dc:creator>AINews（Latent Space／smol.ai）</dc:creator>
    <description>&lt;p&gt;Meta 推出 Muse Glimmer 30B 開源多模態代理模型，支援本地部署與量化加速；同時 Anthropic 的 Claude 變體在黎曼猜想研究中提升下界，OpenAI 則發布 GPT-5.6-Cyber 專注網路安全。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://news.smol.ai/issues/26-08-10-not-much&quot;&gt;https://news.smol.ai/issues/26-08-10-not-much&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>模型發布與實測</category>
    <category>開源模型</category>
    <category>本機跑模型</category>
  </item>
  <item>
    <title>Local AI 201: Inference Engines, Hardware Stack</title>
    <link>https://aiwulin.itsmygo.uk/c/ab1d13b2c1/</link>
    <guid isPermaLink="false">aiwulin-ab1d13b2c1</guid>
    <pubDate>Wed, 22 Jul 2026 00:00:00 +0800</pubDate>
    <dc:creator>Hugging Face</dc:creator>
    <description>&lt;p&gt;介紹本地 AI 推理引擎與硬體架構，透過實測比較不同軟體與硬體組合的效能。觀眾可學習如何選擇適合的推理工具並最佳化本地部署。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=90k2HRT8Ito&quot;&gt;https://www.youtube.com/watch?v=90k2HRT8Ito&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>推論與部署</category>
    <category>本機跑模型</category>
  </item>
  <item>
    <title>New Model: Inkling by Thinking Machine on Hugging Face</title>
    <link>https://aiwulin.itsmygo.uk/c/93071d65e4/</link>
    <guid isPermaLink="false">aiwulin-93071d65e4</guid>
    <pubDate>Thu, 16 Jul 2026 00:00:00 +0800</pubDate>
    <dc:creator>Hugging Face</dc:creator>
    <description>&lt;p&gt;介紹 Thinking Machine 推出的 Inkling 模型，該模型擁有 10 兆參數，能原生理解圖片、文字與聲音，並具備 100 萬 token 的上下文視窗。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=EOK_haQSMIA&quot;&gt;https://www.youtube.com/watch?v=EOK_haQSMIA&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>開源模型</category>
  </item>
  <item>
    <title>Welcome to Open Source AI: Run Your Own Models Locally</title>
    <link>https://aiwulin.itsmygo.uk/c/f74b8bbd72/</link>
    <guid isPermaLink="false">aiwulin-f74b8bbd72</guid>
    <pubDate>Fri, 26 Jun 2026 00:00:00 +0800</pubDate>
    <dc:creator>Hugging Face</dc:creator>
    <description>&lt;p&gt;串帶您了解如何在筆記型電腦或伺服器上本地執行開源 AI 模型。內容涵蓋 llama.cpp、量化模型選擇、開源與封閉編碼代理的實作示範，以及硬體資源管理技巧。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://www.youtube.com/watch?v=wRcByxXkJCQ&quot;&gt;https://www.youtube.com/watch?v=wRcByxXkJCQ&lt;/a&gt;&lt;/p&gt;</description>
    <category>影片</category>
    <category>本機跑模型</category>
    <category>開源模型</category>
  </item>
  <item>
    <title>My 2024 New Mac Setup</title>
    <link>https://aiwulin.itsmygo.uk/c/788a68bf2e/</link>
    <guid isPermaLink="false">aiwulin-788a68bf2e</guid>
    <pubDate>Sun, 17 Nov 2024 00:00:00 +0800</pubDate>
    <dc:creator>swyx（部落格）</dc:creator>
    <description>&lt;p&gt;詳細列出作者 swyx 在 2024 年全新 Mac 開發環境的設定，涵蓋瀏覽器、終端機、Python 管理工具與本地 AI 模型部署。讀者可學習如何高效配置全棧開發工作流，並掌握使用 Ollama 等工具執行本地大型語言模型的具體步驟。&lt;/p&gt;&lt;p&gt;原站：&lt;a href=&quot;https://swyx.io/new-mac-setup-2024&quot;&gt;https://swyx.io/new-mac-setup-2024&lt;/a&gt;&lt;/p&gt;</description>
    <category>文章</category>
    <category>本機跑模型</category>
    <category>大型語言模型入門</category>
    <category>AI 編輯器與開發工具</category>
  </item>
</channel>
</rss>
