跳到主要內容
AI 武林
影片進階EN1.1 萬 次觀看

Skill issue: stop deploying vision language models, use them with Skills — Merve Noyan, Hugging Face

來源 AI Engineer

看影片(在新分頁開啟原站)連到 AI Engineer

摘要

演講者 Merve Noyan 建議開發者停止直接使用視覺語言模型(VLM)進行即時運算,改由小型檢測器處理,並推薦使用 Apache 2.0 授權的開源模型。她介紹了一套工具包,利用 VLM 作為標籤生成器與評判者,結合編碼代理進行「氛圍訓練」,最終訓練出高效能的 RF-DETR 模型。此流程成本低廉,且在路牌檢測與檔案解析任務中表現優異,能捕捉標註模型遺漏的細節。

This talk advocates replacing direct VLM usage with small detectors and a toolkit using VLMs as labelers and judges to train RF-DETR models via 'vibe training' on Apache 2.0 models.

摘要、重點與章節標題由語言模型整理,細節(誰說的、數字、先後)可能有誤;要引用請以原始內容為準。

重點

  • 停止直接使用 VLM 進行即時運算,改用小型檢測器以獲得更高效能。
  • 建立工具包讓編碼代理扮演無知工程師,自動處理模型訓練與部署。
  • 利用 VLM 作為標籤生成器與評判者,透過最小共識機制訓練 RF-DETR 模型。

章節

依話題轉折切分,標題由 AI 產生

  1. 00:00Why developers should stop reaching for a VLM at runtime
  2. 01:51Read the license: move to Apache 2.0 models
  3. 02:46A toolkit for coding agents, the clueless computer vision engineer
  4. 03:41Vibe training: VLM as labeler, VLMs as judges, then train
  5. 05:40The pipeline: overlaid boxes, minimum agreement, RF-DETR
  6. 08:29What it costs: a few dollars end to end
  7. 09:40Results on road signs and document parsing
  8. 11:18Findings: judge imbalance, prompt approval, augmentation blunders
  9. 13:20Favorite models as tools, from segmentation to depth
  10. 15:52Future plans: image guided detection and IoU merging
  11. 17:57Q&A

提到的工具與公司

  • RF-DETR
  • Gemma 4
  • LFM 2.5VL
  • Oppus 4.6
  • GLM 5.2
  • Roboflow

適合誰看

正在開發電腦視覺應用、需要處理即時影像或邊緣運算的開發者。

摘要依據

講者
Merve Noyan
依據
自動字幕

為什麼排在這裡

人氣
0.75
新鮮
0.95

在主題頁與搜尋結果裡,名次由相關、人氣、新鮮三個分數決定;這一頁沒有搜尋的關鍵字,所以沒有相關分數。排序怎麼算

摘要由 AI 根據原文產生,可能有誤;完整內容請看原站。看影片(在新分頁開啟原站)