影片進階EN1.1 萬 次觀看
Skill issue: stop deploying vision language models, use them with Skills — Merve Noyan, Hugging Face
來源 AI Engineer
看影片(在新分頁開啟原站)連到 AI Engineer
摘要
演講者 Merve Noyan 建議開發者停止直接使用視覺語言模型(VLM)進行即時運算,改由小型檢測器處理,並推薦使用 Apache 2.0 授權的開源模型。她介紹了一套工具包,利用 VLM 作為標籤生成器與評判者,結合編碼代理進行「氛圍訓練」,最終訓練出高效能的 RF-DETR 模型。此流程成本低廉,且在路牌檢測與檔案解析任務中表現優異,能捕捉標註模型遺漏的細節。
This talk advocates replacing direct VLM usage with small detectors and a toolkit using VLMs as labelers and judges to train RF-DETR models via 'vibe training' on Apache 2.0 models.
摘要、重點與章節標題由語言模型整理,細節(誰說的、數字、先後)可能有誤;要引用請以原始內容為準。
重點
- 停止直接使用 VLM 進行即時運算,改用小型檢測器以獲得更高效能。
- 建立工具包讓編碼代理扮演無知工程師,自動處理模型訓練與部署。
- 利用 VLM 作為標籤生成器與評判者,透過最小共識機制訓練 RF-DETR 模型。
章節
依話題轉折切分,標題由 AI 產生
- 00:00Why developers should stop reaching for a VLM at runtime
- 01:51Read the license: move to Apache 2.0 models
- 02:46A toolkit for coding agents, the clueless computer vision engineer
- 03:41Vibe training: VLM as labeler, VLMs as judges, then train
- 05:40The pipeline: overlaid boxes, minimum agreement, RF-DETR
- 08:29What it costs: a few dollars end to end
- 09:40Results on road signs and document parsing
- 11:18Findings: judge imbalance, prompt approval, augmentation blunders
- 13:20Favorite models as tools, from segmentation to depth
- 15:52Future plans: image guided detection and IoU merging
- 17:57Q&A
提到的工具與公司
- RF-DETR
- Gemma 4
- LFM 2.5VL
- Oppus 4.6
- GLM 5.2
- Roboflow
適合誰看
正在開發電腦視覺應用、需要處理即時影像或邊緣運算的開發者。
摘要依據
- 講者
- Merve Noyan
- 依據
- 自動字幕
為什麼排在這裡
- 人氣
- 0.75
- 新鮮
- 0.95
在主題頁與搜尋結果裡,名次由相關、人氣、新鮮三個分數決定;這一頁沒有搜尋的關鍵字,所以沒有相關分數。排序怎麼算
相關內容
- Generalized Visual Language Models文章 ・ Lilian Weng(部落格)
- Why Vision Language Models Ignore What They See with Munawar Hayat - #758Podcast ・ The TWIML AI Podcast ・ 58 分鐘
- Accelerating vision-language models with LFM2.5-VL-DSpark文章 ・ Hugging Face Blog
- Coding a Multimodal (Vision) Language Model from scratch in PyTorch with full explanation影片 ・ Umar Jamil ・ 5 小時 46 分(在新分頁開啟原站)
- From VLM/VLA's to Embodied Agents — Armen Aghajanyan, Perceptron AI影片 ・ AI Engineer ・ 21 分鐘(在新分頁開啟原站)
- Optimize, deploy, and benchmark an open-source LLM with vLLM影片 ・ DeepLearningAI ・ 2 分鐘(在新分頁開啟原站)
摘要由 AI 根據原文產生,可能有誤;完整內容請看原站。看影片(在新分頁開啟原站)
