跳到主要內容
AI 武林
影片進階EN4,543 次觀看

From VLM/VLA's to Embodied Agents — Armen Aghajanyan, Perceptron AI

來源 AI Engineer

看影片(在新分頁開啟原站)連到 AI Engineer

其他版本:AIE Talks 摘要頁(在新分頁開啟)

摘要

演講者 Armen Aghajanyan 提出將視覺語言模型、具身推理與控制整合為單一具身基礎模型,解決影片預訓練資料稀疏與上下文膨脹問題。透過稀疏專家混合架構,模型能自動篩選關鍵視覺資訊,並將傳統電腦視覺任務轉化為代理行為,大幅減少對高成本遠端運算元據的依賴。

A talk proposing a unified embodied foundation model that combines perception, reasoning, and control to reduce reliance on expensive robotics data through sparse expert routing and massive multimodal pretraining.

摘要、重點與章節標題由語言模型整理,細節(誰說的、數字、先後)可能有誤;要引用請以原始內容為準。

重點

  • 提出具身基礎模型整合感知、推理與控制能力。
  • 利用稀疏專家架構自動篩選關鍵視覺 token。
  • 以海量多模態資料替代昂貴的遠端運算元據。

章節

依話題轉折切分,標題由 AI 產生

  1. 00:00Perceptron's north star: one model that perceives, reasons, and acts
  2. 01:23Early fusion, and the VLM to VLA to world model ladder
  3. 03:40Challenge one: an hour of video has almost no ground truth
  4. 05:41Challenge two: context bloat from cameras that never switch off
  5. 07:04Data sparse mixture of experts lets the model pick its tokens
  6. 08:54The first embodied foundation model and the petabyte behind it
  7. 09:50Detection as an agentic task: tile, zoom, raise the contrast
  8. 10:45Orchestrators, tactile policies, and self verifying annotation
  9. 12:35New scaling laws: trade teleop hours for video pretraining
  10. 14:11One model emitting control tokens, and open weights in July
  11. 16:17Q&A: temporal context, background robustness, structured extraction

提到的工具與公司

  • Perceptron
  • Gemini 3.1 Pro
  • VLM

適合誰看

正在開發具身 AI、機器人控制或多模態系統架構的研究人員與工程師。

摘要依據

講者
Armen Aghajanyan
依據
自動字幕

為什麼排在這裡

人氣
0.66
新鮮
0.95

在主題頁與搜尋結果裡,名次由相關、人氣、新鮮三個分數決定;這一頁沒有搜尋的關鍵字,所以沒有相關分數。排序怎麼算

摘要由 AI 根據原文產生,可能有誤;完整內容請看原站。看影片(在新分頁開啟原站)