跳到主要內容
AI 武林
影片進階EN2.5 萬 次觀看

How We Built an Agent That Improves Itself — Zubin Aysola, Weights & Biases

來源 AI Engineer

看影片(在新分頁開啟原站)連到 AI Engineer

其他版本:AIE Talks 摘要頁(在新分頁開啟)

摘要

Weights & Biases 分享其自研的 ARIA 軟體代理如何透過生產與離線環境的完全一致程式碼,將實際使用者軌跡轉化為評估任務,自動發現並修復自身缺陷。觀眾可學習如何建立閉環的代理評估框架,利用生產資料驅動離線模擬,並透過 YAML 配置管理多變體實驗,提升系統自我演進能力。

Weights & Biases demonstrates how their ARIA agent uses identical production and offline code to self-evaluate and fix bugs by converting real user traces into benchmark tasks.

摘要、重點與章節標題由語言模型整理,細節(誰說的、數字、先後)可能有誤;要引用請以原始內容為準。

重點

  • 生產與離線環境使用完全一致的程式碼,確保評估結果可比。
  • 將生產軌跡轉化為 YAML 定義的評估任務,自動發現並修復錯誤。
  • 透過大量軌跡生成與變體測試,建立代理自我強化的飛輪效應。

章節

依話題轉折切分,標題由 AI 產生

  1. 00:00Intro: building ARIA
  2. 00:51Why build your own agent harness
  3. 01:31Benchmarks, evals and agents all change together
  4. 02:06Weave for production and offline tracing
  5. 02:56Live demo: ARIA doing autoresearch on itself
  6. 03:40Turning a production trace into an eval task
  7. 04:50Nightly CI evals
  8. 05:15Building a simulation environment
  9. 05:50Identical agents in production and research
  10. 06:40Generate lots of traces
  11. 07:15A model-agnostic harness and YAML variants
  12. 08:04An unconstrained sandbox
  13. 08:54The eval pipeline
  14. 10:14Scoring: pass/fail and relative
  15. 10:54Tasks as YAML: 886 tasks
  16. 12:03Trajectories and the eval flywheel
  17. 12:53Demo results: ARIA fixes its own bug
  18. 14:43Staying in the loop
  19. 16:22Wrap-up

提到的工具與公司

  • ARIA
  • Weave
  • CoreWeave
  • H200
  • Claude

適合誰看

正在開發或最佳化軟體代理、需要建立自動化評估框架的機器學習工程師。

摘要依據

講者
Zubin Aysola
依據
自動字幕

為什麼排在這裡

人氣
0.85
新鮮
0.97

在主題頁與搜尋結果裡,名次由相關、人氣、新鮮三個分數決定;這一頁沒有搜尋的關鍵字,所以沒有相關分數。排序怎麼算

摘要由 AI 根據原文產生,可能有誤;完整內容請看原站。看影片(在新分頁開啟原站)