跳到主要內容
AI 武林
影片進階EN9,965 次觀看

Long-Horizon Agents Need Experiments, Not Just Prompts — Erina Karati

來源 AI Engineer

看影片(在新分頁開啟原站)連到 AI Engineer

其他版本:AIE Talks 摘要頁(在新分頁開啟)

摘要

介紹 Project Paradox 框架與自動研究迴圈,解決長時程代理在記憶來源、不確定性與規劃上的失效問題。透過控制場景、追蹤記錄與平衡評分卡,系統能評估完整執行軌跡並僅保留改善行為的微小策略調整。

The talk presents Project Paradox and an autoresearch loop to fix long-horizon agent failures in memory and planning through controlled experiments.

摘要、重點與章節標題由語言模型整理,細節(誰說的、數字、先後)可能有誤;要引用請以原始內容為準。

重點

  • 長時程代理需透過控制場景與完整軌跡評估,而非單次回應。
  • 建立平衡評分卡與防範機制,避免最佳化單一指標導致不良行為。
  • 凍結系統核心並僅暴露小範圍策略表面供自動研究搜尋。

章節

依話題轉折切分,標題由 AI 產生

  1. 00:00Intro
  2. 00:43Project Paradox
  3. 01:18What the agents can do
  4. 02:53A stateful architecture
  5. 03:33Trust scores and memory importance
  6. 04:28Demo: a picnic
  7. 05:08Where long-horizon behavior breaks
  8. 06:43Bringing in autoresearch
  9. 08:13Autoresearch outside the village
  10. 09:23The loop
  11. 10:43Controlled scenarios
  12. 12:33The mango rumor, fixed
  13. 13:03A balanced scorecard
  14. 14:33Keep the editable surface small
  15. 15:43Example policy changes
  16. 16:33Being careful with claims
  17. 17:03Memory is not enough
  18. 17:48Rollback is not optional
  19. 18:23Beyond games
  20. 19:33A recipe for long-horizon agents
  21. 20:13Closing

適合誰看

負責開發長期狀態維護型 AI 系統(如個人助理、研究代理或工作流自動化)的工程師。

摘要依據

講者
Erina Karati
依據
自動字幕

為什麼排在這裡

人氣
0.78
新鮮
0.97

在主題頁與搜尋結果裡,名次由相關、人氣、新鮮三個分數決定;這一頁沒有搜尋的關鍵字,所以沒有相關分數。排序怎麼算

摘要由 AI 根據原文產生,可能有誤;完整內容請看原站。看影片(在新分頁開啟原站)