看影片(在新分頁開啟原站)連到 AI Engineer
摘要
介紹 Project Paradox 框架與自動研究迴圈,解決長時程代理在記憶來源、不確定性與規劃上的失效問題。透過控制場景、追蹤記錄與平衡評分卡,系統能評估完整執行軌跡並僅保留改善行為的微小策略調整。
The talk presents Project Paradox and an autoresearch loop to fix long-horizon agent failures in memory and planning through controlled experiments.
摘要、重點與章節標題由語言模型整理,細節(誰說的、數字、先後)可能有誤;要引用請以原始內容為準。
重點
- 長時程代理需透過控制場景與完整軌跡評估,而非單次回應。
- 建立平衡評分卡與防範機制,避免最佳化單一指標導致不良行為。
- 凍結系統核心並僅暴露小範圍策略表面供自動研究搜尋。
章節
依話題轉折切分,標題由 AI 產生
- 00:00Intro
- 00:43Project Paradox
- 01:18What the agents can do
- 02:53A stateful architecture
- 03:33Trust scores and memory importance
- 04:28Demo: a picnic
- 05:08Where long-horizon behavior breaks
- 06:43Bringing in autoresearch
- 08:13Autoresearch outside the village
- 09:23The loop
- 10:43Controlled scenarios
- 12:33The mango rumor, fixed
- 13:03A balanced scorecard
- 14:33Keep the editable surface small
- 15:43Example policy changes
- 16:33Being careful with claims
- 17:03Memory is not enough
- 17:48Rollback is not optional
- 18:23Beyond games
- 19:33A recipe for long-horizon agents
- 20:13Closing
適合誰看
負責開發長期狀態維護型 AI 系統(如個人助理、研究代理或工作流自動化)的工程師。
摘要依據
- 講者
- Erina Karati
- 依據
- 自動字幕
為什麼排在這裡
- 人氣
- 0.78
- 新鮮
- 0.97
在主題頁與搜尋結果裡,名次由相關、人氣、新鮮三個分數決定;這一頁沒有搜尋的關鍵字,所以沒有相關分數。排序怎麼算
相關內容
- Long-running Agents文章 ・ Addy Osmani(部落格)
- How to Build Long-Horizon AI Agents — Mitch Troyanovsky, BasisPodcast ・ The MAD Podcast ・ 1 小時 23 分
- Agents, RAG, and Reasoning Models影片 ・ Luis Serrano Academy ・ 27 分鐘(在新分頁開啟原站)
- Effective harnesses for long-running agents文章 ・ Anthropic Engineering Blog
- Why Your Agent Needs Memory, Not Just ContextPodcast ・ The AI Native Dev ・ 44 分鐘
- State-Of-The-Art Prompting For AI AgentsPodcast ・ Lightcone Podcast ・ 31 分鐘
摘要由 AI 根據原文產生,可能有誤;完整內容請看原站。看影片(在新分頁開啟原站)
