跳到主要內容
AI 武林
影片進階EN8,248 次觀看

Scale the Judgment, Not the Model — Andrew Orobator, Reddit

來源 AI Engineer

看影片(在新分頁開啟原站)連到 YouTube・AI Engineer

其他版本:AIE Talks 摘要頁(在新分頁開啟)

摘要

Andrew Orobator 指出在程式碼代理中,模型並非瓶頸,關鍵在於如何將工程師的判斷(如審查標準、安全意識)轉化為程式碼中的技能、工作日誌與角色設定。透過這些機制,代理能像人類一樣累積經驗並建立信任,而非單純依賴更強的模型。

Andrew Orobator argues that scaling coding agents requires encoding human judgment into skills and logs, not just upgrading the model.

摘要、重點與章節標題由語言模型整理,細節(誰說的、數字、先後)可能有誤;要引用請以原始內容為準。

重點

  • 模型能力再強,若缺乏明確的判斷與審查機制,系統仍會失敗。
  • 將審查標準轉化為技能、工作日誌與角色,讓代理能繼承人類經驗。
  • 建立嚴格的驗證梯階與硬閘門,防止代理自行繞過安全控制。

章節

依話題轉折切分,標題由 AI 產生

  1. 00:00The war room for dead feature flags
  2. 00:37Intro
  3. 00:55What this talk isn't
  4. 01:55We're harness engineers now
  5. 02:25Humans absorb judgment; agents need it explicit
  6. 03:00The model isn't the bottleneck
  7. 03:30Judgment trapped in people's heads
  8. 04:40Skills as knowledge lines
  9. 05:15Skills vs. documentation
  10. 05:25Work logs
  11. 06:25This talk was built with a work log
  12. 06:45Personas
  13. 08:05Borrowing eyes you don't have
  14. 08:20Governance
  15. 09:00Verification
  16. 09:40The verification ladder
  17. 10:15Spin at the gate until green
  18. 10:35Make the agent record itself
  19. 11:20Case study: a feature-flag agent
  20. 12:101.26 per pull request
  21. 12:50Agents climb out of the pit of success
  22. 14:25A society of specialist agents
  23. 15:25In, on and off the loop
  24. 15:55Encoded judgment rots
  25. 17:05Managing judgment
  26. 18:15Scale the judgment, not the model

適合誰看

正在開發或整合 AI 程式碼代理、希望建立自動化工作流的工程師與技術管理者。

摘要依據

講者
Andrew Orobator
依據
自動字幕

為什麼排在這裡

人氣
0.71
新鮮
0.95

在主題頁與搜尋結果裡,名次由相關、人氣、新鮮三個分數決定;這一頁沒有搜尋的關鍵字,所以沒有相關分數。排序怎麼算

這個來源最近的內容

AI Engineer 的所有內容

摘要由 AI 根據原文產生,可能有誤;完整內容請看原站。看影片(在新分頁開啟原站)