影片進階EN1.4 萬 次觀看
Connect AI to Billions of Legal Documents — Simon Eskildsen, turbopuffer & Jacob Lauritzen, Legora
來源 AI Engineer
看影片(在新分頁開啟原站)連到 AI Engineer
摘要
分享 Legora 如何將 AI 應用於數億份法律檔案,解決搜尋延遲問題。透過將專案作為儲存基本單位,搭配 TurboPuffer 架構,實現資料物理隔離與低成本儲存。觀眾可學習大型法律資料庫的架構設計、儲存工程最佳化與向量搜尋原理。
This talk explains how Legora uses TurboPuffer to scale AI search on billions of legal documents with low latency and strict data isolation.
摘要、重點與章節標題由語言模型整理,細節(誰說的、數字、先後)可能有誤;要引用請以原始內容為準。
重點
- Legora 透過將專案作為儲存基本單位解決冷熱資料混雜問題。
- 使用 TurboPuffer 架構實現資料物理隔離與客戶管理加密金鑰。
- 以樹狀結構組織向量資料,減少記憶體往返次數提升效能。
章節
依話題轉折切分,標題由 AI 產生
- 00:00What Legora does, and two kinds of legal search
- 02:32One search cluster, then one per region
- 04:00What enterprises mean by physical isolation
- 04:48Moving search into the database they already ran
- 05:38How hot and cold projects thrashed the cache
- 06:40A namespace per project, and what it fixed
- 07:56Writing straight to object storage
- 09:58Why a namespace is the unit of encryption
- 12:07Legal research, and ten billion vectors
- 14:25Puffing data through the memory hierarchy
- 15:30Why a tree beats a graph on object storage
- 17:45How full text search actually works
提到的工具與公司
- Legora
- TurboPuffer
- S3
- NVMe SSD
- BM25
適合誰看
負責設計或管理大型法律資料庫、搜尋引擎或儲存架構的工程師與架構師。
摘要依據
- 講者
- Simon Eskildsen
- 依據
- 自動字幕
為什麼排在這裡
- 人氣
- 0.73
- 新鮮
- 0.93
在主題頁與搜尋結果裡,名次由相關、人氣、新鮮三個分數決定;這一頁沒有搜尋的關鍵字,所以沒有相關分數。排序怎麼算
相關內容
- Retrieval After RAG: Hybrid Search, Agents, and Database Design — Simon Hørup Eskildsen of TurbopufferPodcast ・ Latent Space ・ 1 小時 1 分
- Ep 71: CEO of TurboPuffer Simon Eskildsen on Building Smarter Retrieval, AI App Must-Have Features & Current State of Vector DBsPodcast ・ Unsupervised Learning ・ 51 分鐘
- How Hugging Face Inference Endpoints, Jobs, and Buckets Power Search on Papers with Code文章 ・ Hugging Face Blog
- Billion Scale Vector Storage for RAG影片 ・ Jason Liu ・ 51 分鐘(在新分頁開啟原站)
- 案件太多、敏感卷證又不能上商用AI,法務部啟動三年主權AI計畫文章 ・ iThome
- Ep 39. 和 Alex 聊聊向量数据库与职业规划Podcast ・ 捕蛇者说 ・ 1 小時 19 分
摘要由 AI 根據原文產生,可能有誤;完整內容請看原站。看影片(在新分頁開啟原站)
