跳到主要內容
AI 武林
影片進階EN6 萬 次觀看

From Scratch to SOTA: Training a 3B State-Space Vision Model — Krishna Prasad Srinivasan, Sarvam

來源 AI Engineer

看影片(在新分頁開啟原站)連到 AI Engineer

摘要

Sarvam 的 Krishna Prasad Srinivasan 介紹了一款 30 億參數的視覺語言模型,採用非 Transformer 的 State Space Model 架構,能在單一 GPU 執行並超越百倍大小的模型。該模型專注於印度 22 種語言的區塊級 OCR,透過四階段課程與可驗證的強化學習,解決低資源語言的數位化難題。

A talk introducing a 3B parameter State Space Model that achieves SOTA in document AI for 22 Indian languages on a single GPU.

摘要、重點與章節標題由語言模型整理,細節(誰說的、數字、先後)可能有誤;要引用請以原始內容為準。

重點

  • 採用 State Space Model 架構,解決長檔案處理的記憶體與計算成本問題。
  • 透過四階段課程與可驗證的獎勵機制,在印度 22 種語言上達到 SOTA。
  • 提供區塊級 OCR 與佈局解析,已協助企業數位化超過 3500 萬頁檔案。

章節

依話題轉折切分,標題由 AI 產生

  1. 00:00A 3B model that beats models 100 times larger on document AI
  2. 01:14Sarvam, and why India is missing from the machine readable world
  3. 02:37Why Indic document intelligence is hard
  4. 03:18The contrarian bet: block level OCR with a harness
  5. 04:31Why a state space model instead of a transformer
  6. 06:08Four stage curriculum: 13 trillion text tokens first
  7. 08:05The moat underneath: data engine and evals
  8. 09:13RL with verifiable rewards for OCR
  9. 09:54Results on English benchmarks and 22 Indian languages
  10. 10:50The agentic workbench and 35 million pages in production
  11. 12:39Q&A: low resource transfer, synthetic data, the Indic benchmark, sovereignty

提到的工具與公司

  • Sarvam
  • CRB Bench
  • Gemini
  • Opus
  • Quen

適合誰看

從事檔案數位化、OCR 開發或關注印度低資源語言 AI 的技術人員與企業決策者。

摘要依據

講者
Krishna Prasad Srinivasan
依據
自動字幕

為什麼排在這裡

人氣
0.89
新鮮
0.97

在主題頁與搜尋結果裡,名次由相關、人氣、新鮮三個分數決定;這一頁沒有搜尋的關鍵字,所以沒有相關分數。排序怎麼算

摘要由 AI 根據原文產生,可能有誤;完整內容請看原站。看影片(在新分頁開啟原站)