影片進階EN6 萬 次觀看
From Scratch to SOTA: Training a 3B State-Space Vision Model — Krishna Prasad Srinivasan, Sarvam
來源 AI Engineer
看影片(在新分頁開啟原站)連到 AI Engineer
摘要
Sarvam 的 Krishna Prasad Srinivasan 介紹了一款 30 億參數的視覺語言模型,採用非 Transformer 的 State Space Model 架構,能在單一 GPU 執行並超越百倍大小的模型。該模型專注於印度 22 種語言的區塊級 OCR,透過四階段課程與可驗證的強化學習,解決低資源語言的數位化難題。
A talk introducing a 3B parameter State Space Model that achieves SOTA in document AI for 22 Indian languages on a single GPU.
摘要、重點與章節標題由語言模型整理,細節(誰說的、數字、先後)可能有誤;要引用請以原始內容為準。
重點
- 採用 State Space Model 架構,解決長檔案處理的記憶體與計算成本問題。
- 透過四階段課程與可驗證的獎勵機制,在印度 22 種語言上達到 SOTA。
- 提供區塊級 OCR 與佈局解析,已協助企業數位化超過 3500 萬頁檔案。
章節
依話題轉折切分,標題由 AI 產生
- 00:00A 3B model that beats models 100 times larger on document AI
- 01:14Sarvam, and why India is missing from the machine readable world
- 02:37Why Indic document intelligence is hard
- 03:18The contrarian bet: block level OCR with a harness
- 04:31Why a state space model instead of a transformer
- 06:08Four stage curriculum: 13 trillion text tokens first
- 08:05The moat underneath: data engine and evals
- 09:13RL with verifiable rewards for OCR
- 09:54Results on English benchmarks and 22 Indian languages
- 10:50The agentic workbench and 35 million pages in production
- 12:39Q&A: low resource transfer, synthetic data, the Indic benchmark, sovereignty
提到的工具與公司
- Sarvam
- CRB Bench
- Gemini
- Opus
- Quen
適合誰看
從事檔案數位化、OCR 開發或關注印度低資源語言 AI 的技術人員與企業決策者。
摘要依據
- 講者
- Krishna Prasad Srinivasan
- 依據
- 自動字幕
為什麼排在這裡
- 人氣
- 0.89
- 新鮮
- 0.97
在主題頁與搜尋結果裡,名次由相關、人氣、新鮮三個分數決定;這一頁沒有搜尋的關鍵字,所以沒有相關分數。排序怎麼算
相關內容
- Stanford CS25: Transformers United V6 I From Language Models to Native Multimodal Intelligence影片 ・ Stanford Online ・ 1 小時 5 分
- Coding a Multimodal (Vision) Language Model from scratch in PyTorch with full explanation影片 ・ Umar Jamil ・ 5 小時 46 分(在新分頁開啟原站)
- Stanford CS336 Language Modeling from Scratch | Spring 2026 | Lecture 17: Alignment - Multimodality影片 ・ Stanford Online ・ 1 小時 18 分(在新分頁開啟原站)
- How Physical AI Learns Across Language, Video and Action — Ming-Yu LiuPodcast ・ Machine Learning Street Talk ・ 26 分鐘
- New Model: Inkling by Thinking Machine on Hugging Face影片 ・ Hugging Face ・ 36 分鐘(在新分頁開啟原站)
- Understanding and Implementing Qwen3 From Scratch文章 ・ Sebastian Raschka(部落格)(在新分頁開啟原站)
摘要由 AI 根據原文產生,可能有誤;完整內容請看原站。看影片(在新分頁開啟原站)
