影片進階EN6,936 次觀看
Small Language Model Alignment - Finetune SLMs to ALWAYS pick the best answer (Unsloth DPO)
看影片(在新分頁開啟原站)連到 YouTube・Neural Breakdown with AVB
摘要
介紹如何透過 DPO 方法微調小型語言模型,使其能穩定選擇最佳答案。觀眾將學習使用 Unsloth 與 Huggingface TRL 工具進行偏好最佳化訓練。
This video teaches how to fine-tune small language models using DPO to always select the best answer with Unsloth and Huggingface TRL.
這筆內容還沒有取得字幕或內文,這段摘要只根據標題與說明欄產生,可能不夠準確;實際內容請以原站為準。
提到的工具與公司
- Unsloth
- Huggingface TRL
- DPO
- Huggingface
適合誰看
適合有機器學習基礎、想實踐模型微調的開發者。
摘要依據
- 講者
- AVB
- 依據
- 標題與說明欄(還沒有取得字幕或內文)
為什麼排在這裡
- 人氣
- 0.31
- 新鮮
- 0.60
在主題頁與搜尋結果裡,名次由相關、人氣、新鮮三個分數決定;這一頁沒有搜尋的關鍵字,所以沒有相關分數。排序怎麼算
相關內容
- How to finetune LLMs on custom data domains (CPT tutorial with Unsloth)影片 ・ Neural Breakdown with AVB
- DPO 算法原理与代码实现:让 LLM 对齐变得简单文章 ・ chaofa用代码打点酱油(袁朝发)
- Direct Preference Optimization (DPO) explained: Bradley-Terry model, log probabilities, math影片 ・ Umar Jamil ・ 49 分鐘
- Whisper Data Preparation and Fine tuning with Unsloth影片 ・ Trelis Research ・ 41 分鐘
- Finetuning Open-Source LLMs影片 ・ Sebastian Raschka ・ 20 分鐘
- Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps文章 ・ Hugging Face Blog
摘要由 AI 根據標題與說明欄產生(還沒有取得原文),可能有誤;完整內容請看原站。看影片(在新分頁開啟原站)
