RT 1884878912812990545
Original · 1884452216800686579 · Aman Chadha @i_amanchadha

🧠 DeepSeek-R1 Primer • http://r1.aman.ai

- DeepSeek R1, the open-source reasoning model, is poised to democratize Large Reasoning Models (LRMs). R1 offers benchmark performance comparable with OpenAI's o1 at a fraction of the cost.
- This primer dissects the innovative architecture behind DeepSeek-R1 covering the following key aspects:

🔹 Architectural Foundations
- Mixture of Experts (MoE)
- Multihead Latent Attention (MLA)
- FP8 Quantization
- Multi-Token Prediction (MTP)
🔹 Training Pipeline: from Pre-Training to Reasoning
- Stage 1: Cold Start with Supervised Fine-Tuning (SFT)
- Stage 2: Reinforcement Learning (RL); Rewards
🔹 Group Relative Policy Optimization (GRPO)
- How GRPO Works
- Step-by-Step Breakdown
- PPO vs. DPO vs. KTO vs. APO vs. GRPO
🔹 Emergent Reasoning Behaviors
🔹 Distillation: Reasoning in Compact Models
🔹 Results
🔹 Open Questions
🔹 Open-R1
- Objectives of Open-R1
- Impact on the Community

🎨 Additionally, we cover DeepSeek's Janus-Pro (http://janus-pro.aman.ai), their cutting-edge multimodal model for understanding and generating across various media types.

Written in collaboration with @VinijaJain.

#AI #LLMs #Reasoning

RT @i_amanchadha: 🧠 DeepSeek-R1 Primer • https://t.co/YfEosETfrG

- DeepSeek R1, the open-source reasoning model, is poised to democratize…

Fuente verbatim: corpus/posts/1884878912812990545.md · en X · acto Enero 2025