Original ·1884452216800686579· Aman Chadha @i_amanchadha🧠 DeepSeek-R1 Primer • http://r1.aman.ai
- DeepSeek R1, the open-source reasoning model, is poised to democratize Large Reasoning Models (LRMs). R1 offers benchmark performance comparable with OpenAI's o1 at a fraction of the cost.
- This primer dissects the innovative architecture behind DeepSeek-R1 covering the following key aspects:
🔹 Architectural Foundations
- Mixture of Experts (MoE)
- Multihead Latent Attention (MLA)
- FP8 Quantization
- Multi-Token Prediction (MTP)
🔹 Training Pipeline: from Pre-Training to Reasoning
- Stage 1: Cold Start with Supervised Fine-Tuning (SFT)
- Stage 2: Reinforcement Learning (RL); Rewards
🔹 Group Relative Policy Optimization (GRPO)
- How GRPO Works
- Step-by-Step Breakdown
- PPO vs. DPO vs. KTO vs. APO vs. GRPO
🔹 Emergent Reasoning Behaviors
🔹 Distillation: Reasoning in Compact Models
🔹 Results
🔹 Open Questions
🔹 Open-R1
- Objectives of Open-R1
- Impact on the Community
🎨 Additionally, we cover DeepSeek's Janus-Pro (http://janus-pro.aman.ai), their cutting-edge multimodal model for understanding and generating across various media types.
Written in collaboration with @VinijaJain.
#AI #LLMs #Reasoning
RT @i_amanchadha: 🧠 DeepSeek-R1 Primer • https://t.co/YfEosETfrG
- DeepSeek R1, the open-source reasoning model, is poised to democratize…