---
id: "1884452216800686579"
user: "i_amanchadha"
name: "Aman Chadha"
status: ok
url: "https://x.com/i_amanchadha/status/1884452216800686579"
---

🧠 DeepSeek-R1 Primer • http://r1.aman.ai

- DeepSeek R1, the open-source reasoning model, is poised to democratize Large Reasoning Models (LRMs). R1 offers benchmark performance comparable with OpenAI's o1 at a fraction of the cost.
- This primer dissects the innovative architecture behind DeepSeek-R1 covering the following key aspects:

🔹 Architectural Foundations
- Mixture of Experts (MoE)
- Multihead Latent Attention (MLA)
- FP8 Quantization
- Multi-Token Prediction (MTP)
🔹 Training Pipeline: from Pre-Training to Reasoning
- Stage 1: Cold Start with Supervised Fine-Tuning (SFT)
- Stage 2: Reinforcement Learning (RL); Rewards
🔹 Group Relative Policy Optimization (GRPO)
- How GRPO Works
- Step-by-Step Breakdown
- PPO vs. DPO vs. KTO vs. APO vs. GRPO
🔹 Emergent Reasoning Behaviors
🔹 Distillation: Reasoning in Compact Models
🔹 Results
🔹 Open Questions
🔹 Open-R1
- Objectives of Open-R1
- Impact on the Community

🎨 Additionally, we cover DeepSeek's Janus-Pro (http://janus-pro.aman.ai), their cutting-edge multimodal model for understanding and generating across various media types.

Written in collaboration with @VinijaJain.

#AI #LLMs #Reasoning
