DeepSeek-R1: Incentivizing Reasoning Capability via RL
DeepSeek · 2025
"Reinforcement learning can elicit strong reasoning in LLMs at a fraction of frontier training cost."
DeepSeek, a Chinese AI lab, released an open-source reasoning model that matched top proprietary 'reasoning' models (like OpenAI's o1) on many benchmarks, while claiming a dramatically lower training and inference cost. The paper's core claim is that strong step-by-step reasoning ability can be taught to a language model largely through reinforcement learning — rewarding the model for correct final answers and consistent reasoning — rather than primarily through expensive, human-curated supervised fine-tuning data, which had been the more common assumption.
The mechanism, called Group Relative Policy Optimization (GRPO), trains the model by generating multiple candidate reasoning attempts for a given problem, scoring them with rule-based rewards (mainly: did it reach the correct final answer, and did it stay in one consistent language rather than mixing languages mid-response), and reinforcing the patterns that led to higher-scoring attempts. Pure reinforcement learning from scratch tends to produce unstable, repetitive, or language-mixing behavior early on, so the approach uses a small amount of 'cold-start' example data to stabilize the model before RL takes over as the primary training signal. The result is a model that learns to self-correct and verify its own reasoning steps largely through trial-and-reward rather than being shown millions of hand-labeled reasoning examples.
What training approach did DeepSeek-R1 rely on primarily to teach the model to reason, rather than mainly using human-labeled reasoning examples?
Read more about the topic
The explanation above is written with AI assistance. These are the originals — go to them to check it.
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL (original paper)arXiv, 2025
- DeepSeek-R1 explainerHugging Face
Software 2.0
"Neural networks are a new software paradigm where code is learned from data, not written by hand."