Skip to content
← The Scroll
Contemporary · Paper

DeepSeek-R1: Incentivizing Reasoning Capability via RL

DeepSeek · 2025

"Reinforcement learning can elicit strong reasoning in LLMs at a fraction of frontier training cost."

The idea

DeepSeek, a Chinese AI lab, released an open-source reasoning model that matched top proprietary 'reasoning' models (like OpenAI's o1) on many benchmarks, while claiming a dramatically lower training and inference cost. The paper's core claim is that strong step-by-step reasoning ability can be taught to a language model largely through reinforcement learning — rewarding the model for correct final answers and consistent reasoning — rather than primarily through expensive, human-curated supervised fine-tuning data, which had been the more common assumption.

Why it works

The mechanism, called Group Relative Policy Optimization (GRPO), trains the model by generating multiple candidate reasoning attempts for a given problem, scoring them with rule-based rewards (mainly: did it reach the correct final answer, and did it stay in one consistent language rather than mixing languages mid-response), and reinforcing the patterns that led to higher-scoring attempts. Pure reinforcement learning from scratch tends to produce unstable, repetitive, or language-mixing behavior early on, so the approach uses a small amount of 'cold-start' example data to stabilize the model before RL takes over as the primary training signal. The result is a model that learns to self-correct and verify its own reasoning steps largely through trial-and-reward rather than being shown millions of hand-labeled reasoning examples.

The takeaway — recall it first
Check your understanding

What training approach did DeepSeek-R1 rely on primarily to teach the model to reason, rather than mainly using human-labeled reasoning examples?

Further reading

Read more about the topic

The explanation above is written with AI assistance. These are the originals — go to them to check it.

  • DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL (original paper)arXiv, 2025
  • DeepSeek-R1 explainerHugging Face
Up NextSuggested: Continues the theme of AI / ML Papers

Software 2.0

"Neural networks are a new software paradigm where code is learned from data, not written by hand."

Andrej Karpathy · EssayContinue→
Listen
0 / 5