Skip to content
← Home
Contemporary · Paper

DeepSeek-R1: Incentivizing Reasoning Capability via RL

DeepSeek · 2025

"Reinforcement learning can elicit strong reasoning in LLMs at a fraction of frontier training cost."

The idea

DeepSeek, a Chinese AI lab, released an open-source reasoning model that matched top proprietary 'reasoning' models (like OpenAI's o1) on many benchmarks, while claiming a dramatically lower training and inference cost. The paper's core claim is that strong step-by-step reasoning ability can be taught to a language model largely through reinforcement learning — rewarding the model for correct final answers and consistent reasoning — rather than primarily through expensive, human-curated supervised fine-tuning data, which had been the more common assumption.

Why it works
The takeaway — recall it first
Further reading

Read more about the topic

Up NextSuggested: Continues the theme of AI / ML Papers

Software 2.0

"Neural networks are a new software paradigm where code is learned from data, not written by hand."

Andrej Karpathy · EssayContinue→
Listen
0 / 2