DeepSeek-R1: Incentivizing Reasoning Capability via RL
DeepSeek · 2025
"Reinforcement learning can elicit strong reasoning in LLMs at a fraction of frontier training cost."
DeepSeek, a Chinese AI lab, released an open-source reasoning model that matched top proprietary 'reasoning' models (like OpenAI's o1) on many benchmarks, while claiming a dramatically lower training and inference cost. The paper's core claim is that strong step-by-step reasoning ability can be taught to a language model largely through reinforcement learning — rewarding the model for correct final answers and consistent reasoning — rather than primarily through expensive, human-curated supervised fine-tuning data, which had been the more common assumption.
Read more about the topic
Software 2.0
"Neural networks are a new software paradigm where code is learned from data, not written by hand."