AIToday
Daily Dose of Data SciencePublished: Apr 28, 2026, 10:00 JST1 min read

DeepSeek's R1 uses verifiable rewards and simplified training to match human-preference-based reasoning models with a fraction of the memory

DeepSeek's R1 uses verifiable rewards and simplified training to match human-preference-based reasoning models with a fraction of the memory

3 Key Points

  1. In January 2025, DeepSeek released R1 using RLVR (Reinforcement Learning with Verifiable Rewards) instead of the human-preference training pipeline. For math problems, the verifier checked if the model's answer matched the known solution; for code, a compiler ran the output and returned pass or fail (binary rewards: 1 for correct, 0 for wrong).

  2. DeepSeek's approach uses GRPO (Group Relative Policy Optimization), which removes the critic model and learned reward model entirely. Instead of four full-size models in memory (policy, reference policy, reward model, critic), the system requires just two (the policy being trained and a reference copy for KL regularization), cutting memory demands substantially.

  3. DeepSeek R1-Zero, trained with GRPO and verifiable rewards with no supervised fine-tuning, went from 15.6% to 77.9% on AIME 2024 math problems; with majority voting, it hit 86.7%, matching OpenAI's o1. The model developed self-verification, reflection, and chain-of-thought reasoning purely from the binary signal.

Ask the AI about this article →

Daily Dose of Data ScienceRead Original Article

Get AI news like this every morning

For example, today's edition would include:

  • World Labs unveils Atlas, a 3D world model from one imageSiliconANGLE AI · 1h ago
  • TCL CSOT bets on InP laser chips as supply tightensDIGITIMES Asia · 1h ago
  • Google launches AI image tool Google PicsITmedia AI+ · 1h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Next articleApplied Intuition, a $15B physical AI company, discusses deployment of autonomous systems across vehicles, mining, construction, agriculture, and defense