Kuaishou's SRPO Cuts LLM Reasoning Training to a Tenth the Steps, Matches DeepSeek-R1-Zero
Kwaipilot's new reinforcement learning framework matches DeepSeek-R1-Zero on math and code benchmarks simultaneously—using just one-tenth the compute. The team open-sourced both the method and the model.