Advances in reinforcement learning and updates to video generation benchmarks
Technical discussions focused on implementations of Reinforcement Learning with Verifiable Rewards (RLVR) and Group Relative Policy Optimization (GRPO). Meanwhile, new video generation models entered public evaluation leaderboards.
3 independent accounts
7 posts
3,216 interactions
deepseek r1reinforcement learning with verifiable rewards