CAST: Non-Privileged Clipped Asymmetric Self-Teaching with Advantage Flipping for GRPO
Published in arXiv preprint, 2026
Reinforcement learning with verifiable rewards (RLVR), especially Group Relative Policy Optimization (GRPO), often suffers from sparse outcome-level rewards and vanishing group-relative advantages. CAST introduces an answer-free self-teacher that provides dense token-level guidance while preserving the verifier-grounded GRPO objective, with bidirectional local advantage sign reversal for improved mathematical reasoning.
Recommended citation: Yang Li, Gongle Xue, Yijia Guo, Yuheng Yuan, Liwen Hu, Lei Ma. "CAST: Non-Privileged Clipped Asymmetric Self-Teaching with Advantage Flipping for GRPO." arXiv preprint arXiv:2606.00172, 2026.
Download Paper
