Deep Dive into KL Estimators in RL: k3 Gradient Is Wrong, Sequence KL Gradient Is Wrong Too—Corrected Derivation
TutorialSource: xAuthor: HeMuyu0327Hotness: 137Published Jul 25, 2026
The author uncovers that popular KL estimators in RL losses, such as the k3 estimator used in DeepSeek, have incorrect gradients: k3's gradient actually estimates forward KL instead of reverse KL. Sequence-level KL also has a gradient mismatch. A blog post provides correct definitions and computational tricks to reduce complexity using .cumsum and a stop-gradient hack.
- reinforcement learning
- KL divergence
- gradient
- estimator
Comments
Log in to comment
No comments yet. Be the first.