← AI News

Deep Dive into KL Estimators in RL: k3 Gradient Is Wrong, Sequence KL Gradient Is Wrong Too—Corrected Derivation

TutorialSource: xAuthor: HeMuyu0327Hotness: 137Published Jul 25, 2026

The author uncovers that popular KL estimators in RL losses, such as the k3 estimator used in DeepSeek, have incorrect gradients: k3's gradient actually estimates forward KL instead of reverse KL. Sequence-level KL also has a gradient mismatch. A blog post provides correct definitions and computational tricks to reduce complexity using .cumsum and a stop-gradient hack.

  • reinforcement learning
  • KL divergence
  • gradient
  • estimator
View source →

Comments

Log in to comment

No comments yet. Be the first.