Divergence and loss · Entropy

Cross-entropy descent

Gradient descent that nudges a model's logits by −lr·(q − p) drives the cross-entropy H(P,Q) = H(P) + D(P‖Q) down toward the floor H(P) and never below it, because the gap above the floor is the KL divergence, which is zero only when Q = P.

Builds on KL divergence. Taught in Bits & Surprise.

Drag the model's bars or press train and watch real gradient descent through a softmax fit a four-class truth: the loss curve falls to the dashed H(P) line and stops there. A learning-rate slider sets the step size; single-step to see each update.

The claim above is what this widget demonstrates. Disagree with it?

Where next