Gradient descent that nudges a model's logits by −lr·(q − p) drives the cross-entropy H(P,Q) = H(P) + D(P‖Q) down toward the floor H(P) and never below it, because the gap above the floor is the KL divergence, which is zero only when Q = P.
Drag the model's bars or press train and watch real gradient descent through a softmax fit a four-class truth: the loss curve falls to the dashed H(P) line and stops there. A learning-rate slider sets the step size; single-step to see each update.