10.09.2026

Talk: On the Natural Gradient of the Evidence Lower Bound in non-cylindrical generative Models

by David Söding, Hamburg University of Technology - on Thursday, September 10th, 2026 at 11:00 am - in room 5.002, HIPone, Blohmstr. 15, 21079 Hamburg.

David is currently doing his Bachelor’s thesis on this topic, which has resulted in a joint publication with Nihat Ay and Adwait . 

Abstract:

Training a generative model means maximising the evidence, but that's intractable, so we maximise the ELBO instead and hope the two behave alike. How you take the step matters. Plain gradient descent measures distance in whatever parameters you happen to have chosen, so it's blind to the actual geometry of the distributions. Natural gradient descent instead uses the Fisher-Rao metric, measuring steps differently by how much the probability distribution actually changes. That difference turns out to matter for the ELBO-versus-evidence question, because the two optimisers map gradients onto the model manifold in different ways. Work at this institute has already shown that under NGD the two objectives have exactly the same gradient, but only for cylindrical models, a condition almost nothing satisfies in practice.

So my question is what happens once you drop this condition. Does the advantage survive in some weaker form? I tested this by simulating both optimisers on both objectives on small binary Bayesian networks, small enough that the evidence and the gap can still be computed exactly.