Gradient descent variants that fix slow model training
A WooCommerce recommendation engine kept timing out the server, and the weights oscillated instead of settling. Swapping the basic update rule for Momentum and then Adam is what fixed it, and this explains why adaptive step sizes matter on messy data.