Gradient Descent and Optimizer Geometry: Momentum, Adam, and Loss Surfaces
Gradient Descent and Optimizer Geometry: Momentum, Adam, and Loss Surfaces
Search
Ask the AI

Gradient Descent and Optimizer Geometry: Momentum, Adam, and Loss Surfaces

Why do gradient descent, momentum, and Adam take different paths from the same starting point? This experiment uses a two-dimensional quadratic with a known minimum and a calculable stability boundary. The aim is to connect parameter movement to gradients and optimizer state, not to declare a universally best optimizer.

Scope: a fixed objective, float64, start [2.2, -2.0], and 12 deterministic updates. There is no sampling, neural network, or GPU. The program has been checked against all 39 rows behind the original contour plot. Updated 2026-09-08 to correct the previous code/figure configuration mismatch and unsupported claims about monotonic Adam updates and mandatory warmup.

The Run notes above refer to the older deep-learning-math-lab package for the original contour plot. The added state traces, negative controls, and checks use the separate optimizer-audit-lab package linked in section 5. Use each package’s own working directory and commands.

1. Calculate Curvature Before Interpreting a Path

L(x, y) = 0.5 * (8*x*x + y*y) + 0.8*x*y
grad L = [8*x + 0.8*y, y + 0.8*x]
H = [[8, 0.8], [0.8, 1]]
theta* = [0, 0], L(theta*) = 0

The Hessian eigenvalues are approximately 0.909735 and 8.090265. Both are positive, so the origin is the unique minimum. The condition number is 8.892987: enough anisotropy to inspect, without calling it an extreme pathological case. The cross term means the principal directions are not exactly the coordinate axes.

The Hessian is constant everywhere. A turn in the trajectory therefore does not show an optimizer adapting to changing curvature; the current gradient and stored state are changing. This is deterministic gradient descent, not a minibatch SGD experiment.

Twelve updates of GD, momentum, and Adam from the same point
The original contour plot is retained. Each path starts at [2.2, -2] and takes 12 updates with the configuration below. Path shape is not an optimizer ranking.
Method Fixed configuration and update 12
GD lr=0.08; final [0.101159, -0.896491]; loss=0.370230. No loss-increasing updates in this window.
Momentum lr=0.045, beta=0.75, zero initial velocity; final [0.022978, -0.122060]; loss=0.007317. Loss increases at updates 3, 4, 8, and 9.
Adam lr=0.28, betas=(0.9, 0.999), eps=1e-8, zero initial moments with bias correction; final [-0.529374, 0.462135]; loss=1.032020.

These are the existing illustration’s settings, not best configurations found with equal tuning budgets. Momentum’s lowest final loss describes this window only. Generalization, runtime, and other objectives were not tested.

2. The First Step and the Stability Boundary

theta0 = [2.2, -2.0]
g0 = [16.0, -0.24]
theta1 = theta0 - 0.08 * g0 = [0.92, -1.9808]
L(theta0) = 17.84
L(theta1) = 3.88951552

For this positive-definite quadratic, theta_t = (I - lr*H)^t theta0. An eigenvector component is multiplied by 1 - lr*lambda at each update. A negative factor means crossing zero in that direction; an absolute factor greater than one means growing error. Crossing the minimum is not the same as diverging.

The high-curvature component first changes sign above 1/lambda_max = 0.123605. Convergence from every initial point requires the strict interval 0 < lr < 2/lambda_max = 0.247211. At the upper boundary that component does not decay. Beyond it, the nonzero high-curvature component of this starting point grows.

GD rate Eigenmode factors and measured behavior
0.08 Low/high curvature factors: 0.927221 / 0.352779. Both positive; loss at update 50 is 0.00118688. No alternating eigenmode signs.
0.20 Factors: 0.818053 / -0.618053. The high-curvature component alternates with shrinking amplitude; loss at update 50 is 4.31e-9.
0.26 Factors: 0.763469 / -1.103469. The high-curvature amplitude grows; loss at update 50 is 293959.665.

These three 50-update checks are separate from the 12-update comparison above. At rate 0.20, coordinate oscillation and decreasing loss can coexist. Do not confuse those two observations.

3. What Momentum Stores

velocity_t = beta * velocity_(t-1) + gradient_t
theta_t = theta_(t-1) - lr * velocity_t

Here velocity is a gradient accumulator before learning-rate scaling. Its inertia can carry parameters across the valley after the current gradient changes direction. The four loss increases in the recorded path do not support a claim that momentum removes all oscillation.

Another convention uses u_t = beta*u_(t-1) + lr*gradient_t; theta -= u_t. These are equivalent for a constant rate and compatible state initialization, but generally not under a schedule. The package includes a separate scalar check with constant gradient 1, beta=0.9, and zero initial state: rates [0.1, 0.01] give final parameters -0.119 and -0.200, respectively. This check is not part of the quadratic trajectory.

When porting an implementation, check where the rate enters the recurrence, dampening, and first-buffer initialization. The PyTorch SGD documentation explains these conventions. No PyTorch parity test was run here.

4. Adam Does Not Guarantee Monotonic Loss

m_t = beta1*m_(t-1) + (1-beta1)*g_t
v_t = beta2*v_(t-1) + (1-beta2)*g_t*g_t
m_hat = m_t / (1-beta1**t)
v_hat = v_t / (1-beta2**t)
theta_t = theta_(t-1) - lr*m_hat/(sqrt(v_hat)+eps)

The v state tracks an exponential average of squared gradients, an uncentered second moment, not variance after subtracting the squared mean. This convention follows Algorithm 1 in the Adam paper. It scales coordinates using gradient history; it does not directly invert the Hessian.

Adam update Recorded coordinates and loss
1 [1.920000, -1.720000]; loss=13.582880. Both coordinate changes are close to 0.28 in magnitude despite initial gradients of 16 and -0.24.
9 [-0.084274, 0.049923]; loss=0.026289, the minimum observed within these 12 updates.
10 to 12 Loss rises through 0.250965, 0.619280, and 1.032020. Continuing moves away from the earlier minimum; this short rebound does not establish long-run divergence.

Starting from zero moments, bias correction gives m_hat=g and v_hat=g*g on the first update. The subtracted vector is lr*g/(abs(g)+eps), with each magnitude bounded by lr. It is approximately [0.28, -0.28] here. Deliberately omitting bias correction instead gives [0.885438, -0.885437]. That negative control exposes an implementation difference, not an inevitable exploding first step in correct Adam.

No warmup was used in this example. It refutes a universal requirement for warmup, not its usefulness in large-model training. The first-update bound must not be asserted for every subsequent update.

5. The Complete Executable Program

This is the same optimizer_example.py shipped with both localized articles. It copies and converts input to float64, rejects unknown methods and invalid hyperparameters, and returns parameter paths and update states. It supports this two-dimensional objective only, not a general training API.

Expand the complete NumPy program
"""A fixed quadratic experiment, not a training-framework implementation."""
import numpy as np

H = np.array([[8.0, 0.8], [0.8, 1.0]])
START = np.array([2.2, -2.0])
CONFIGS = {
    "gd": {"lr": 0.08},
    "momentum": {"lr": 0.045, "beta": 0.75},
    "adam": {"lr": 0.28, "beta1": 0.9, "beta2": 0.999, "eps": 1e-8},
}


def loss(theta):
    x, y = theta
    return float(0.5 * (8 * x * x + y * y) + 0.8 * x * y)


def grad(theta):
    x, y = theta
    return np.array([8 * x + 0.8 * y, y + 0.8 * x])


def scalar(value, name, lower=0, upper=np.inf, strict=True):
    a = np.asarray(value)
    if a.ndim or a.dtype.kind not in "iuf" or not np.isfinite(a):
        raise ValueError(name + " must be a finite real scalar")
    value = float(a)
    if value >= upper or (value <= lower if strict else value < lower):
        raise ValueError(name + " outside its valid interval")
    return value


def run_optimizer(method, start, lr, steps=12, beta=0.75,
                  beta1=0.9, beta2=0.999, eps=1e-8):
    if method not in CONFIGS:
        raise ValueError("method must be gd, momentum, or adam")
    raw = np.asarray(start)
    if raw.shape != (2,) or raw.dtype.kind not in "iuf":
        raise ValueError("start must be a real two-vector")
    theta = raw.astype(np.float64, copy=True)
    if not np.isfinite(theta).all():
        raise ValueError("start must be finite")
    if isinstance(steps, (bool, np.bool_)) or not isinstance(steps, (int, np.integer)) or steps < 0:
        raise ValueError("steps must be a nonnegative integer")
    lr, eps = scalar(lr, "lr"), scalar(eps, "eps")
    beta = scalar(beta, "beta", upper=1, strict=False)
    beta1 = scalar(beta1, "beta1", upper=1, strict=False)
    beta2 = scalar(beta2, "beta2", upper=1, strict=False)
    velocity = np.zeros(2) if method == "momentum" else None
    m, v = (np.zeros(2), np.zeros(2)) if method == "adam" else (None, None)
    path, records = [theta.copy()], []
    for t in range(1, steps + 1):
        before, g = theta.copy(), grad(theta)
        m_hat = v_hat = None
        if method == "gd":
            update = lr * g
        elif method == "momentum":
            velocity = beta * velocity + g
            update = lr * velocity
        else:
            m = beta1 * m + (1 - beta1) * g
            v = beta2 * v + (1 - beta2) * g * g
            m_hat, v_hat = m / (1 - beta1 ** t), v / (1 - beta2 ** t)
            update = lr * m_hat / (np.sqrt(v_hat) + eps)
        theta = theta - update
        if not np.isfinite(theta).all() or not np.isfinite(loss(theta)):
            raise FloatingPointError("nonfinite iterate or loss")
        record = {"step": t, "loss_before": loss(before), "loss_after": loss(theta)}
        for name, value in (("theta_before", before), ("gradient", g), ("update", update),
                            ("theta_after", theta), ("velocity", velocity), ("m", m),
                            ("v", v), ("m_hat", m_hat), ("v_hat", v_hat)):
            record[name] = None if value is None else value.copy()
        records.append(record)
        path.append(theta.copy())
    return np.array(path), records


if __name__ == "__main__":
    for method, config in CONFIGS.items():
        path, _ = run_optimizer(method, START, **config)
        print(f"{method}: theta={path[-1]}, loss={loss(path[-1]):.6f}")

Download the program, CSV traces, checks, and plotting script. The README includes isolated-environment installation steps. Numerical checks need only NumPy; plotting additionally needs Matplotlib.

python optimizer_example.py
python audit_optimizers.py --output reproduced
python plot_optimizers.py --output reproduced/loss-by-update.png

6. Plot the Same Data, Including the Rebound

Measured loss over twelve updates, including momentum and Adam rebounds
Measured curves using the same start and settings as the contour plot, with a logarithmic loss axis. Adam’s rebound after update 9 is retained. This replaces the earlier schematic animation that could not be checked step by step.

paths.csv contains 39 rows: each method’s initial point and 12 updates. states.csv contains 36 update records. Its gradient is evaluated at theta_before; update is the subtracted vector; theta_after is the resulting point. Blank state fields mean not applicable, not an assumed zero value.

7. How the Numbers Were Checked

Check Measured result and scope
Legacy trace All 39 new path rows match the existing CSV at six decimal places. The old archive and contour image were not overwritten.
GD recurrence 51 states per rate, including the initial point, at three rates compared with independent matrix powers. Maximum coordinate error: about 1.14e-13.
Momentum recurrence 13 states compared with powers of an independent 4-by-4 transition matrix. Maximum coordinate error: about 8.88e-16.
Gradient Six components at three fixed points checked by central differences with step 1e-5. Maximum absolute error: about 1.40e-10.
Negative control and inputs Missing Adam bias correction changes the first update. Seventeen invalid inputs are rejected; integer inputs, zero updates, repeatability, and caller-array preservation are also checked.
Environment Python 3.13.9, NumPy 2.3.5, float64. No GPU, training dataset, timing measurement, or framework parity test.

audit.json records the full values, settings, and source SHA-256 hashes. Matrix powers and finite differences provide different computational checks, not a proof for arbitrary inputs. Last-bit results may differ across numerical environments.

8. Memory, Decay, and Training Boundaries

Define the accounting before comparing memory. For P parameters, with parameters, gradients, and states all using s bytes per element, count raw array payloads only:

Method Array payload, not peak device memory
GD/SGD without momentum Parameters and gradients total 2Ps, with no parameter-sized history buffer. Here P=2 and s=8: 32 bytes.
Momentum One extra Ps buffer gives 3Ps total: 48 bytes here.
Adam Two extra Ps states, m and v, give 4Ps total: 64 bytes here. Optimizer state does not “triple”; the momentum-free baseline has no such history state.

This excludes trace copies, temporary arrays, Python objects, activations, mixed-precision master weights, and allocator overhead. The educational program saves per-update copies for inspection, so its actual process memory exceeds this accounting.

For decay, isolate one step with fresh Adam state, zero objective gradient, parameters [2, -0.5], lr=0.1, decay coefficient 0.1, and eps=1e-8:

Coupled L2: g = 0.1 * theta
theta_next = theta - 0.1*g/(abs(g)+1e-8)
           = [1.900000005, -0.400000020]

Decoupled decay:
theta_next = (1 - 0.1*0.1) * theta
           = [1.980000000, -0.495000000]

This one-step algebra isolates decay semantics; it is not a complete AdamW performance test. Passing an L2 gradient through adaptive normalization differs from shrinking parameters directly. See the decoupled weight decay paper for the distinction. The example does not prescribe a unique optimizer for every Transformer.

9. Use the Experiment for Debugging

First verify a single update on a fixed objective and starting point. Then compare gradients, states, and parameters throughout the trajectory, not just final loss. For alternating GD coordinates, inspect eigenmode factor magnitudes to separate convergent oscillation from divergence. For an Adam rebound, retain the minimum’s step and the full state trace, then vary the rate, budget, or initialization one at a time.

Real-model comparisons additionally need controlled splits, seeds, batch sizes, tuning budgets, and validation metrics. This small experiment teaches update-rule debugging, not generalization. The next article on convolution and receptive fields connects local operations to matrix implementations and checkable numerical examples.

Leave a Reply

Scroll down