Why do gradient descent, momentum, and Adam take different paths from the same starting point? This experiment uses a two-dimensional quadratic with a known minimum and a calculable stability boundary. The aim is to connect parameter movement to gradients and optimizer state, not to declare a universally best optimizer.
Scope: a fixed objective, float64, start [2.2, -2.0], and 12 deterministic updates. There is no sampling, neural network, or GPU. The program has been checked against all 39 rows behind the original contour plot. Updated 2026-09-08 to correct the previous code/figure configuration mismatch and unsupported claims about monotonic Adam updates and mandatory warmup.
The Run notes above refer to the older deep-learning-math-lab package for the original contour plot. The added state traces, negative controls, and checks use the separate optimizer-audit-lab package linked in section 5. Use each package’s own working directory and commands.
1. Calculate Curvature Before Interpreting a Path
L(x, y) = 0.5 * (8*x*x + y*y) + 0.8*x*y
grad L = [8*x + 0.8*y, y + 0.8*x]
H = [[8, 0.8], [0.8, 1]]
theta* = [0, 0], L(theta*) = 0
The Hessian eigenvalues are approximately 0.909735 and 8.090265. Both are positive, so the origin is the unique minimum. The condition number is 8.892987: enough anisotropy to inspect, without calling it an extreme pathological case. The cross term means the principal directions are not exactly the coordinate axes.
The Hessian is constant everywhere. A turn in the trajectory therefore does not show an optimizer adapting to changing curvature; the current gradient and stored state are changing. This is deterministic gradient descent, not a minibatch SGD experiment.

| Method | Fixed configuration and update 12 |
|---|---|
| GD | lr=0.08; final [0.101159, -0.896491]; loss=0.370230. No loss-increasing updates in this window. |
| Momentum | lr=0.045, beta=0.75, zero initial velocity; final [0.022978, -0.122060]; loss=0.007317. Loss increases at updates 3, 4, 8, and 9. |
| Adam | lr=0.28, betas=(0.9, 0.999), eps=1e-8, zero initial moments with bias correction; final [-0.529374, 0.462135]; loss=1.032020. |
These are the existing illustration’s settings, not best configurations found with equal tuning budgets. Momentum’s lowest final loss describes this window only. Generalization, runtime, and other objectives were not tested.
2. The First Step and the Stability Boundary
theta0 = [2.2, -2.0]
g0 = [16.0, -0.24]
theta1 = theta0 - 0.08 * g0 = [0.92, -1.9808]
L(theta0) = 17.84
L(theta1) = 3.88951552
For this positive-definite quadratic, theta_t = (I - lr*H)^t theta0. An eigenvector component is multiplied by 1 - lr*lambda at each update. A negative factor means crossing zero in that direction; an absolute factor greater than one means growing error. Crossing the minimum is not the same as diverging.
The high-curvature component first changes sign above 1/lambda_max = 0.123605. Convergence from every initial point requires the strict interval 0 < lr < 2/lambda_max = 0.247211. At the upper boundary that component does not decay. Beyond it, the nonzero high-curvature component of this starting point grows.
| GD rate | Eigenmode factors and measured behavior |
|---|---|
| 0.08 | Low/high curvature factors: 0.927221 / 0.352779. Both positive; loss at update 50 is 0.00118688. No alternating eigenmode signs. |
| 0.20 | Factors: 0.818053 / -0.618053. The high-curvature component alternates with shrinking amplitude; loss at update 50 is 4.31e-9. |
| 0.26 | Factors: 0.763469 / -1.103469. The high-curvature amplitude grows; loss at update 50 is 293959.665. |
These three 50-update checks are separate from the 12-update comparison above. At rate 0.20, coordinate oscillation and decreasing loss can coexist. Do not confuse those two observations.
3. What Momentum Stores
velocity_t = beta * velocity_(t-1) + gradient_t
theta_t = theta_(t-1) - lr * velocity_t
Here velocity is a gradient accumulator before learning-rate scaling. Its inertia can carry parameters across the valley after the current gradient changes direction. The four loss increases in the recorded path do not support a claim that momentum removes all oscillation.
Another convention uses u_t = beta*u_(t-1) + lr*gradient_t; theta -= u_t. These are equivalent for a constant rate and compatible state initialization, but generally not under a schedule. The package includes a separate scalar check with constant gradient 1, beta=0.9, and zero initial state: rates [0.1, 0.01] give final parameters -0.119 and -0.200, respectively. This check is not part of the quadratic trajectory.
When porting an implementation, check where the rate enters the recurrence, dampening, and first-buffer initialization. The PyTorch SGD documentation explains these conventions. No PyTorch parity test was run here.
4. Adam Does Not Guarantee Monotonic Loss
m_t = beta1*m_(t-1) + (1-beta1)*g_t
v_t = beta2*v_(t-1) + (1-beta2)*g_t*g_t
m_hat = m_t / (1-beta1**t)
v_hat = v_t / (1-beta2**t)
theta_t = theta_(t-1) - lr*m_hat/(sqrt(v_hat)+eps)
The v state tracks an exponential average of squared gradients, an uncentered second moment, not variance after subtracting the squared mean. This convention follows Algorithm 1 in the Adam paper. It scales coordinates using gradient history; it does not directly invert the Hessian.
| Adam update | Recorded coordinates and loss |
|---|---|
| 1 | [1.920000, -1.720000]; loss=13.582880. Both coordinate changes are close to 0.28 in magnitude despite initial gradients of 16 and -0.24. |
| 9 | [-0.084274, 0.049923]; loss=0.026289, the minimum observed within these 12 updates. |
| 10 to 12 | Loss rises through 0.250965, 0.619280, and 1.032020. Continuing moves away from the earlier minimum; this short rebound does not establish long-run divergence. |
Starting from zero moments, bias correction gives m_hat=g and v_hat=g*g on the first update. The subtracted vector is lr*g/(abs(g)+eps), with each magnitude bounded by lr. It is approximately [0.28, -0.28] here. Deliberately omitting bias correction instead gives [0.885438, -0.885437]. That negative control exposes an implementation difference, not an inevitable exploding first step in correct Adam.
No warmup was used in this example. It refutes a universal requirement for warmup, not its usefulness in large-model training. The first-update bound must not be asserted for every subsequent update.
5. The Complete Executable Program
This is the same optimizer_example.py shipped with both localized articles. It copies and converts input to float64, rejects unknown methods and invalid hyperparameters, and returns parameter paths and update states. It supports this two-dimensional objective only, not a general training API.
Expand the complete NumPy program
"""A fixed quadratic experiment, not a training-framework implementation."""
import numpy as np
H = np.array([[8.0, 0.8], [0.8, 1.0]])
START = np.array([2.2, -2.0])
CONFIGS = {
"gd": {"lr": 0.08},
"momentum": {"lr": 0.045, "beta": 0.75},
"adam": {"lr": 0.28, "beta1": 0.9, "beta2": 0.999, "eps": 1e-8},
}
def loss(theta):
x, y = theta
return float(0.5 * (8 * x * x + y * y) + 0.8 * x * y)
def grad(theta):
x, y = theta
return np.array([8 * x + 0.8 * y, y + 0.8 * x])
def scalar(value, name, lower=0, upper=np.inf, strict=True):
a = np.asarray(value)
if a.ndim or a.dtype.kind not in "iuf" or not np.isfinite(a):
raise ValueError(name + " must be a finite real scalar")
value = float(a)
if value >= upper or (value <= lower if strict else value < lower):
raise ValueError(name + " outside its valid interval")
return value
def run_optimizer(method, start, lr, steps=12, beta=0.75,
beta1=0.9, beta2=0.999, eps=1e-8):
if method not in CONFIGS:
raise ValueError("method must be gd, momentum, or adam")
raw = np.asarray(start)
if raw.shape != (2,) or raw.dtype.kind not in "iuf":
raise ValueError("start must be a real two-vector")
theta = raw.astype(np.float64, copy=True)
if not np.isfinite(theta).all():
raise ValueError("start must be finite")
if isinstance(steps, (bool, np.bool_)) or not isinstance(steps, (int, np.integer)) or steps < 0:
raise ValueError("steps must be a nonnegative integer")
lr, eps = scalar(lr, "lr"), scalar(eps, "eps")
beta = scalar(beta, "beta", upper=1, strict=False)
beta1 = scalar(beta1, "beta1", upper=1, strict=False)
beta2 = scalar(beta2, "beta2", upper=1, strict=False)
velocity = np.zeros(2) if method == "momentum" else None
m, v = (np.zeros(2), np.zeros(2)) if method == "adam" else (None, None)
path, records = [theta.copy()], []
for t in range(1, steps + 1):
before, g = theta.copy(), grad(theta)
m_hat = v_hat = None
if method == "gd":
update = lr * g
elif method == "momentum":
velocity = beta * velocity + g
update = lr * velocity
else:
m = beta1 * m + (1 - beta1) * g
v = beta2 * v + (1 - beta2) * g * g
m_hat, v_hat = m / (1 - beta1 ** t), v / (1 - beta2 ** t)
update = lr * m_hat / (np.sqrt(v_hat) + eps)
theta = theta - update
if not np.isfinite(theta).all() or not np.isfinite(loss(theta)):
raise FloatingPointError("nonfinite iterate or loss")
record = {"step": t, "loss_before": loss(before), "loss_after": loss(theta)}
for name, value in (("theta_before", before), ("gradient", g), ("update", update),
("theta_after", theta), ("velocity", velocity), ("m", m),
("v", v), ("m_hat", m_hat), ("v_hat", v_hat)):
record[name] = None if value is None else value.copy()
records.append(record)
path.append(theta.copy())
return np.array(path), records
if __name__ == "__main__":
for method, config in CONFIGS.items():
path, _ = run_optimizer(method, START, **config)
print(f"{method}: theta={path[-1]}, loss={loss(path[-1]):.6f}")
Download the program, CSV traces, checks, and plotting script. The README includes isolated-environment installation steps. Numerical checks need only NumPy; plotting additionally needs Matplotlib.
python optimizer_example.py
python audit_optimizers.py --output reproduced
python plot_optimizers.py --output reproduced/loss-by-update.png
6. Plot the Same Data, Including the Rebound

paths.csv contains 39 rows: each method’s initial point and 12 updates. states.csv contains 36 update records. Its gradient is evaluated at theta_before; update is the subtracted vector; theta_after is the resulting point. Blank state fields mean not applicable, not an assumed zero value.
7. How the Numbers Were Checked
| Check | Measured result and scope |
|---|---|
| Legacy trace | All 39 new path rows match the existing CSV at six decimal places. The old archive and contour image were not overwritten. |
| GD recurrence | 51 states per rate, including the initial point, at three rates compared with independent matrix powers. Maximum coordinate error: about 1.14e-13. |
| Momentum recurrence | 13 states compared with powers of an independent 4-by-4 transition matrix. Maximum coordinate error: about 8.88e-16. |
| Gradient | Six components at three fixed points checked by central differences with step 1e-5. Maximum absolute error: about 1.40e-10. |
| Negative control and inputs | Missing Adam bias correction changes the first update. Seventeen invalid inputs are rejected; integer inputs, zero updates, repeatability, and caller-array preservation are also checked. |
| Environment | Python 3.13.9, NumPy 2.3.5, float64. No GPU, training dataset, timing measurement, or framework parity test. |
audit.json records the full values, settings, and source SHA-256 hashes. Matrix powers and finite differences provide different computational checks, not a proof for arbitrary inputs. Last-bit results may differ across numerical environments.
8. Memory, Decay, and Training Boundaries
Define the accounting before comparing memory. For P parameters, with parameters, gradients, and states all using s bytes per element, count raw array payloads only:
| Method | Array payload, not peak device memory |
|---|---|
| GD/SGD without momentum | Parameters and gradients total 2Ps, with no parameter-sized history buffer. Here P=2 and s=8: 32 bytes. |
| Momentum | One extra Ps buffer gives 3Ps total: 48 bytes here. |
| Adam | Two extra Ps states, m and v, give 4Ps total: 64 bytes here. Optimizer state does not “triple”; the momentum-free baseline has no such history state. |
This excludes trace copies, temporary arrays, Python objects, activations, mixed-precision master weights, and allocator overhead. The educational program saves per-update copies for inspection, so its actual process memory exceeds this accounting.
For decay, isolate one step with fresh Adam state, zero objective gradient, parameters [2, -0.5], lr=0.1, decay coefficient 0.1, and eps=1e-8:
Coupled L2: g = 0.1 * theta
theta_next = theta - 0.1*g/(abs(g)+1e-8)
= [1.900000005, -0.400000020]
Decoupled decay:
theta_next = (1 - 0.1*0.1) * theta
= [1.980000000, -0.495000000]
This one-step algebra isolates decay semantics; it is not a complete AdamW performance test. Passing an L2 gradient through adaptive normalization differs from shrinking parameters directly. See the decoupled weight decay paper for the distinction. The example does not prescribe a unique optimizer for every Transformer.
9. Use the Experiment for Debugging
First verify a single update on a fixed objective and starting point. Then compare gradients, states, and parameters throughout the trajectory, not just final loss. For alternating GD coordinates, inspect eigenmode factor magnitudes to separate convergent oscillation from divergence. For an Adam rebound, retain the minimum’s step and the full state trace, then vary the rate, budget, or initialization one at a time.
Real-model comparisons additionally need controlled splits, seeds, batch sizes, tuning budgets, and validation metrics. This small experiment teaches update-rule debugging, not generalization. The next article on convolution and receptive fields connects local operations to matrix implementations and checkable numerical examples.