The useful question when choosing a deep learning framework is not “which one is strongest?” It is: what does each implementation compute when the model, parameters, inputs, and objective are the same? Two programs that both train are not a controlled comparison if their weight layout or loss reduction differs.
Revised 2026-09-08: this tutorial now includes a measured PyTorch / TensorFlow CPU comparison. It corrects earlier blanket claims about GPU speedups, clearing gradients on every minibatch, and inferring checkpoint contents from a filename. No GPU benchmark or production experience is claimed here.
1. What Frameworks Provide, and What They Cannot Decide
Frameworks provide tensor operations, automatic differentiation, layers, optimizers, and device execution. Autodiff differentiates the computation you actually execute. It cannot decide whether labels match, a denominator is appropriate, or data leaked across a split. An incorrect objective can have correctly computed derivatives.
We will use one linear layer with three samples, two inputs, and three outputs. Samples occupy rows, arithmetic is float64 on CPU, and the objective averages all nine squared output errors. This differs from the layout with samples in columns and half-squared sample-mean objective in the matrix calculus tutorial; migrating between the examples requires an explicit conversion.
X = [[1, 2], [-1, 0.5], [0.25, -2]] # (B=3, D=2)
W = [[0.2,-0.4], [0.7,0.3], [-0.3,0.5]] # (M=3, D=2)
b = [0.05, -0.1, 0.1]
Y = [[0.4,-0.2,0.1], [-0.5,0.5,-0.3], [0.2,0,0.7]]
prediction = X @ W.T + b # (3, 3)
loss = np.mean((prediction - Y)**2)
The first prediction row is [-0.55, 1.20, 0.80]. The three row-wise sums of squared errors are 3.3525, 2.2475, and 3.57125. Their mean over nine entries is exactly 7337/7200, approximately 1.019027778. These fixed numbers make a comparison interpretable.
2. Linear and Dense Store Weights Differently
| Check | Contract in this example |
|---|---|
| Input/output | Both implementations take a 3×2 input and produce a 3×3 output. Independently randomized initializations are not compared. |
| PyTorch | nn.Linear(2, 3) stores a 3×2 weight and computes X @ weight.T + bias. |
| TensorFlow | Dense(3) stores a 2×3 kernel and computes X @ kernel + bias. Assign kernel = W.T, building the layer before setting its weights. |
| Gradients | Transpose the TensorFlow kernel gradient back to the W layout before comparing. The bias gradient has length 3 and the input gradient is 3×2. |
The API contracts are documented in PyTorch Linear and TensorFlow Dense. A non-square weight is intentional: confusing 3×2 and 2×3 is visible to shape checks. A mistaken transpose of a square matrix can preserve shape while changing the computation.
3. A Complete, Runnable Comparison
Download the framework contract lab and run these commands from the extracted directory. The measured environment was arm64 macOS, Python 3.13.9, NumPy 2.3.5, PyTorch 2.14.0, TensorFlow 2.21.0, and Keras 3.15.1. Framework downloads are substantial; no training dataset download is required.
python3 -m venv .venv
.venv/bin/python -m pip install -r requirements.txt
.venv/bin/python compare_layers.py
.venv/bin/python audit_frameworks.py --output reproduced
The success markers are FRAMEWORK_COMPARE_OK and FRAMEWORK_AUDIT_OK. On another platform, first check compatible wheels for your Python version and architecture. These pins describe the tested environment, not a universal installation command for every device.
Expand the complete compare_layers.py program
"""Compare one fixed CPU linear layer, not framework speed or general accuracy."""
import json
import numpy as np
import tensorflow as tf
import torch
def fixture():
return tuple(np.array(a, dtype=np.float64) for a in (
[[1.0, 2.0], [-1.0, 0.5], [0.25, -2.0]],
[[0.2, -0.4], [0.7, 0.3], [-0.3, 0.5]],
[0.05, -0.1, 0.1],
[[0.4, -0.2, 0.1], [-0.5, 0.5, -0.3], [0.2, 0.0, 0.7]]))
def torch_layer(W, b):
layer = torch.nn.Linear(2, 3, dtype=torch.float64, device='cpu')
with torch.no_grad():
layer.weight.copy_(torch.from_numpy(W))
layer.bias.copy_(torch.from_numpy(b))
return layer
def tensorflow_layer(W, b):
layer = tf.keras.layers.Dense(3, dtype='float64',
kernel_initializer='zeros', bias_initializer='zeros')
layer.build((None, 2))
layer.set_weights([W.T, b])
return layer
def run_torch(X, W, b, Y):
layer = torch_layer(W, b)
x = torch.tensor(X, dtype=torch.float64, requires_grad=True)
prediction = layer(x)
loss = ((prediction - torch.from_numpy(Y)) ** 2).mean()
loss.backward()
return {'prediction': prediction.detach().numpy(), 'loss': loss.item(),
'dW': layer.weight.grad.numpy(), 'db': layer.bias.grad.numpy(),
'dX': x.grad.numpy()}
def run_tensorflow(X, W, b, Y):
with tf.device('/CPU:0'):
layer = tensorflow_layer(W, b)
x = tf.Variable(X, dtype=tf.float64)
with tf.GradientTape() as tape:
prediction = layer(x)
loss = tf.reduce_mean((prediction - Y) ** 2)
dK, db, dX = tape.gradient(loss, [layer.kernel, layer.bias, x])
return {'prediction': prediction.numpy(), 'loss': float(loss.numpy()),
'dW': dK.numpy().T, 'db': db.numpy(), 'dX': dX.numpy()}
def main():
X, W, b, Y = fixture()
prediction = X @ W.T + b
error = prediction - Y
G = 2 * error / error.size
reference = {'prediction': prediction, 'loss': float(np.mean(error ** 2)),
'dW': G.T @ X, 'db': G.sum(axis=0), 'dX': G @ W}
report = {'loss_definition': 'mean over all 9 squared output errors',
'numpy_loss': reference['loss'], 'frameworks': {}}
for name, run in (('pytorch', run_torch), ('tensorflow', run_tensorflow)):
result = run(X, W, b, Y)
errors = {}
for key, expected in reference.items():
np.testing.assert_allclose(result[key], expected, atol=1e-12, rtol=1e-12)
errors[key] = float(np.max(np.abs(result[key] - expected)))
report['frameworks'][name] = {'loss': result['loss'], 'max_abs_errors': errors}
print(json.dumps(report, indent=2))
print('FRAMEWORK_COMPARE_OK')
if __name__ == '__main__':
main()
Both sets of parameters are explicitly overwritten. PyTorch’s backward() populates leaf tensors’ .grad fields; TensorFlow records the forward computation inside GradientTape and returns requested derivatives through gradient(). The input is tracked here too, so the check is not limited to weights and biases.
For the element-mean objective, G = 2E/(BM), dW = G.T @ X, db = G.sum(axis=0), and dX = G @ W. The program compares autodiff with these NumPy formulas. The companion audit separately uses fraction arithmetic and scalar loops, so agreement between frameworks is not the only reference.
4. Measured Results and Their Limits
| Test | Observed result |
|---|---|
| Exact reference | Each framework checks 9 predictions, 1 loss, 6 weight gradients, 3 bias gradients, and 6 input gradients: 50 numeric comparisons across both. Maximum absolute error against exact fraction values rounded to float64 was about 4.44e-16. Absolute and relative assertion tolerances were each 1e-12. |
| 20 updates | Both use plain SGD with learning rate 0.125. Training MSE moves from 1.019027778 to 0.712134108 after the first update and 0.041723010 after update 20. Maximum cross-framework prediction difference across the 21 recorded states was zero. |
| Save/reload | A PyTorch state_dict is loaded into a newly constructed compatible layer. A Keras model using built-in layers is saved as .keras and reloaded. Each reproduces its pre-save predictions on the fixed input with maximum difference zero. |
These results do not establish bitwise equality for all models, devices, or precisions. A decrease in training loss is not evidence of better held-out performance. This run has no classification dataset, independent test set, GPU timing, exported-runtime comparison, or optimizer-state resume test.

View the full-size experiment plot; per-coordinate comparison CSV; 21 training records; environment, source hashes, and audit JSON.
5. Three Counterexamples Worth Reproducing
Accumulation: clearing on every microbatch is not a universal rule
Recomputing the same batch and calling backward twice without updating parameters doubles PyTorch’s parameter gradients. Across the nine parameter entries in this run, the discrepancy from exactly twice the first gradient was zero. Clear gradients at the boundary of the effective batch you intend to optimize, not automatically at every possible subdivision.
Split three rows into microbatches of sizes 2 and 1. The correct objective is (2/3)L_two + (1/3)L_one. Averaging the two means equally overweights the last sample. In this experiment, the wrong gradient differs from the full-batch gradient by at most 0.277777778; weighting by sample count gives maximum difference zero. This equivalence uses the separable squared loss and identical model state here. It is not a general guarantee for cross-sample statistics, changing randomness, or other batch-coupled behavior.
Broadcasting: a lower loss can mean the labels are wrong
Replacing the 3×3 target with Y[:, :1], a 3×1 target, still executes in both framework expressions used here. The loss becomes approximately 0.607361111, smaller than the correct loss. It compares against repeated first-column labels, not an improved model. Assert exact prediction/target shape agreement before allowing any intentional broadcasting.
Missing derivatives: None is not a numeric zero
A constant TensorFlow input is not automatically watched as a differentiation source. In this run, the input derivative for an unwatched tf.constant(X) was None. Calling tape.watch(x) restored it and matched the PyTorch input gradient. None means the request did not produce that derivative; replacing it with zero can hide a tracking error. See the TensorFlow autodiff guide for additional cases.
6. Choose Against Project Requirements
To reproduce a published model, begin with its supported framework and locked dependencies, then run its smallest test. To integrate with a service, check the target runtime’s operators, dynamic shapes, and devices first. “Research uses A; industry uses B” is not a substitute for these constraints.
For learning, manually follow data selection, forward computation, loss, backward computation, and parameter update at least once. But fit() is not inherently an undebuggable black box: Keras supports custom training loops and custom layers and model subclasses. This experiment uses its TensorFlow backend; Keras is not limited to that backend.
Nor does TensorFlow mean “static only” or PyTorch mean “eager only.” TensorFlow 2 defaults to eager execution and can build graphs through tf.function. PyTorch compilation and export have their own contracts. TorchScript is marked deprecated in the official documentation; a new project should check torch.export and its target runtime rather than copy an old tracing command mechanically. Compilation and export were not exercised in this run.
7. Saving, Resuming, and Deploying Need Different Checks
.pt is an extension, not a guarantee about file contents. This audit explicitly uses torch.save(layer.state_dict(), ...), then reconstructs a compatible layer before loading. Other serialization choices have different contents and dependencies. See the PyTorch serialization notes; do not infer “no structure” solely from the suffix.
The Keras round trip uses built-in layers and does not cover custom-object registration; see Keras saving and serialization. Both tests load only files just created by the experiment. Loading options are not a promise that arbitrary external model files are trustworthy.
Resuming training also requires checking optimizer, scheduler, random-number, and data-iteration state. Deployment requires testing actual runtime inputs, outputs, operators, and numerical differences. A Python service can use a framework directly; ONNX conversion is not mandatory for every deployment. ONNX, GGUF, and hardware-specific artifacts are not interchangeable containers for every network.
Continue with browser ONNX deployment checks and quantization validation and sampling bias. These are further reading, not evidence that this experiment already validated those workflows.
8. No CUDA Is Not a Failed CPU Installation
A CPU is sufficient for this example. GPU speedups depend on workload size, operators, transfers, and synchronization; there is no measured “dozens of times faster” result here. A false torch.cuda.is_available() only describes CUDA access in the current process. It says neither that CPU execution is broken nor that every other accelerator backend is unavailable. Apple devices have a separate MPS backend.
Expand the device check with a guarded CUDA query
"""Report availability without querying a CUDA device when none is available."""
import json
import tensorflow as tf
import torch
def report():
cuda = torch.cuda.is_available()
mps = getattr(torch.backends, 'mps', None)
return {
'torch_version': torch.__version__,
'torch_cuda_build': torch.version.cuda,
'cuda_available': cuda,
'cuda_device': torch.cuda.get_device_name(0) if cuda else None,
'mps_built': bool(mps and mps.is_built()),
'mps_available': bool(mps and mps.is_available()),
'tensorflow_version': tf.__version__,
'tensorflow_gpus': [d.name for d in tf.config.list_physical_devices('GPU')],
'tutorial_execution_device': 'CPU',
}
if __name__ == '__main__':
print(json.dumps(report(), indent=2))
The recorded environment had CUDA unavailable, no CUDA build version, MPS built but unavailable to this process, and no TensorFlow GPU devices. Both CPU experiments still passed. This describes one environment, not every Mac of the same model. Diagnose installation or driver issues from the complete error and device context, not from one Boolean followed by a blanket reinstall.
The README includes the file map, optional plotting commands, and a complete dependency snapshot. After reproducing the results, deliberately change one contract: transpose the kernel incorrectly, change the loss denominator, or alter the microbatch sizes. Identify why the check changes before switching tools. That is the step from calling layers to diagnosing training behavior.