PyTorch Tensors & Autograd
Everything in PyTorch is built on two ideas. Tensors are n-dimensional arrays like NumPy's, but able to live on GPUs and other accelerators. Autograd is the engine that records operations on tensors and automatically computes gradients by backpropagation. Neural network layers, optimizers, and loss functions are all conveniences on top of these two primitives.
Understanding tensors (shapes, dtypes, devices, broadcasting, and views) prevents most day-to-day PyTorch bugs, like shape mismatches, silent broadcasting errors, and "expected all tensors to be on the same device" exceptions. Understanding autograd explains training: why you call zero_grad(), what .backward() does, and when to use no_grad().
TL;DR
- A tensor has a shape, a dtype (
float32,bfloat16,int64…), and a device (cpu,cuda,mps). - Operations follow NumPy-style broadcasting. Check shapes deliberately.
- Views (
view, slicing,transpose) share memory with the original;clone()copies. - Tensors with
requires_grad=Truerecord operations in a dynamic graph;loss.backward()fills.gradfields. - Gradients accumulate, so zero them each step (
optimizer.zero_grad()). - Use
torch.no_grad()ortorch.inference_mode()for evaluation and inference, and.detach()to cut tensors out of the graph.
Quick Example
Fitting a line with raw tensors and autograd, no nn modules:
The training loop page shows the idiomatic version with nn.Module and optimizers.
Core Concepts
Creating Tensors
Shapes, dtypes, and Devices
- Shape:
t.shape/t.size(). Reshape withreshape,view,unsqueeze/squeeze,permute, andflatten. - dtype:
float32is the default for floats.float16/bfloat16are used for mixed precision, andint64for class indices and labels. Mismatched dtypes cause errors, or silent upcasting. - Device:
t.to("cuda"),t.to(device),model.to(device). All tensors in an operation must be on the same device. Create tensors directly on the target device (device=device) to avoid copies.
Broadcasting
When shapes differ, PyTorch aligns them from the trailing dimension and expands size-1 dimensions:
Use keepdim=True in reductions and assert shapes in critical code.
Views vs Copies
view, slicing, transpose, permute, and expand return views that share storage: modifying one modifies the other. clone() makes an independent copy. Some operations need contiguous memory: view fails after transpose, while reshape (copying if needed) or .contiguous().view(...) works. In-place operations (methods ending in _, like add_) save memory but can break autograd if they overwrite values needed for gradients.
Autograd
When a tensor has requires_grad=True (all nn.Parameters do), each operation records a node in a dynamic computation graph (define-by-run). Calling .backward() on a scalar (usually the loss) traverses the graph in reverse, applying the chain rule, and accumulates gradients into each leaf tensor's .grad.
- Graphs are rebuilt every forward pass, so you can use Python control flow (loops, conditionals) freely.
- After
backward(), intermediate buffers are freed by default. Passretain_graph=Trueonly if you truly need a second backward pass. torch.autograd.grad(outputs, inputs)computes gradients without touching.grad, which is useful for gradient penalties and meta-learning.
Disabling Gradient Tracking
Best Practices
Be Explicit About Devices
Define a single device variable, move the model once, and create or move data per batch. Avoid .cuda() sprinkled through code, which breaks on CPU and Apple Silicon (MPS) machines.
Assert Shapes at Boundaries
Shape bugs often don't raise errors; broadcasting silently produces wrong results. Add assertions or use named dimensions in comments (# (B, T, C)) around model inputs, losses, and reshapes.
Use item() and detach() for Logging
Logging loss itself keeps the whole graph alive and leaks memory across iterations. Log loss.item() (a Python float) or loss.detach().
Prefer Out-of-Place Operations Unless Memory-Bound
In-place operations can cause autograd errors ("a leaf Variable that requires grad is being used in an in-place operation") or subtle bugs. Use them deliberately, for memory savings.
Common Mistakes
Forgetting to Zero Gradients
Call optimizer.zero_grad() (with set_to_none=True, the default in recent versions) each step, unless you're intentionally accumulating gradients.
Mixing Devices
RuntimeError: Expected all tensors to be on the same device usually means the model is on GPU but a batch, mask, or freshly created tensor is on CPU. Create auxiliary tensors with device=x.device.
Converting to NumPy Inside the Graph
Calling .numpy() on a tensor that requires grad raises an error, and round-tripping through NumPy mid-computation breaks gradient flow. Stay in torch operations for anything that needs gradients.
FAQ
What's the difference between a PyTorch tensor and a NumPy array?
Both are n-dimensional arrays with similar APIs. PyTorch tensors can run on GPUs and other accelerators, and they support automatic differentiation. CPU tensors and NumPy arrays can share memory via torch.from_numpy and .numpy().
What does loss.backward() actually do?
It walks the computation graph recorded during the forward pass in reverse, computing gradients of the loss with respect to every tensor with requires_grad=True, and adds them to those tensors' .grad attributes. The optimizer then uses these gradients to update parameters.
When should I use torch.no_grad() vs inference_mode()?
Both disable gradient tracking. inference_mode is stricter and slightly faster, so it's ideal for serving and evaluation where tensors never re-enter autograd. no_grad is more flexible, for example for manual parameter updates or when outputs might later be used with autograd.
Why do gradients accumulate instead of being overwritten?
Accumulation enables techniques like gradient accumulation across micro-batches (simulating bigger batches) and summing gradients from multiple losses. The cost is that you must explicitly reset them each optimization step.
Related Topics
- PyTorch — The framework overview
- PyTorch Training Loop — Putting tensors and autograd to work
- PyTorch Datasets & DataLoaders — Feeding batches of tensors
- Transformer Architecture — A model built from these primitives
- TensorFlow — The other major deep learning framework
- Data Science — The broader field