Study companion and working knowledge base for the Deep Learning textbook by Goodfellow, Bengio & Courville (MIT Press, 2016), read free at deeplearningbook.org. Indexes all 20 chapters, carries a 2016-to-2026 delta layer naming what the book got right, what was superseded (transformers, AdamW, diffusion, double descent) and what still holds, and ships four deterministic tools: a prerequisite-aware reading-path planner, a training-failure diagnostic, a capacity-and-regularization planner, and a parameter/FLOP/activation-memory calculator. Use when studying or teaching this book, planning a route through it, deciding whether a chapter's advice is still current, or translating its math into a training decision. It points at the official chapters — it never reproduces them.
66
80%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Passed
No findings from the security scan
Source book: Deep Learning, Ian Goodfellow, Yoshua Bengio & Aaron Courville (MIT Press, 2016) · 20 chapters, 3 parts · read free at deeplearningbook.org · companion compiled 2026-08-25.
This is a companion, not a copy. The book is copyrighted, and its site states that the HTML-only format exists to discourage copying under the authors' MIT Press contract. Nothing here reproduces its text. Every chapter file is original synthesis — what the chapter establishes, how to use it, where it has aged — plus a link to the official chapter. Read the book at the link; use this to navigate it, keep it current, and turn it into decisions. See references/rights_and_use.md.
regularization, saddle points, partition function; resolved
through the Topic Index, then that chapter file is read before answering.chNN — load that chapter's file.scripts/reading_path_planner.py.When asked about something outside these 20 chapters, say so and route to the delta reference rather than improvising the book's position on material published after it.
Name the task, the performance measure, and the experience in one sentence before any model code. Most failed projects failed at P: an unstated metric, or a proxy whose relationship to the real objective was never checked.
Choose the output distribution, then take its negative log. Gaussian → MSE, Bernoulli → binary cross-entropy, categorical → cross-entropy, Laplace → MAE. "Which loss?" is always the question "which distribution?" in disguise. Modern contrastive and preference objectives sit outside this frame — a real limit of the book, not a gap in your understanding.
D(p‖q) ≠ D(q‖p). Forward KL is mode-covering (blurry averages); reverse KL is mode-seeking (sharp but partial). This single fact predicts VAE blur, GAN mode collapse, and the characteristic over-confidence of mean-field variational posteriors.
High training error → capacity or optimization is the bottleneck; more data will not help.
Low training error with a large validation gap → data or regularization. This is the highest-value
heuristic in the book. scripts/training_diagnostics.py runs it.
Regularization trades variance for bias. But the classical U-shaped capacity curve is incomplete: past the interpolation threshold, test error can fall again (double descent, 2019–2020, post-dating the book). Practical consequence: when a large model overfits, try more data, more regularization or longer training before shrinking it.
Convolution asserts translation equivariance and locality. Recurrence asserts that the past compresses into a state. A distributed representation asserts that factors combine combinatorially. When the assertion is false, the architecture cannot be rescued by tuning — and when it is true, it beats capacity. This is also why Vision Transformers need more data than ConvNets: they discard the prior and buy it back with examples.
Backprop is the chain rule scheduled well: one forward-pass-equivalent of compute, and memory proportional to stored activations. Depth fails through vanishing/exploding gradients and ill-conditioning, which is why residual connections, normalization and clipping exist.
For undirected models, the likelihood gradient needs samples from the model itself. Four escape routes: sample it (CD/PCD), sidestep it algebraically (pseudolikelihood, score matching), learn around it (NCE), or estimate it for evaluation (AIS). Score matching's descendants are today's diffusion models — which is why Part III repays reading even though its models did not survive.
Gradient norm exploding → clip. Norm large but loss flat → ill-conditioning. Norm near zero with high loss → saturation or dead units. NaN → numerics first. Change one thing per experiment.
| # | Title | Key content |
|---|---|---|
| ch01 | Introduction | representation learning, depth as composition, curse of dimensionality |
| ch02 | Linear Algebra | norms, SVD, eigendecomposition, conditioning, PCA |
| ch03 | Probability & Information Theory | distributions, entropy, KL, cross-entropy |
| ch04 | Numerical Computation | under/overflow, conditioning, gradient descent, KKT |
| ch05 | Machine Learning Basics | capacity, bias–variance, No Free Lunch, MLE, manifolds |
| ch06 | Deep Feedforward Networks | output/hidden units, universal approximation, backprop |
| ch07 | Regularization | norm penalties, augmentation, early stopping, dropout |
| ch08 | Optimization | SGD, momentum, init, Adam, batch norm, saddles |
| ch09 | Convolutional Networks | sparse interactions, sharing, equivariance, pooling |
| ch10 | Sequence Modeling | BPTT, vanishing gradients, LSTM/GRU, attention |
| ch11 | Practical Methodology | metrics, baselines, the data-vs-capacity rule, debugging |
| ch12 | Applications | scaling, compression, vision, speech, NLP (dated) |
| ch13 | Linear Factor Models | PPCA, factor analysis, ICA, sparse coding |
| ch14 | Autoencoders | undercomplete, sparse, denoising, contractive |
| ch15 | Representation Learning | transfer, distributed codes, disentanglement |
| ch16 | Structured Probabilistic Models | directed/undirected, energy-based, d-separation |
| ch17 | Monte Carlo Methods | importance sampling, MCMC, Gibbs, mixing |
| ch18 | Confronting the Partition Function | CD/PCD, pseudolikelihood, score matching, NCE, AIS |
| ch19 | Approximate Inference | ELBO, EM, mean field, amortization |
| ch20 | Deep Generative Models | Boltzmann machines, VAE, GAN, autoregressive |
S=engineering/deep-learning-book/skills/deep-learning-book/scripts
python3 $S/reading_path_planner.py --goal "train a transformer" --background applied --hours-per-week 5
python3 $S/training_diagnostics.py --train-loss 0.02 --val-loss 1.9 --grad-norm 0.4 --epochs 30
python3 $S/capacity_planner.py --params 12000000 --train-examples 50000 --train-error 0.01 --val-error 0.22
python3 $S/model_arithmetic.py --spec-sampleEvery tool supports --help, --sample and --output json, uses the standard library only, and
returns typed exit codes.
This companion covers the 2016 edition's 20 chapters and the delta between them and 2026
practice. It does not cover: reinforcement learning beyond passing mention, LLM training
infrastructure, RLHF/DPO alignment, agentic systems, MLOps tooling, or fairness and safety
evaluation — none of which the book treats. For production ML engineering use
engineering-team/senior-ml-engineer; for LLM cost work use engineering/llm-cost-optimizer.
When a question lands outside the book, say the book does not cover it and cite the delta reference for what replaced its position. A companion that quietly extrapolates is worse than one that names its boundary.
19392f7
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.