epoch 0 | train
A B

what to notice

Watch the per-layer gradient bars: units that land in the flat half of the relu stop passing gradient entirely, and at this learning rate whole layers can go quiet. The loss flattens early not because the problem is solved but because parts of the network stopped being trainable. Switch to tanh and compare.

instruments

loss

log scale
– train -- test

gradient magnitude by layer

Gradients shrink as they travel back – the products-of-slopes fact from Maths of Learning ch3, live.

permalink

learning rate0.6