4.1. ImageNet Classification
Training and validation error on ImageNet.
Figure 4. 左:普通网络;右:残差网络。细线为训练误差,粗线为验证误差。
3.1. Residual Learning
Let H(x) denote the desired mapping. Learning F(x) = H(x) − x rewrites the output as F(x) + x. When the shortcut and residual branch have different dimensions, a projection Ws aligns them before addition.
段落与符号一起理解,保留推导中的变量关系。
3.2. Identity Mapping
by Shortcuts
Selected equations from the paper.
框选公式,识别后可复制 LaTeX 或导出 Word。
4.1. ImageNet Classification
Training and validation error on ImageNet.
The 34-layer plain net has higher validation error than the shallower 18-layer plain net.
Figure 4. 左:普通网络;右:残差网络。细线为训练误差,粗线为验证误差。
浅蓝色曲线代表什么?为什么左图加深网络后误差更高,右图却更低?
浅蓝色代表 18 层网络。 左图是普通网络,右图是 ResNet;同色的细线、粗线分别表示训练误差和验证误差。
关键在于网络结构。 左图中,34 层普通网络连训练误差也更高;右图加入残差连接后,34 层网络表现更好。这是论文用来说明深层网络优化问题的一组对比。
带着公式,一起读懂
设 H(x) 为期望学习的映射。令 F(x) = H(x) − x,则输出可以写为 F(x) + x。当捷径分支与残差分支的维度不同时,先用投影 Wₛ 对齐维度,再相加。
- residual mapping
- 残差映射
- projection shortcut
- 投影捷径
不能只用“过拟合”解释左图:34 层普通网络连训练误差也更高。汇报时把左右图放在一起,说明残差连接改善的是深层网络的优化。





