§5 结论
5. Conclusions
To be able to handle dynamic, changing conditions, we want to move from deep statistical models which are able to perform system 1 tasks to deep structural models also able to perform system 2 tasks by taking advantage of the computational workhorse of system 1 abilities.
为了能够应对动态、变化的条件,我们希望从「能执行系统 1 任务的深度统计模型」,走向「也能执行系统 2 任务的深度结构模型」——同时利用系统 1 能力这个计算主力。
Today's deep networks may benefit from additional structure and inductive biases to do substantially better on system 2 tasks, natural language understanding, out-of-distribution systematic generalization and efficient transfer learning. We have tried to clarify what some of these inductive biases may be, but much work needs to be done to improve that understanding and find appropriate ways to incorporate these priors in neural architectures and training frameworks.
今日的深度网络可能受益于额外的结构与归纳偏置,以便在系统 2 任务、自然语言理解、分布外系统性泛化与高效迁移学习上做得更好。我们已经尝试阐明其中一些归纳偏置可能是什么,但要改进这种理解、并找到把这些先验纳入神经架构与训练框架的恰当方式,仍需大量工作。
We have motivated these inductive biases in terms of expected (and observed in recent work) gains in terms of out-of-distribution generalization and fast adaptation in transfer settings rather than the standard test set from the same distribution as the training set.
我们对这些归纳偏置的动机论证,是基于分布外泛化与迁移设置中快速适应的预期(以及近期工作中观察到的)收益——而非基于与训练集同分布的标准测试集。
The general insight here is that the proposed inductive biases should help organize knowledge into the stable reusable parts that are likely to be useful in new settings and tasks (such as causal mechanisms), separating them from the more volatile pieces of information (the values of variables) that can be changed by agents (through causal intervention) or those that are affected by these changes and may vary across environments or tasks.
这里的一般洞见是:所提归纳偏置应帮助把知识组织为稳定可复用部分(它们在新设置与新任务中很可能有用,如因果机制),并把它们与更易变的信息片段(变量的值,可被智能体通过因果干预改变)或受这些变化影响、可能随环境或任务而变的内容分离开来。
Some of the more salient inductive biases we propose deserve especially more attention in deep learning research include (a) the fairly direct connection between high-level variables and natural language or more generally how humans communicate knowledge among them, i.e., we can verbalize our thoughts to a large extent and this can provide rich insights about underlying inductive biases such as these: (b) the modular decomposition of knowledge into independent reusable pieces that can be composed on the fly to address new contexts, (c) the causal interpretation of actions by agents and of changes in distribution, with agents generally intending to affect a single or very few (generally latent) variables, and (d) the sparsity of dependencies between high-level variables (and thus the small number of variables that are linked by causal mechanisms imagined by humans to explain their environment).
我们提出的一些更突出的归纳偏置尤其值得深度学习研究关注,包括:(a) 高层级变量与自然语言之间相当直接的联系——更一般地说,人类如何在彼此间交流知识——即我们能在很大程度上言语化自己的思想,这能为底层归纳偏置提供丰富洞见;(b) 把知识模块化分解为独立可复用片段,可在运行中组合以应对新上下文;(c) 对智能体行动与分布变化的因果解读——智能体通常意图影响单个或极少数(通常是潜在的)变量;(d) 高层级变量之间依赖的稀疏性——从而人类用以解释其环境的因果机制所链接的变量数量很少。
Finally, we would also like to mention that inductive biases are not the only way to bridge the gap to high-level human cognition: we may gain by improving our optimization algorithms, by scaling up neural networks and by moving to other frameworks that better capture uncertainty about the world (e.g., by learning a Bayesian posterior over neural network models as compared to learning a point estimates).
最后,我们还想说明:归纳偏置并不是弥合通向人类高层级认知差距的唯一途径——我们还可能通过改进优化算法、扩展神经网络规模、以及转向更好捕获世界不确定性的其他框架(例如学习神经网络模型上的贝叶斯后验,而非学习点估计)来获益。
It would also be intriguing to think of ways to combine all these different elements together.
思考如何把所有不同要素组合在一起,也将是引人入胜的。
The authors are grateful to Alex Lamb, Rosemary Nan Ke and Olexa Bilaniuk for leading many projects some of which are also discussed here. The authors are grateful to Mike Mozer and Bernhard Schölkopf for many brainstorming discussions. The authors would also like to thank Stefan Bauer, Aniket Didolkar, Nasim Rahaman, Kanika Madan, Philippe Beaudoin, Charles Blundell, Dianbo Liu, Moksh Jain for useful feedback. The authors would also like to acknowledge comments by the reviewers which helped in improving the manuscript.
作者感谢 Alex Lamb、Rosemary Nan Ke 与 Olexa Bilaniuk 主导了本文中讨论的多个项目;感谢 Mike Mozer 与 Bernhard Schölkopf 进行了大量头脑风暴式讨论;并感谢 Stefan Bauer、Aniket Didolkar、Nasim Rahaman、Kanika Madan、Philippe Beaudoin、Charles Blundell、Dianbo Liu、Moksh Jain 提供的有用反馈;同时感谢审稿人的意见,帮助改进了稿件。