← 中英对照目录 · ← 书架
§3 基于高层级认知的归纳偏置(一)
3.1 Conscious vs Unconscious Processing · 3.2 Attention as Dynamic Information Flow · 3.3 Blend of Serial and Parallel Computations
§3 导言:AI 研究与认知神经科学的协同
Synergy between AI research and cognitive neuroscience
Our aim is to take inspiration from (and further develop) research into the cognitive science of conscious processing, to deliver greatly enhanced AI, with abilities observed in humans thanks to high-level reasoning leading among other things to greater abilities to face unusual or novel situations by reasoning, compositionally reusing existing knowledge and being able to communicate about that.
我们的目标,是从意识加工的认知科学研究中汲取灵感(并进一步发展它),以交付大大增强的 AI——它拥有人类因高层级推理而展现的能力,尤其是通过推理、以组合方式复用既有知识、并能对此进行交流,来应对非常规或全新情境的更强能力。
At the same time, new AI models could drive new insights into the neural mechanisms underlying conscious processing, instantiating a virtuous circle. Machine learning procedures have the advantage that they can be tested for their effective learning abilities, and in our case in terms of out-of-distribution abilities or in the context of causal environments changing due to interventions.
与此同时,新的 AI 模型也能为意识加工背后的神经机制带来新洞见,形成良性循环。机器学习方法具有一个优点:它们可以就其有效学习能力被测试——在本文语境中,即分布外能力,或在因干预而变化的因果环境中的表现。
3.1 大脑中的意识加工与无意识加工
3.1 Conscious vs Unconscious Processing in Brains
Imagine that you are driving a car from your office, back home. You do not need to pay a lot of attention to the road and you can talk to the passenger. Now imagine encountering a road block due to construction: you have to pay more attention, you have to be on lookout for new information, if the passengers starts talking to you, then you may have to tell the person, "please let me drive".
想象你正开车从办公室回家。你不需要对路况投入太多注意,可以和乘客聊天。现在想象遇到施工路障:你必须更集中注意力、留意新信息;如果乘客开始和你说话,你可能得对他说「让我专心开车」。
It is interesting to consider that when humans are confronted with a new situation, very commonly they require their conscious attention. In the driving example, when there is a road block you need to pay attention in order to think through what to do next, and you probably don't want to be disturbed, because your conscious attention can only focus on one thing at a time.
值得注意的是:当人类面对新情境时,通常需要意识性注意。在开车例子中,遇到路障时需要集中注意思考下一步怎么办,你很可能不希望被打扰——因为意识性注意一次只能聚焦一件事。
There is something in the way humans process information which seems to be different – both functionally and in terms of the neural signature in the brain – when we deal with conscious processing and novel situations (changes in the distribution) which require our conscious attention, compared to our habitual routines.
人类加工信息的方式中,有些东西在处理「需要意识注意的、新颖情境(即分布变化)」与处理「习惯性常规」时似乎是不同的——无论从功能上,还是从大脑中的神经特征上。
In those novel situations, we generally have to think, focus and attend to specific elements of our perception, actions or memories and sometimes inhibit our reactions based on context (e.g., facing new traffic rules or a road block). Why would humans have evolved to deal with such an ability with changes in distribution? Maybe simply because life experience is highly non-stationary.
在这些新颖情境中,我们通常必须思考、集中注意于感知、行动或记忆的特定要素,有时还要根据上下文抑制我们的反应(例如面对新的交规或路障)。为什么人类会进化出这种应对分布变化的能力?也许仅仅因为人生经验是高度非平稳的
System 1 and System 2. Cognitive scientists distinguish habitual versus controlled processing, where the former correspond to default behaviors, whereas the latter require attention and mental effort. Daniel Kahneman introduced the framework of fast and slow thinking, and describes the system 1 and system 2 styles of processing in our brain.
系统 1 与系统 2。认知科学家区分「习惯性」与「受控」加工:前者对应默认行为,后者需要注意与心智努力。丹尼尔·卡尼曼提出了快思考与慢思考的框架,描述大脑中的系统 1 与系统 2 两种加工方式。
Some tasks can be achieved using only system 1 abilities whereas others also require system 2 and conscious processing. There are also notions of explicit (verbalizable) knowledge and explicit processing (which roughly correspond to system 2) and implicit (intuitive) knowledge and corresponding system 1 neural computations.
有些任务仅凭系统 1 能力就能完成,而另一些还需要系统 2 与意识加工。还存在「显式(可言语化)知识」与「显式加工」(大致对应系统 2)、以及「隐式(直觉)知识」与对应的系统 1 神经计算等概念。
The default (or unconscious) processing of system 1 can take place very rapidly (as fast as about 100ms) and mobilize many areas of the brain in parallel. On the other hand, controlled (or conscious) processing involves a sequence of thoughts, usually verbalizable, typically requiring seconds to achieve.
系统 1 的默认(或无意识)加工可以非常快地发生(最快约 100 毫秒),并并行调动大脑的许多区域。而受控(或有意识)加工涉及一系列思想,通常是可言语化的,典型需要数秒才能完成。
Whereas we can act in fast and precise habitual ways without having to think consciously, the reverse is not true: controlled processing (i.e., system 2 cognition) generally requires unconscious processing to perform much of its work. It is as if the conscious part of the computation was just the top-level program and the tip of the iceberg.
虽然我们可以不假思索地以快速、精确的习惯方式行动,但反过来却不行:受控加工(即系统 2 认知)通常需要无意识加工来承担其大部分工作。仿佛计算中有意识的部分只是最顶层的程序和冰山一角。
Yet, it seems to be a very powerful one, which makes it possible for us to solve new problems creatively by recombining old pieces of knowledge, to reason, to imagine explanations and future outcomes, to plan and to apply or discover causal dependencies. It is also at that level that we interface with other humans through natural language.
然而,它似乎非常强大:它使我们能够通过重组旧知识片段创造性地解决新问题、进行推理、想象解释与未来结果、做规划、应用或发现因果依赖。也正是这个层面,我们通过自然语言与其他人交互。
And when a word refers to a complex concept for which we do not have a clear verbalizable and precise explanation (like how we manage to drive our bike), we can still name it and reason about how it relates with other pieces of knowledge, etc.
当一个词指向一个我们并没有清晰可言语化、精确解释的复杂概念(比如我们如何骑自行车)时,我们仍能给它命名,并推理它与其他知识片段的关系。
Even imagination and planning (which are hallmarks of system 2 abilities) require system 1 computations to sample candidate solutions to a problem (from a possibly astronomical number, which we never have to explicitly examine).
即便是想象与规划(系统 2 能力的标志)也需要系统 1 计算来从可能天文数字级的候选中采样问题的候选解——我们永远不必逐一显式检验它们。
Our brain seems to thus harbour two very different types of knowledge: the kind we can explicitly reason about and communicate verbally (system 2 knowledge) and the kind that is intuitive and implicit (system 1 knowledge). When we learn something new, it typically starts being represented explicitly, and then as we practice it more, it may migrate to a different, implicit form.
我们的大脑似乎容纳着两种截然不同的知识:可以显式推理并口头交流的那类(系统 2 知识),以及直觉、隐式的那类(系统 1 知识)。当我们学习新东西时,它通常先是显式表征的,然后随着更多练习,可能迁移为另一种隐式形式。
When you learn the grammar of a new language, you may be given some set of rules, which you try to apply on the fly, but that requires a lot of effort and is done painfully slowly. As you practice this skill, it can gradually migrate to a habitual form, you make less mistakes (for the common cases), you can read / translate / write more fluently, and you may even eventually forget the original rules.
当你学习一门新语言的语法时,可能被给了一组规则,你尝试即时应用它们,但这需要大量努力、进行得异常缓慢。随着你练习这项技能,它会逐渐迁移为习惯形式:你犯的错更少(常见情形下)、读/译/写更流利,甚至最终可能忘掉最初的规则。
When a new rule is introduced, you may have to move back some of that processing to system 2 computation to avoid inconsistencies. It looks as if one of the key roles of conscious processing is to integrate different sources of knowledge (from perception and memory) in a coherent way.
当引入新规则时,你可能不得不把部分加工移回系统 2 计算以避免不一致。看起来,意识加工的一个关键作用,正是以连贯的方式整合来自知觉与记忆的不同知识源
全局工作空间理论(GWT)
The Global Workspace Theory
The above division of labour is at the heart of the cognitive neuroscience Global Workspace Theory (or GWT) from Baars and its extension, the Global Neuronal Workspace model. The GWT suggests an architecture allowing specialist components to interact.
上述分工正是 Baars 的认知神经科学全局工作空间理论(GWT)及其扩展「全局神经元工作空间模型」的核心。GWT 提出了一种允许各专业组件相互交互的架构。
The key claim of the GWT is the existence of a shared representation—sometimes called a blackboard, sometimes a workspace—that can be modified by any selected specialist and whose content is broadcast to all specialists. That selection is based on a form of attention and can correspond to dynamically selecting (based on the input) a module or a few modules in a modular neural net that are most appropriate for a particular context and task.
GWT 的关键主张是存在一种共享表征——有时称为黑板、有时称为工作空间——任何被选中的专业组件都能修改它,其内容广播给所有专业组件。这种选择基于某种注意力形式,可对应于在模块化神经网络中基于输入动态选出最适合特定上下文与任务的一个或少数几个模块。
The basic idea of deep learning frameworks inspired by the GWT is to explore a similar communication and coordination scheme for a neural net comprising of distinct modules. The GWT theory posits that conscious processing revolves around a communication bottleneck between selected parts of the brain which are called upon when addressing a current task.
受 GWT 启发的深度学习框架的基本思想,是为由不同模块构成的神经网络探索类似的通信与协调方案。GWT 理论假设:意识加工围绕大脑被选中部分之间的通信瓶颈展开,这些部分在处理当前任务时被调用。
There is a threshold of relevance beyond which information which was previously handled unconsciously gains access to this bottleneck, instantiated in a working memory. When that happens, that information is broadcast to the whole brain, allowing the different relevant parts of it to synchronize, forcing each module to learn to exchange with other modules in a way that allows swapping one module for another as source or destination of communicated content, i.e., with a shared "language".
存在一个相关性阈值,超过它,先前被无意识处理的信息就能获得进入这个瓶颈(实例化为工作记忆)的权限。一旦发生,该信息就被广播到整个大脑,让相关部分同步——迫使每个模块学会与其他模块交换信息,其方式允许把一个模块替换为另一个模块作为通信内容的源或目的地,即使用一种共享的「语言」。
These shared representations can be interpreted by many other modules. This gives rise to semantic representations that are not tied to a particular modality but can be triggered by any of the sensory channels. As we argue throughout this paper, this makes it possible to flexibly obtain new combinations of pieces of knowledge, enabling a compositional advantage aligned with the needs of systematic generalization out-of-distribution.
这些共享表征可以被许多其他模块解读,从而产生不绑定于特定模态、可由任何感觉通道触发的语义表征。正如我们在全文中论证的,这使我们能灵活获得知识片段的新组合,从而实现与分布外系统性泛化需求一致的组合优势。
3.2 注意力作为动态信息流
3.2 Attention as dynamic information flow
The GWT suggests a fleeting memory capacity in which only one consistent content can be dominant at any given moment, which suggests a sharper form of attention than the soft attention currently dominant in deep learning and described below. Attention is about sequentially selecting what computation to perform on what quantities.
GWT 提示一种瞬时记忆容量:任一时刻只能有一个一致内容占主导——这暗示一种比当前深度学习主流(下文描述的软注意力)更「锐利」的注意力形式。注意力关乎顺序选择在什么量上执行什么计算
Let us consider a machine translation task from English to French. To obtain a good translation generating the next French word, we normally focus especially on the "right" few words in the source English sentence that may be relevant to do the translation. This is the motivation that stimulated our work on content-based soft self-attention but may also be at the heart of conscious processing in humans as well as in future deep learning systems with both system 1 and system 2 abilities.
考虑一个英译法的机器翻译任务。要生成下一个法语单词并得到好翻译,我们通常特别关注源英语句子中与翻译相关的那几个「对的」词。这正是推动我们做基于内容的软自注意力的动机——它也可能处于人类意识加工、以及未来兼具系统 1 与系统 2 能力的深度学习系统的核心。
Content-Based Soft Attention. Soft attention forms a soft selection of one element (or multiple elements) from a set of elements at the previous level of computations, i.e we are taking a convex combination of the values of the elements at the previous level. These convex weights are coming from a softmax that is conditioned on how each of the elements' key vector matches some query vector.
基于内容的软注意力。软注意力对上一层计算元素集合中的一个(或多个)元素做软选择——即取上一层元素值的凸组合。这些凸权重来自 softmax,它基于每个元素的键向量与某个查询向量的匹配程度。
In a way, attention is parallel, because computing these attention weights considers all the possible elements in some set, yielding a score for each of them, to decide which of them are going to receive the most attention. With stochastic hard-attention one samples from a distribution over elements to choose the attended content, whereas with soft attention one mixes these contents with different positive convex weights.
从某种意义上说,注意力是并行的——因为计算这些注意权重会考虑集合中的所有元素、为每个元素打分,以决定哪些获得最多关注。随机硬注意力从元素分布中采样来选择被注意的内容,而软注意力以不同的正凸权重混合这些内容。
Content-based attention also introduces a non-local inductive bias into neural network processing, allowing it to infer long-range dependencies that might be difficult to discern if computations are biased by local proximity. Attention is at the heart of the current state-of-the-art NLP systems and is the standard tool for memory-augmented neural networks.
基于内容的注意力还向神经网络加工引入了非局部归纳偏置,使其能推断长程依赖——如果计算被局部邻近性偏置,这些依赖可能难以察觉。注意力处于当前最先进 NLP 系统的核心,也是记忆增强神经网络的标准工具。
Attention and memory can also help address the problem of credit assignment through long-term dependencies by creating dynamic skip connections through time (i.e., a memory access) which unlock the problems of vanishing gradients and learning long-term dependencies.
注意力与记忆还能通过创建跨时间的动态跳跃连接(即一次记忆访问)来帮助解决长程依赖下的信用分配问题,从而解开梯度消失与长程依赖学习的难题。
Attention also transforms neural networks from machines that are processing vectors (e.g., each layer of a deep net), to machines that are processing sets, more particularly sets of key/value pairs, as with Transformers.
注意力还把神经网络从「处理向量」的机器(如深度网络的每一层)转变为「处理集合」的机器——更具体地说,是处理键/值对的集合,正如 Transformer 所做的那样。
Soft attention uses the product of a query (or read key) represented as a matrix Q of dimensionality Nᵣ×d, with d the dimension of each key, with a set of Nₒ objects each associated with a key (or write-key) as a row in matrix Kᵀ (Nₒ×d), and after normalization with a softmax yields outputs in the convex hull of the values (or write-values) Vᵢ (row i of matrix V). The result is Attention(Q,K,V) = softmax(QKᵀ/√d)V.
软注意力使用查询(或读键)——表示为维度 Nᵣ×d 的矩阵 Q,d 是每个键的维度——与一组 Nₒ 个对象(每个对象关联一个键/写键,作为矩阵 Kᵀ(Nₒ×d)的一行)相乘,经 softmax 归一化后,输出落在值(或写值)Vᵢ(矩阵 V 的第 i 行)的凸包内。结果即 Attention(Q,K,V) = softmax(QKᵀ/√d)V。
With soft attention, one obtains a convex combination of the values in the rows of V, whereas stochastic hard attention would sample one of the value vectors with probability equal to that weight. If the soft attention is focused on one element for a particular row (i.e., the softmax is saturated), we get deterministic hard attention: only one of the objects is selected and its value copied to row j of the result.
软注意力得到 V 各行值的凸组合;随机硬注意力则按该权重概率采样一个值向量。如果某一行软注意力集中于一个元素(即 softmax 饱和),就得到确定性硬注意力:只选择一个对象,其值复制到结果第 j 行。
Note that the d dimensions in the key can be split into heads which then have their attention matrix and write values computed separately. Note that hard attention is more biologically plausible (we only see one interpretation of the Necker cube at once, and have one thought at a time) but soft attention enables end-to-end training and has been the most commonly used in deep learning architectures up to now, e.g., with transformers.
注意:键的 d 维可以拆分为多个「头」,各自分别计算注意力矩阵与写入值。还要注意:硬注意力在生物学上更合理(我们一次只能看到内克尔立方体的一种解读、一次只有一个念头),但软注意力支持端到端训练,且是迄今深度学习架构中最常用的(如 Transformer)。
However, there is recent evidence that if the communication bottleneck is discretized, better OOD generalization is observed, may be because the resulting simpler lingua franca would make it easier to swap one module for another in the attention-controlled communication between modules.
然而,最近有证据表明:如果通信瓶颈被离散化,会观察到更好的 OOD 泛化——也许是因为由此产生的更简单的通用语言,使模块间注意力控制的通信中「用一个模块替换另一个模块」更容易。
Attention as dynamic connections. We can think of attention as a way to create a dynamic connection between different blocks of computation, whereas in the traditional neural net setting, connections are fixed. On the receiving end (downstream module) of an attention-selected input, it is difficult to tell from the selected value vector from where it comes (among the selected upstream modules which competed for attention).
注意力即动态连接。我们可以把注意力理解为在不同计算块之间创建动态连接的方式,而传统神经网络中连接是固定的。在注意力所选输入的接收端(下游模块),仅从所选值向量很难判断它来自哪里(在竞争注意力的上游模块中)。
To resolve this, it would make sense that the information being propagated along with the selected value includes a notion of key or type or name, i.e., of where the information comes from, hence creating a form of indirection (a reference to where the information came from, which can be passed to downstream computations).
要解决这一点,随所选值传播的信息应包含「键/类型/名称」的标记——即信息来自哪里——从而创建一种间接引用(indirection)(对信息来源的引用,可传递给下游计算)。
Attention implements variable binding. When the inputs and outputs of each of the modules are a set of objects or entities (each associated with a key and value vector), we have a generic object-processing machine which can operate on "variables" in a sense analogous to variables in a programming language: as interchangeable arguments of functions.
注意力实现变量绑定。当每个模块的输入输出都是对象/实体集合(每个关联一个键向量与值向量)时,我们就有了一个通用对象处理机器,它能以类似于编程语言中变量的方式操作「变量」——即作为可互换的函数参数。
Because each object has a key embedding (which one can understand both as a name and as a type), the same computation can be applied to any variable which fits an expected "distributed type" (specified by a query vector). Each attention head then corresponds to a typed argument of the function computed by the factor.
因为每个对象都有键嵌入(既可以理解为名称,也可以理解为类型),同一个计算可应用于任何符合预期「分布式类型」(由查询向量指定)的变量。于是每个注意力头就对应因子所计算函数的一个带类型参数
When the key of an object matches the query of head k, it can be used as the k-th input vector argument for the desired computation. Whereas in regular neural networks (without attention) neurons operate on fixed input variables (the neurons which are feeding them from the previous layer), the key-value attention mechanisms make it possible to select on the fly which variable instance (i.e. which entity or object) is going to be used as input for each of the arguments of some computation, with a different set of query embeddings for each argument head.
当对象的键与第 k 个头的查询匹配时,它就能用作所需计算的第 k 个输入向量参数。在常规神经网络(无注意力)中,神经元在固定输入变量(上一层喂给它们的神经元)上运算;而键值注意力机制使我们能即时选择哪个变量实例(即哪个实体或对象)被用作某个计算每个参数的输入,每个参数头使用不同的查询嵌入集。
The computations performed on the selected inputs can be seen as functions with typed arguments, and attention is used to bind their formal argument to the selected input, albeit in a soft differentiable way (that mixes multiple possibilities) in the case of soft attention. Type constraints have already been found useful in identification for causal discovery.
对所选输入执行的计算可看作「带类型参数」的函数,注意力被用来把形参绑定到所选输入——尽管软注意力是以软的、可微的方式(混合多种可能性)。类型约束已被发现在因果发现的可识别性中有用。
Current attention-based neural network already implement key-value-query soft attention mechanism (as above). What is missing is an ability to handle discrete types, hard (but possibly stochastic) choices of arguments, and more powerful inference machinery that uses not just type matching but is also able to reason about which modules and variables should be composed in a given context.
当前基于注意力的神经网络已经实现了键-值-查询软注意力机制(如上)。所缺的是:处理离散类型的能力、对参数做硬(但可随机)选择的能力,以及更强力的推理机制——它不仅用类型匹配,还能推理在给定上下文中应组合哪些模块与变量。
3.3 串行与并行计算的混合
3.3 Blend of Serial and Parallel Computations
From a computational perspective, one hypothesis about the dynamics of communication between different modules is that different modules generally act in parallel and receive inputs from other modules. However, when they do need to communicate information with another arbitrary module, the information has to go through a routing bottleneck (the global workspace) controlled by an attention mechanism.
从计算的角度看,关于不同模块间通信动力学的一个假设是:不同模块通常并行运转、接收来自其他模块的输入。然而当它们确实需要与另一任意模块通信信息时,信息必须经过由注意力机制控制的路由瓶颈(全局工作空间)。
Because so few elements can be put in coherence at each step of the GWT selection, the inference process generally requires several such steps, leading to the highly sequential nature of system 2 computation (compared with the highly parallel nature of system 1 computation).
由于 GWT 选择的每一步只能让极少数元素进入一致状态,推理过程通常需要多个这样的步骤,导致系统 2 计算的高度序列化性质(与系统 1 计算的高度并行性质相对)。
The contents which have thus been selected are essentially the only ones which can be committed to memory, starting with short-term memory. Working memory refers to the ability of the brain to operate on a few recently accessed elements (i.e., those in short-term memory).
这样被选中的内容,本质上是可以被提交到记忆(首先是短期记忆)的唯一内容。工作记忆指大脑对少数最近访问的元素(即短期记忆中的那些)进行运算的能力。
These elements can be remembered and have a heavy influence on the next thought, action or perception, as well as on what learning focuses on, possibly playing a role similar to desired outputs, goals or targets in supervised learning for system 1 computations.
这些元素能被记住,并对下一个想法、行动或知觉、以及对学习聚焦于什么产生重大影响——可能扮演类似监督学习中「期望输出、目标或标签」之于系统 1 计算的角色。
Partial State. From an RL perspective, it is interesting to note that if the GWT holds an important part of the state (including imagined future states, when planning), it does not describe all the aspects of the environment, only a handful of them, as already explored in the RL literature. This is different from standard RL approaches where the input (or the sequence of past inputs) is mapped to a fixed-size (estimated and latent) state vector.
部分状态(Partial State)。从强化学习(RL)角度看,值得注意的是:如果 GWT 持有状态的重要部分(规划时包括想象的未来状态),它并不描述环境的全部方面,只描述其中一小撮——正如 RL 文献已探索的那样。这与标准 RL 方法不同:标准方法把输入(或过去输入的序列)映射到固定大小的(估计的、潜在的)状态向量。
The GWT suggests instead that, besides long-term memory content (which mostly does not change), the rapidly changing state should be seen as a very small set of entities (e.g., objects or particular attributes of objects, and their relation), with an information content similar to that of a single sentence.
GWT 则提示:除了长期记忆内容(大多不变),快速变化的状态应被看作一个非常小的实体集合(如对象、对象的特定属性及其关系),其信息内容类似于一个句子。
This suggests neural net architectures in which very few modules and specific (variable, value) pairs are selected at every inference step, based on those that were recently selected, the current sensory input and the current contents of memory (which can also compete for write-access to the workspace).
这提示一种神经网络架构:每一步推理只选择极少数模块与特定的(变量、值)对,选择依据是最近被选中的那些、当前感觉输入、以及记忆的当前内容(记忆内容也能竞争工作空间的写权限)。
Only the selected modules would be under pressure to adapt when the result of the combination needs to be tuned, leading to selective adaptation similar to that explored by (Bengio et al., 2019) (see Section 4.3 above) where just a few relevant modules need to adapt to a change in distribution.
当组合结果需要调优时,只有被选中的模块承受适应压力,导致选择性适应——类似于 Bengio 等人 2019 年探索的机制(见上文 4.3 节):只需少数相关模块适应分布变化。
System 2 to System 1 Consolidation. As an agent, a human being is facing frequent changes because of their actions or the actions of other agents in the environment. Most of the time, humans follow their habitual policy, but tend to use system 2 cognition when having to deal with unfamiliar settings. It allows humans to generalize out-of-distribution in surprisingly powerful ways, and understanding this style of processing would help us build these abilities in AI as well.
系统 2 到系统 1 的固化(Consolidation)。作为智能体,人类因自身行动或环境中其他智能体的行动而频繁面对变化。大多数时候,人类遵循习惯性策略,但在必须处理陌生情境时倾向于使用系统 2 认知。这使人类能以惊人强大的方式实现分布外泛化——理解这种加工方式将帮助我们在 AI 中构建这些能力。
This is illustrated with our early example of driving in an area with unfamiliar traffic regulations, which requires full conscious attention (Section 3.1). This observation suggests that system 2 cognition is crucial in order to achieve the kind of flexibility and robustness to changes in distribution required in the natural world.
这由我们前面「在交规陌生的地区开车需要全神贯注」的例子(3.1 节)加以说明。这一观察表明:要实现自然世界所需的、对分布变化的灵活性与鲁棒性,系统 2 认知至关重要。
It looks like current deep learning systems are fairly good at perception and system 1 tasks. They can rapidly produce an answer (if you have parallel computing like that of GPUs) through a complex calculation which is difficult (or impossible) to dissect into the application of a few simple verbalizable operations. They require a lot of practice to learn and can become razor sharp good at the kinds of data they are trained on.
看起来当前深度学习系统在感知与系统 1 任务上相当出色:它们能通过复杂计算(若有 GPU 那样的并行计算)快速产生答案——而这种计算难以(或不可能)拆解为若干简单可言语化操作的应用。它们需要大量练习来学习,并在所训练的数据类型上变得极其锋利。
On the other hand, humans enjoy system 2 abilities which permit fast learning (I can tell you a new rule in one sentence and you do not have to practice it in order to be able to apply it, albeit awkwardly and slowly at first) and systematic generalization, both of which should be important characteristics of the next generation of deep learning systems.
另一方面,人类享有系统 2 能力:它们允许快速学习(我用一句话告诉你一条新规则,你无需练习就能应用它——虽然起初笨拙而缓慢)与系统性泛化——这两者都应是下一代深度学习系统的重要特征。
Between-Modules Interlingua and Communication Topology. If the brain is composed of different modules, it is interesting to think about what code or lingua franca is used to communicate between them, such that it can lead to interchangeable pieces of knowledge being dynamically selected and combined to solve a new problem.
模块间通用语言与通信拓扑。如果大脑由不同模块组成,就有趣了:它们之间用什么代码或通用语言通信,才能让可互换的知识片段被动态选择和组合来解决新问题?
The GWT bottleneck may thus also play a role in forcing the emergence of such a lingua franca: the same information received by module A (e.g. "there is a fire") can come from any other module (say B, which detected a fire by smell, or C which detected a fire by sight). Hence B and C need to use a compatible representation which is broadcast via the GWT bottleneck for A's use.
因此 GWT 瓶颈也扮演着「迫使这种通用语言涌现」的角色:模块 A 收到的同一信息(如「有火情」)可以来自任何其他模块(如 B——它通过嗅觉发现火情,或 C——它通过视觉发现火情)。于是 B 和 C 需要使用一种兼容表征,通过 GWT 瓶颈广播供 A 使用。
Again, we see the crucial importance of attention mechanisms to force the emergence of shared representations and indirect references exchanged between the modules via the conscious bottleneck. However, the GWT bottleneck is by far not the only way for modules to communicate with each other.
我们再次看到注意力机制的关键重要性:它迫使共享表征与间接引用通过意识瓶颈在模块间交换。然而,GWT 瓶颈远非模块间通信的唯一方式。
Regarding the topology of the communication channels between modules, it is known that modules in the brain satisfy some spatial topology such that the computation is not all-to-all between all the modules. It is plausible that the brain uses both fixed local or spatially nearby connections as well as the global broadcasting system with top-down influence.
关于模块间通信通道的拓扑,已知大脑中的模块满足某种空间拓扑——模块之间并非全连接。合理的是:大脑既使用固定的局部/空间邻近连接,也使用带自上而下影响的全局广播系统。
We also know that there are hierarchical communication routes in the visual cortex (on the path from pixels to object recognition), and we know how successful that has been in computer vision with convnets. Combining these different kinds of inter-module communication modalities in deep network thus seems well advised as well:
我们还知道视觉皮层中存在层级通信路径(从像素到目标识别的路径上),也知道这在计算机视觉的卷积网络中多么成功。因此,在深度网络中组合这些不同类型的模块间通信方式,似乎也是明智的:
(1) Modules which are near each other in the brain layout can probably communicate directly without the need to clog the global broadcast channel (and this would not be reportable consciously). (2) Modules which are arbitrarily far from each other in the spatial layout of the brain can exchange information via the global workspace, following the theatre analogy of Baars' GWT.
(1) 在大脑布局中彼此邻近的模块大概可以直接通信,无需堵塞全局广播通道(且这不会被意识报告);(2) 在大脑空间布局中任意远的模块,则遵循 Baars GWT 的剧场类比,通过全局工作空间交换信息。
The other advantage of this communication route is of course the exchangeability of the sources of information being broadcast, which we hypothesize leads to better systematic generalization.
这条通信路径的另一个优势,自然是被广播信息来源的可交换性——我们假设它带来更好的系统性泛化。
The role of working memory in the GWT is not just as a communication buffer. It also serves as a blackboard (or analogously the "registers" in CPUs) where operations can be done locally to improve coherence. This enables a coherence-seeking mechanism: the different modules (especially the active ones) should adopt a configuration of their internal variables (and especially the more abstract entities they communicate to other modules) which is consistent with what other active modules "believe".
工作记忆在 GWT 中不只是通信缓冲区。它还是黑板(类似于 CPU 中的「寄存器」),可在上面做局部运算以提升一致性。这启用了一种寻求一致性机制:不同模块(尤其是活跃的模块)应让自己的内部变量(尤其是它们与其他模块通信的更抽象实体)配置,与其他活跃模块所「相信」的内容一致。
It is possible, that a large part of the functional role of conscious processing is for that purpose, which is consistent with the view of the working memory as a central element of the inference machinery seeking to obtain coherent configurations of the variables interacting according to some piece of knowledge (such as a factor of the factor graph, a causal dependency).
很可能,意识加工的大部分功能角色正是为此——这与把工作记忆视为推理机制中心元素的看法一致:该机制寻求让变量按照某条知识(如因子图的一个因子、一条因果依赖)交互时取得一致的配置。
← 主页