← 中英对照目录 · ← 书架
第四章 · Section 4
Structures Underlying Human-Like General Intelligence · 类人通用智能的结构
AGI is a very broad pursuit, not tied to the creation of systems emulating human-type general intelligence. However, if one temporarily restricts attention to AGI systems intended to vaguely emulate human functionality, then one can make significantly more intellectual progress in certain interesting directions. For example, by piecing together insights from the various architectures mentioned above, one can arrive at a rough idea regarding what are the main aspects that need to be addressed in creating a "human-level AGI" system. I will present here my rough understanding of the key aspects of human-level AGI in a series of seven figures, each adapted from a figure used to describe (all or part of) one of the AGI approaches listed above. The collective of these seven figures I will call the "integrative diagram." When the term "architecture" is used in the context of these figures, it refers to an abstract cognitive architecture that may be realized in hardware, software, wetware or perhaps some other way. This "integrative diagram" is not intended as a grand theoretical conclusion, but rather as a didactic overview of the key elements involved in human-level general intelligence, expressed in a way that is not extremely closely tied to any one AGI architecture or theory, but represents a fair approximation of the AGI field's overall understanding.
AGI 是一个非常广泛的追求,并不绑定于创造"模拟人类型通用智能"的系统。但如果暂时把注意力限定在"大致模拟人类功能"的 AGI 系统上,那么在某些有趣方向上就能取得多得多的智识进展。例如,把上述各种架构的洞见拼在一起,就能大致了解创造"人类级 AGI"系统需要解决的主要面向是什么。 我将用一系列七幅图来呈现我对人类级 AGI 关键面向的粗略理解——每幅图都改编自上述某个 AGI 路径描述其(全部或部分)所用的一幅图。这七幅图的集合我称之为"整合图"(integrative diagram)。在这些图的语境中使用"架构"一词时,指的是一个抽象的认知架构——它可以在硬件、软件、湿件或其他方式中实现。这一"整合图"并非意在作为宏大的理论结论,而是对人类级通用智能所涉关键要素的一份教学式概览——其表达方式不与任何单一 AGI 架构或理论极其紧密绑定,而是对 AGI 领域整体理解的相当忠实的近似。
图 1:类人心智的高层结构
Figure 1: High-Level Structure of a Human-Like Mind
First, figure 1 gives a high-level breakdown of a human-like mind into components, based on Aaron Sloman's high-level cognitive-architectural sketch. This diagram represents, roughly speaking, "modern common sense" about the architecture of a human-like mind. The separation between structures and processes, embodied in having separate boxes for Working Memory vs. Reactive Processes, and for Long Term Memory vs. Deliberative Processes, could be viewed as somewhat artificial, since in the human brain and most AGI architectures, memory and processing are closely integrated. However, the tradition in cognitive psychology is to separate out Working Memory and Long Term Memory from the cognitive processes acting thereupon, so I have adhered to that convention. The other changes from Sloman's diagram are the explicit inclusion of language, representing the hypothesis that language processing is handled in a somewhat special way in the human brain; and the inclusion of a reinforcement component parallel to the perception and action hierarchies, as inspired by intelligent control systems theory and deep learning theory. Of course Sloman's high level diagram in its original form is intended as inclusive of language and reinforcement, but I felt it made sense to give them more emphasis.
首先,图 1 基于 Aaron Sloman 的高层认知架构草图,把类人心智高层分解为若干组件。粗略地说,这幅图代表对"类人心智架构"的"现代常识"。结构与过程的分离——体现在"工作记忆 vs. 反应式过程""长时记忆 vs. 审慎式过程"各自独立方框——可被视为有些人为,因为在人类大脑与大多数 AGI 架构中,记忆与加工紧密整合。但认知心理学的传统是把工作记忆与长时记忆与作用于其上的认知过程分开,所以我遵循了该惯例。对 Sloman 图的其他改动是:显式纳入语言——代表"人类大脑以某种特殊方式处理语言加工"的假说;以及纳入一个与感知/行动层级平行的强化组件——受智能控制系统理论与深度学习理论启发。当然,Sloman 高层图的原始形式本就意在包含语言与强化,但我觉得给予它们更多强调是合理的。
图 2:工作记忆与反应式加工
Figure 2: Working Memory and Reactive Processing (modeled on LIDA)
Figure 2, modeling working memory and reactive processing, is essentially the LIDA diagram as given in prior papers by Stan Franklin, Bernard Baars and colleagues. The boxes in the upper left corner of the LIDA diagram pertain to sensory and motor processing, which LIDA does not handle in detail, and which are modeled more carefully by deep learning theory. The bottom left corner box refers to action selection, which in the integrative diagram is modeled in more detail by Psi. The top right corner box refers to Long-Term Memory, which the integrative diagram models in more detail as a synergetic multi-memory system (Figure 4).
图 2 建模工作记忆与反应式加工,本质上就是 Stan Franklin、Bernard Baars 及同事以往论文中给出的 LIDA 图。LIDA 图左上角的方框涉及感觉与运动加工——LIDA 不详细处理它,而深度学习理论更仔细地建模它。左下角方框指行动选择——整合图中由 Psi 更详细地建模它。右上角方框指长时记忆——整合图把它更详细地建模为一个协同的多记忆系统(图 4)。
图 3:动机化行动
Figure 3: Architecture of Motivated Action (modified Psi)
Figure 3, modeling motivation and action selection, is a lightly modified version of the Psi diagram from Joscha Bach's book Principles of Synthetic Intelligence. The main difference from Psi is that in the integrative diagram the Psi motivated action framework is embedded in a larger, more complex cognitive model. Psi comes with its own theory of working and long-term memory, which is related to but different from the one given in the integrative diagram – it views the multiple memory types distinguished in the integrative diagram as emergent from a common memory substrate. Psi's handling of working memory lacks the detailed, explicit workflow of LIDA, though it seems broadly conceptually consistent with LIDA. In Figure 3, the box labeled "Other parts of working memory" is labeled "Protocol and situation memory" in the original Psi diagram. The Perception, Action Execution and Action Selection boxes have fairly similar semantics to the similarly labeled boxes in the LIDA-like Figure 2, so that these diagrams may be viewed as overlapping. The LIDA model doesn't explain action selection and planning in as much detail as Psi, so the Psi-like Figure 3 could be viewed as an elaboration of the action-selection portion of the LIDA-like Figure 2. In Psi, reinforcement is considered as part of the learning process involved in action selection and planning; in Figure 3 an explicit "reinforcement box" has been added to the original Psi diagram, to emphasize this.
图 3 建模动机与行动选择,是 Joscha Bach《合成智能原理》一书中 Psi 图的轻度修改版。与 Psi 的主要区别是:整合图中 Psi 的动机化行动框架被嵌入一个更大、更复杂的认知模型。Psi 自带自己的工作记忆与长时记忆理论——它与整合图给出的理论相关但不同:它把整合图中区分的多种记忆类型视为从一个共同记忆基质涌现出来。Psi 对工作记忆的处理缺乏 LIDA 那样详细、显式的工作流,尽管它在概念上似乎与 LIDA 大体一致。 在图 3 中,标为"工作记忆其他部分"的方框,在原始 Psi 图中标为"协议与情境记忆"。感知、行动执行与行动选择方框与类 LIDA 图 2 中同名的方框语义相当相似,因此这两幅图可被视为重叠。LIDA 模型不像 Psi 那样详细解释行动选择与规划,所以类 Psi 的图 3 可被视为类 LIDA 图 2 的行动选择部分的细化。在 Psi 中,强化被视为行动选择与规划所涉学习过程的一部分;在图 3 中,原始 Psi 图被加了一个显式的"强化方框"以强调这一点。
图 4:长时记忆与审慎/元认知加工
Figure 4: Long-Term Memory and Deliberative and Metacognitive Thinking
Figure 4, modeling long-term memory and deliberative processing, is derived from my own prior work studying the "cognitive synergy" between different cognitive processes associated with different types of memory, and seeking to embody this synergy into the OpenCog system. The division into types of memory is fairly standard in the cognitive science field. Declarative, procedural, episodic and sensorimotor memory are routinely distinguished; we like to distinguish attentional memory and intentional (goal) memory as well, and view these as the interface between long-term memory and the mind's global control systems. One focus of our AGI design work has been on designing learning algorithms, corresponding to these various types of memory, that interact with each other in a synergetic way, helping each other to overcome their intrinsic combinatorial explosions. There is significant evidence that these various types of long-term memory are differently implemented in the brain, but the degree of structure and dynamical commonality underlying these different implementations remains unclear.
图 4 建模长时记忆与审慎加工,源自作者以往研究——研究"与不同类型记忆关联的不同认知过程之间的认知协同",并寻求把这种协同体现进 OpenCog 系统。把记忆划分为类型在认知科学领域相当标准。陈述性、程序性、情景性与感觉运动记忆通常被区分;我们还喜欢区分注意记忆与意向(目标)记忆,并把它们视为长时记忆与心智全局控制系统之间的接口。我们 AGI 设计工作的一大焦点,是设计对应于这些不同类型记忆的学习算法——它们以协同方式彼此交互,互相帮助克服各自固有的"组合爆炸"。有显著证据表明这些不同类型的长时记忆在大脑中以不同方式实现,但这些不同实现之下结构/动力学的共通程度仍不清楚。
Each of these long-term memory types has its analogue in working memory as well. In some cognitive models, the working memory and long-term memory versions of a memory type and corresponding cognitive processes, are basically the same thing. OpenCog is mostly like this – it implements working memory as a subset of long-term memory consisting of items with particularly high importance values. The distinctive nature of working memory is enforced via using slightly different dynamical equations to update the importance values of items with importance above a certain threshold. On the other hand, many cognitive models treat working and long term memory as more distinct than this, and there is evidence for significant functional and anatomical distinctness in the brain in some cases. So for the purpose of the integrative diagram, it seemed best to leave working and long-term memory subcomponents as parallel but distinguished.
每种长时记忆类型在工作记忆中也有对应物。在某些认知模型中,某记忆类型及其对应认知过程的工作记忆版本与长时记忆版本,基本上是同一回事。OpenCog 大致如此——它把工作记忆实现为"由重要值特别高的项目组成"的长时记忆子集,并通过用略微不同的动力学方程更新重要值超阈值项目,来强制实现工作记忆的特殊性质。另一方面,许多认知模型把工作记忆与长时记忆视为比这更不同,而且有证据表明在某些情况下大脑中有显著的功能与解剖区分。因此,为了整合图的目的,最好把工作记忆与长时记忆子组件保持为"平行但区分"。
Figure 4 may be interpreted to encompass both workaday deliberative thinking and metacognition ("thinking about thinking"), under the hypothesis that in human beings and human-like minds, metacognitive thinking is carried out using basically the same processes as plain ordinary deliberative thinking, perhaps with various tweaks optimizing them for thinking about thinking. If it turns out that humans have, say, a special kind of reasoning faculty exclusively for metacognition, then the diagram would need to be modified. Modeling of self and others is understood to occur via a combination of metacognition and deliberative thinking, as well as via implicit adaptation based on reactive processing.
图 4 可被诠释为既包含日常的审慎思维,也包含元认知("对思考的思考")——其假说基础是:在人类与类人心智中,元认知思维基本上用与普通审慎思维相同的过程执行,或许带一些为"思考思考"优化的调整。如果结果证明人类有一种专门用于元认知的特殊推理官能,那么这幅图就需要修改。对自我与他人的建模,被理解为通过元认知与审慎思维、以及基于反应式加工的内隐适应共同发生。
图 5:多模态感知
Figure 5: Architecture for Multimodal Perception
Figure 5 models perception, according to the concept of deep learning. Vision and audition are modeled as deep learning hierarchies, with bottom-up and top-down dynamics. The lower layers in each hierarchy refer to more localized patterns recognized in, and abstracted from, sensory data. Output from these hierarchies to the rest of the mind is not just through the top layers, but via some sort of sampling from various layers, with a bias toward the top layers. The different hierarchies cross-connect, and are hence to an extent dynamically coupled together. It is also recognized that there are some sensory modalities that aren't strongly hierarchical, e.g touch and smell – these may also cross-connect with each other and with the more hierarchical perceptual subnetworks. Of course the suggested architecture could include any number of sensory modalities; the diagram is restricted to four just for simplicity.
图 5 按深度学习的概念建模感知。视觉与听觉被建模为深度学习层级,具有自下而上与自上而下的动力学。每个层级中的较低层指从感觉数据中识别并抽象出的更局部化模式。这些层级向心智其余部分的输出,不只是通过顶层,而是通过某种从各层采样的方式——偏向顶层。不同层级交叉连接,因此在某种程度上动态耦合在一起。也认识到有些感觉模态不是强层级的,如触觉与嗅觉(后者最好建模为类似不对称 Hopfield 网络、易陷入频繁混沌动力学的东西)——它们也可能彼此交叉连接、并与更层级化的感知子网络连接。当然,建议的架构可包含任意数量的感觉模态;图中只保留四个仅为简化。
The self-organized patterns in the upper layers of perceptual hierarchies may become quite complex and may develop advanced cognitive capabilities like episodic memory, reasoning, language learning, etc. A pure deep learning approach to intelligence argues that all the aspects of intelligence emerge from this kind of dynamics (among perceptual, action and reinforcement hierarchies). My own view is that the heterogeneity of human brain architecture argues against this perspective, and that deep learning systems are probably better as models of perception and action than of general cognition. However, the integrative diagram is not committed to my perspective on this – a deep-learning theorist could accept the integrative diagram, but argue that all the other portions besides the perceptual, action and reinforcement hierarchies should be viewed as descriptions of phenomena that emerge in these hierarchies due to their interaction.
感知层级上层自组织的模式可能变得相当复杂,并发展出情景记忆、推理、语言学习等高级认知能力。一种纯粹的深度学习智能观主张:智能的所有面向都从这类动力学(感知、行动与强化层级之间)中涌现。我自己的看法是:人类大脑架构的异质性反驳了这一视角,深度学习系统或许更适合作为感知与行动的模型,而非通用认知的模型。然而,整合图并不承诺我的这一看法——一位深度学习理论家可以接受整合图,但主张:除感知、行动与强化层级之外的所有其他部分,都应被视为"这些层级因交互而涌现出的现象"的描述。
图 6:行动与强化
Figure 6: Architecture for Action and Reinforcement
Figure 6 shows an action subsystem and a reinforcement subsystem, parallel to the perception subsystem. Two action hierarchies, one for an arm and one for a leg, are shown for concreteness, but of course the architecture is intended to be extended more broadly. In the hierarchy corresponding to an arm, for example, the lowest level would contain control patterns corresponding to individual joints, the next level up to groupings of joints (like fingers), the next level up to larger parts of the arm (hand, elbow). The different hierarchies corresponding to different body parts cross-link, enabling coordination among body parts; and they also connect at multiple levels to perception hierarchies, enabling sensorimotor coordination. Finally there is a module for motor planning, which links tightly with all the motor hierarchies, and also overlaps with the more cognitive, inferential planning activities of the mind, in a manner that is modeled different ways by different theorists. The reinforcement hierarchy in Figure 6 provides reinforcement to actions at various levels on the hierarchy, and includes dynamics for propagating information about reinforcement up and down the hierarchy.
图 6 展示与感知子系统平行的行动子系统与强化子系统。为具体起见,展示了两个行动层级——一个对应手臂、一个对应腿——但架构显然意在更广泛扩展。例如在对应手臂的层级中:最低层包含对应单个关节的控制模式,上一层对应关节组(如手指),再上一层对应手臂的更大部位(手、肘)。对应不同身体部位的不同层级交叉连接,实现身体部位间的协调;它们也在多个层级连接到感知层级,实现感觉运动协调。最后还有一个运动规划模块——它与所有运动层级紧密连接,也与心智更具认知性、推理性的规划活动重叠(不同理论家以不同方式建模这种重叠)。 图 6 中的强化层级为层级上各水平的行动提供强化,并包含在层级上下传播强化信息的动力学。
图 7:语言处理
Figure 7: Architecture for Language Processing
Figure 7 deals with language, treating it as a special case of coupled perception and action. The traditional architecture of a computational language comprehension system is a pipeline, which is equivalent to a hierarchy with the lowest-level linguistic features (e.g. sounds, words) at the bottom, and the highest level features (semantic abstractions) at the top, and syntactic features in the middle. Feedback connections enable semantic and cognitive modulation of lower-level linguistic processing. Similarly, language generation is commonly modeled hierarchically, with the top levels being the ideas needing verbalization, and the bottom level corresponding to the actual sentence produced. In generation the primary flow is top-down, with bottom-up flow providing modulation of abstract concepts by linguistic surface forms.
图 7 处理语言,把它视为"耦合的感知与行动"的一个特例。计算语言理解系统的传统架构是一个流水线(pipeline)——它等价于一个层级:底层是最低层语言特征(如声音、单词),顶层是最高层特征(语义抽象),句法特征在中间。反馈连接使低层语言加工能被语义与认知调制。类似地,语言生成通常也被层级化建模:顶层是需要言语化的想法,底层是实际产出的句子。生成中的主要流动是自上而下的,自下而上的流动为抽象概念提供语言表层形式的调制。
This completes the posited, rough integrative architecture diagram for human-like general intelligence, split among 7 different pictures, formed by judiciously merging together architecture diagrams produced via a number of cognitive theorists with different, overlapping foci and research paradigms. One may wonder: Is anything critical left out of the diagram? A quick perusal of the table of contents of cognitive psychology textbooks suggests that if anything major is left out, it's also unknown to current cognitive psychology. However, one could certainly make an argument for explicit inclusion of certain other aspects of intelligence, that in the integrative diagram are left as implicit emergent phenomena. For instance, creativity is obviously very important to intelligence, but, there is no "creativity" box in any of these diagrams – because in our view, and the view of the cognitive theorists whose work we've directly drawn on here, creativity is best viewed as a process emergent from other processes that are explicitly included in the diagrams.
这就完成了为类人通用智能假设的、粗略的整合架构图——它被分成 7 幅图,由许多焦点与范式不同但重叠的认知理论家所产的架构图审慎合并而成。人们或许会问:图中是否遗漏了任何关键内容?快速翻阅认知心理学教科书的目录会发现:如果有重大内容被遗漏,那也是当前认知心理学也不知道的。然而,确实可以主张显式纳入智能的某些其他面向——它们在整合图中被留作内隐的涌现现象。例如,创造力显然对智能非常重要,但任何一幅图中都没有"创造力"方框——因为在我们、以及此处直接借鉴的认知理论家看来,创造力最好被视为"从图中显式纳入的其他过程中涌现出来的过程"。
A high-level "cognitive architecture diagram" like this is certainly not a design for an AGI. Rather, it is more like a pointer in the direction of a requirements specification. These are, to a rough approximation, the aspects that must be taken into account, by anyone who wants to create a human-level AGI; and this is how these aspects appear to interact in the human mind. Different AGI approaches may account for these aspects and their interactions in different ways – e.g. via explicitly encoding them, or creating a system from which they can emerge, etc.
这样一张高层的"认知架构图"当然不是 AGI 的设计方案。相反,它更像一个指向"需求规格"方向的指针。粗略地讲,这些就是任何想创造人类级 AGI 的人都必须考虑进去的面向;而这正是这些面向在人类心智中交互的方式。不同 AGI 路径可以用不同方式处理这些面向及其交互——例如通过显式编码它们,或创建一个能让它们涌现出来的系统,等等。
← 主页