第五章 · Section 5
Metrics and Environments for Human-Level AGI · 人类级 AGI 的度量与环境
Science hinges on measurement; so if AGI is a scientific pursuit, it must be possible to measure what it means to achieve it.
Given the variety of approaches to AGI, it is hardly surprising that there are also multiple approaches to quantifying and measuring the achievement of AGI. However, things get a little simpler if one restricts attention to the subproblem of creating "human-level" AGI.
When one talks about AGI beyond the human level, or AGI that is very qualitatively different from human intelligence, then the measurement issue becomes very abstract – one basically has to choose a mathematical measure of general intelligence, and adopt it as a measure of success. This is a meaningful approach, yet also worrisome, because it's difficult to tell, at this stage, what relation any of the existing mathematical measures of general intelligence is going to have to practical systems.
When one talks about human-level AGI, however, the measurement problem gets a lot more concrete: one can use tests designed to measure human performance, or tests designed relative to human behavior. The measurement issue then decomposes into two subproblems: quantifying achievement of the goal of human-level AGI, and measuring incremental progress toward that goal. The former subproblem turns out to be considerably more straightforward.
科学以测量为枢纽;所以如果 AGI 是一项科学追求,就必须能够测量"达成它"意味着什么。
鉴于 AGI 路径的多样性,存在多种量化和测量 AGI 达成的方法也就不足为奇。然而,如果把注意力限定在"创造人类级 AGI"这个子问题上,事情会简单一些。
当谈论超出人类水平的 AGI,或与人类智能有性质上巨大差异的 AGI 时,测量问题变得非常抽象——基本上必须选择一种通用智能的数学度量,并把它当作成功标准。这是一种有意义的路径,但也令人担忧——因为在现阶段,很难说现有的任何通用智能数学度量会与实际系统有什么关系。
而当谈论人类级 AGI 时,测量问题具体得多:可以用为测量人类表现而设计的测试,或相对人类行为设计的测试。测量问题于是分解为两个子问题:量化"人类级 AGI"目标的达成,以及测量朝该目标的增量进展。前一个子问题被证明要直接得多。
5.1 度量与环境
Metrics and Environments
The issue of metrics is closely tied up with the issue of "environments" for AGI systems. For AGI systems that are agents interacting with some environment, any method of measuring the general intelligence of these agents will involve the particulars of the AGI systems' environments. If an AGI is implemented to control video game characters, then its intelligence must be measured in the video game context. If an AGI is built with solely a textual user interface, then its intelligence must be measured purely via conversation, without measuring, for example, visual pattern recognition.
And the importance of environments for AGI goes beyond the value of metrics. Even if one doesn't care about quantitatively comparing two AGI systems, it may still be instructive to qualitatively observe the different ways they face similar situations in the same environment. Using multiple AGI systems in the same environment also increases the odds of code-sharing and concept-sharing between different systems. It makes it easier to conceptually compare what different systems are doing and how they're working.
度量问题与 AGI 系统的"环境"问题紧密相连。对于作为智能体与环境交互的 AGI 系统,任何测量这些智能体通用智能的方法,都会涉及 AGI 系统环境的特定细节。如果 AGI 被实现来控制视频游戏角色,那么它的智能必须在视频游戏语境中测量。如果 AGI 只用文本用户界面构建,那么它的智能必须纯粹通过对话测量——例如不去测量视觉模式识别。
环境对 AGI 的重要性还不止于度量价值。即使不关心定量比较两个 AGI 系统,定性观察它们在同一环境中面对相似情境的不同方式,也可能很有启发性。在同一环境中使用多个 AGI 系统,也增加了不同系统间共享代码与共享概念的可能性,让概念上比较不同系统在做什么、如何工作变得更容易。
It is often useful to think in terms of "scenarios" for AGI systems, where a "scenario" means an environment plus a set of tasks defined in that environment, plus a set of metrics to measure performance on those tasks. At this stage, it is unrealistic to expect all AGI researchers to agree to conduct their research relative to the same scenario. The early-stage manifestations of different AGI approaches tend to fit naturally with different sorts of environments and tasks. However, to whatever extent it is sensible for multiple AGI projects to share common environments or scenarios, this sort of cooperation should be avidly pursued.
用"场景"来思考 AGI 系统常常很有用——"场景"指一个环境、在该环境中定义的一组任务,加上一组测量这些任务表现的度量。在现阶段,期望所有 AGI 研究者同意在同一个场景下开展研究是不现实的。不同 AGI 路径的早期形态,往往自然地契合不同类型的环境与任务。然而,只要多个 AGI 项目共享公共环境或场景是明智的,就应该热切地追求这种合作。
5.2 量化"人类级 AGI"里程碑
Quantifying the Milestone of Human-Level AGI
A variety of metrics, relative to various different environments, may be used to measure achievement of the goal of "human-level AGI." Examples include:
• the classic Turing Test, conceived as (roughly) "fooling a panel of college-educated human judges, during a one hour long interrogation, that one is a human being".
• the Virtual World Turing Test occurring in an online virtual world, where the AGI and the human controls are controlling avatars (this is inclusive of the standard Turing Test if one assumes the avatars can use language).
• Shane Legg's AIQ measure, which is a computationally practical approximation to the algorithmic information theory based formalization of general intelligence given by Legg and Hutter.
• Text compression – the idea being that any algorithm capable of understanding text should be transformable into an algorithm for compressing text based on the patterns it recognizes therein. This is the basis of the Hutter Prize, a cash prize which rewards data compression improvements on a specific 100 MB English text file.
• the Online University Student Test, where an AGI has to obtain a college degree at an online university, carrying out the same communications with the professors and the other students as a human student would (including choosing its curriculum).
• the Robot University Student Test, where an AGI has to obtain a college degree at a physical university, carrying out the same communications as a human student would, and also moving about the campus and handling relevant physical objects in a sufficient manner to complete the coursework.
• the Artificial Scientist Test, where an AGI that can do high-quality, original scientific research, including choosing the research problem, reading the relevant literature, writing and publishing the paper, etc. (this may be refined to a Nobel Prize Test).
相对于各种不同环境,有多种度量可用于测量"人类级 AGI"目标的达成。例子包括:
• 经典图灵测试——大致设想为"在一小时长的审问中,骗过一群受过大学教育的人类裁判,让它们相信你是人"。
• 虚拟世界图灵测试——发生在一个在线虚拟世界中,AGI 与人类控制者都在控制化身(若假定化身能使用语言,则它包含标准图灵测试)。
• Shane Legg 的 AIQ 度量——它是 Legg-Hutter 基于算法信息论形式化的通用智能的一种可计算实用近似。
• 文本压缩——其想法是:任何能理解文本的算法都应能被转化为基于它所识别模式的文本压缩算法。这是 Hutter 奖的基础——一个由 Marcus Hutter 出资的现金奖,奖励对特定 100 MB 英文文本文件的数据压缩改进。
• 在线大学生测试——AGI 必须在一所在线大学获得学位,像人类学生一样与教授和其他学生进行相同交流(包括选择课程等)。
• 机器人大学生测试——AGI 必须在一所实体大学获得学位,像人类学生一样交流,并且要在校园中移动、以足够完成课业的方式处理相关物理对象。
• 人工科学家测试——AGI 能做高质量、原创的科学研究,包括选择研究问题、阅读相关文献、撰写并发表论文等(可细化为诺贝尔奖测试)。
Each of these approaches has its pluses and minuses. None of them can sensibly be considered necessary conditions for human-level intelligence, but any of them may plausibly be considered sufficient conditions. The latter three have the disadvantage that they may not be achievable by every human – so they may set the bar a little too high. The former two have the disadvantage of requiring AGI systems to imitate humans, rather than just honestly being themselves; and it may be that accurately imitating humans when one does not have a human body or experience, requires significantly greater than human level intelligence.
Regardless of the practical shortcomings of the above measures, though, I believe they are basically adequate as precisiations of "what it means to achieve human-level general intelligence."
这些方法各有优缺点。它们中没有哪一个能合理地被视为人类级智能的必要条件,但任何一个都可能合理地被视为充分条件。后三个(在线大学生/机器人大学生/人工科学家)的缺点是:并非每个人类都能达成它们——所以它们可能把标杆设得有点过高。前两个(图灵测试/虚拟世界图灵测试)的缺点是:它们要求 AGI 系统模仿人类,而非诚实地做自己;而且在没有人类身体或经验时精确模仿人类,可能需要显著高于人类水平的智能。
不过,撇开上述度量的实际缺陷,我相信它们作为"达成人类级通用智能意味着什么"的精确化,基本上是充分的。
5.3 测量朝人类级 AGI 的增量进展
Measuring Incremental Progress Toward Human-Level AGI
While postulating criteria for assessing achievement of full human-level general intelligence seems relatively straightforward, positing good tests for intermediate progress toward the goal of human-level AGI seems much more difficult.
That is: it is not clear how to effectively measure whether one is, say, 50 percent of the way to human-level AGI? Or, say, 75 or 25 percent?
What I have found via a long series of discussions on this topic with a variety of AGI researchers is that:
• It's possible to pose many "practical tests" of incremental progress toward human-level AGI, with the property that if a proto-AGI system passes the test using a certain sort of architecture and/or dynamics, then this implies a certain amount of progress toward human-level AGI based on particular theoretical assumptions about AGI.
• However, in each case of such a practical test, it seems intuitively likely to a significant percentage of AGI researchers that there is some way to "game" the test via designing a system specifically oriented toward passing that test, and which doesn't constitute dramatic progress toward AGI.
虽然提出评估"达成完整人类级通用智能"的标准似乎相对直接,但为"朝人类级 AGI 目标的中期进展"提出好的测试似乎要困难得多。
也就是说:如何有效测量"你是否已经走了一半通往人类级 AGI 的路",并不清楚?是 75% 还是 25%?
通过就此话题与多位 AGI 研究者的一系列长期讨论,我发现:
• 可以提出许多朝人类级 AGI 增量进展的"实用测试",其性质是:如果一个原初 AGI 系统用某种架构/动力学通过了测试,那么基于对 AGI 的特定理论假设,它意味着朝人类级 AGI 有了一定的进展。
• 然而,对于每个这样的实用测试,相当比例的 AGI 研究者直觉上认为存在某种"钻测试空子"的方式——即设计一个专门面向通过该测试、却并不构成朝 AGI 戏剧性进展的系统。
A series of practical tests of this nature were discussed and developed at a 2009 gathering at the University of Tennessee, Knoxville, called the "AGI Roadmap Workshop," which led to an article in AI Magazine titled Mapping the Landscape of Artificial General Intelligence. Among the tests discussed there were:
• The Wozniak "coffee test": go into an average American house and figure out how to make coffee, including identifying the coffee machine, figuring out what the buttons do, finding the coffee in the cabinet, etc.
• Story understanding – reading a story, or watching it on video, and then answering questions about what happened (including questions at various levels of abstraction)
• Passing the elementary school reading curriculum (which involves reading and answering questions about some picture books as well as purely textual ones)
• Learning to play an arbitrary video game based on experience only, or based on experience plus reading instructions
• Passing child psychologists' typical evaluations aimed at judging whether a human preschool student is normally intellectually capable
2009 年,在田纳西大学诺克斯维尔分校举行的一场名为"AGI 路线图研讨会"的聚会上,讨论并发展了一系列这类实用测试,最终产生了发表在 AI Magazine 上的文章《绘制通用人工智能的地景》。会上讨论的测试包括:
• Wozniak"咖啡测试":走进一座普通的美国房子,弄明白如何煮咖啡——包括识别咖啡机、弄明白按钮的作用、在橱柜里找到咖啡等。
• 故事理解——读一个故事或看它的视频,然后回答关于发生了什么的问题(包括各种抽象层次的问题)。
• 通过小学阅读课程(涉及阅读并回答关于绘本与纯文本书的问题)。
• 仅凭经验、或经验加阅读说明来学习玩任意视频游戏。
• 通过儿童心理学家的典型评估——用于判断一个人类学龄前学生智力是否正常。
One thing we found at the AGI Roadmap Workshop was that each of these tests seems to some AGI researchers to encapsulate the crux of the AGI problem, and to be unsolvable by any system not far along the path to human-level AGI – yet seems to other AGI researchers, with different conceptual perspectives, to be something probably game-able by narrow-AI methods. And of course, given the current state of science, there's no way to tell which of these practical tests really can be solved via a narrow-AI approach, except by having a lot of researchers and engineers try really hard over a long period of time.
我们在 AGI 路线图研讨会上发现的一件事是:这些测试中的每一个,在一些 AGI 研究者看来都浓缩了 AGI 问题的核心、任何尚未在通往人类级 AGI 道路上走很远的系统都无法解决;然而在另一些概念视角不同的 AGI 研究者看来,它们又可能是窄 AI 方法可以钻空子的东西。当然,鉴于当前科学状况,除非让大量研究者与工程师在很长一段时间内非常努力地尝试,否则无法判断这些实用测试中哪些真能被窄 AI 方法解决。
5.3.1 评估机器学习能力的通用性
Metrics Assessing Generality of Machine Learning Capability
Complementing the above tests that are heavily inspired by human everyday life, there are also some more computer science oriented evaluation paradigms aimed at assessing AI systems going beyond specific tasks. For instance, there is a literature on multitask learning, where the goal for an AI is to learn one task quicker given another task solved previously. There is a literature on shaping, where the idea is to build up the capability of an AI by training it on progressively more difficult versions of the same tasks. Also, Achler has proposed criteria measuring the "flexibility of recognition" and posited this as a key measure of progress toward AGI.
While we applaud the work done in these areas, we also note it is an open question whether exploring these sorts of processes using mathematical abstractions, or in the domain of various machine-learning or robotics test problems, is capable of adequately addressing the problem of AGI. The potential problem with this kind of approach is that generalization among tasks, or from simpler to more difficult versions of the same task, is a process whose nature may depend strongly on the overall nature of the set of tasks and task-versions involved. Real-world humanly-relevant tasks have a subtlety of interconnectedness and developmental course that is not captured in current mathematical learning frameworks nor standard AI test problems.
与上述重度受人类日常生活启发的测试互补,还有一些更偏计算机科学的评估范式,旨在评估超越特定任务的 AI 系统。例如有关于多任务学习的文献——其目标是让 AI 在先解决一个任务后能更快学会另一个任务。有关于"塑形"(shaping)的文献——其想法是通过在越来越难的同任务版本上训练来构建 AI 能力。此外,Achler 提出了测量"识别灵活性"的标准,并把它作为朝 AGI 进展的关键度量。
虽然我们赞赏这些领域的工作,但也指出:用数学抽象、或在各种机器学习/机器人测试问题的领域内探索这类过程,能否充分解决 AGI 问题,仍是一个开放问题。这类方法潜在的问题是:任务间的泛化,或从同一任务更简单版本到更难版本的泛化,其本质可能强烈依赖于所涉任务集合与任务版本集合的整体性质。现实世界的、与人类相关的任务,其互联性与发展路径的微妙之处,是当前数学学习框架与标准 AI 测试问题都未捕捉到的。
To put it a little differently, it is possible that all of the following hold:
• the universe of real-world human tasks may possess a host of "special statistical properties" that have implications regarding what sorts of AI programs will be most suitable
• exploring and formalizing and generalizing these statistical properties is an important research area; however,
• an easier and more reliable approach to AGI testing is to create a testing environment that embodies these properties implicitly, via constituting an emulation of the most cognitively meaningful aspects of the real-world human learning environment
Another way to think about these issues is to contrast the above-mentioned "AGI Roadmap Workshop" ideas with the "General Game Player (GGP)" AI competition, in which AIs seek to learn to play games based on formal descriptions of the rules. Clearly doing GGP well requires powerful AGI; and doing GGP even mediocrely probably requires robust multitask learning and shaping. But it is unclear whether GGP constitutes a good approach to testing early-stage AI programs aimed at roughly humanlike intelligence. This is because, unlike the tasks involved in, say, making coffee in an arbitrary house, or succeeding in preschool or university, the tasks involved in doing simple instances of GGP seem to have little relationship to humanlike intelligence or real-world human tasks.
换一种说法,以下所有情况都可能成立:
• 现实世界人类任务的宇宙,可能拥有大量"特殊的统计性质",它们暗示着什么样的 AI 程序将最合适;
• 探索、形式化并泛化这些统计性质是一个重要的研究领域;然而,
• 一种更简单、更可靠的 AGI 测试路径,是创建一个内隐体现这些性质的测试环境——通过构建对现实世界人类学习环境最具认知意义方面的模拟来实现。
看待这些问题的另一个方式是:把上文提到的"AGI 路线图研讨会"想法与"通用游戏玩家(GGP)"AI 竞赛对照——GGP 中 AI 试图基于规则的正式描述学习玩游戏。显然,把 GGP 做得好需要强大的 AGI;把 GGP 做得平庸也可能需要稳健的多任务学习与塑形。但 GGP 是否是测试"面向大致类人智能的早期 AI 程序"的好方法,并不清楚。这是因为——与"在任意房子里煮咖啡""在学前班或大学成功"等任务不同——做简单 GGP 实例所涉的任务,似乎与类人智能或现实世界的人类任务几乎没有关系。
So, an important open question is whether the class of statistical biases present in the set of real-world human environments tasks, has some sort of generalizable relevance to AGI beyond the scope of human-like general intelligence, or is informative only about the particularities of human-like intelligence. Currently we seem to lack any solid, broadly accepted theoretical framework for resolving this sort of question.
因此,一个重要的开放问题是:现实世界人类环境/任务集合中存在的统计偏差类别,是否对人类类通用智能范围之外的 AGI 具有某种可泛化的相关性,还是只对人类类智能的特殊性有信息量。目前我们似乎缺乏任何扎实、被广泛接受的理论框架来解决这类问题。
5.3.2 为什么测量朝 AGI 的增量进展这么难?
Why Is Measuring Incremental Progress Toward AGI So Hard?
A question raised by these various observations is whether there is some fundamental reason why it's hard to make an objective, theory-independent measure of intermediate progress toward advanced AGI, which respects the environment and task biased nature of human intelligence as well as the mathematical generality of the AGI concept. Is it just that we haven't been smart enough to figure out the right test – or is there some conceptual reason why the very notion of such a test is problematic?
Why might a solid, objective empirical test for intermediate progress toward humanly meaningful AGI be such a difficult project? One possible reason could be the phenomenon of "cognitive synergy" briefly noted above. In this hypothesis, for instance, it might be that there are 10 critical components required for a human-level AGI system. Having all 10 of them in place results in human-level AGI, but having only 8 of them in place results in having a dramatically impaired system – and maybe having only 6 or 7 of them in place results in a system that can hardly do anything at all.
这些观察提出的一个问题是:为什么很难做出一个客观、独立于理论的"朝高级 AGI 中期进展"的度量——它既尊重人类智能"环境与任务偏向"的本质,又尊重 AGI 概念的数学通用性?是因为我们还不够聪明、没找到正确的测试——还是有某种概念上的原因,使"这样的测试"这一概念本身就有问题?
为什么一个扎实、客观的"朝对人类有意义的 AGI 中期进展"经验测试会是一个如此困难的项目?一个可能的原因是上文简要提到的"认知协同"现象。在这一假说下,例如,一个人类级 AGI 系统可能需要 10 个关键组件。10 个全部就位才产生人类级 AGI;只有 8 个就位会产生严重受损的系统;而只有 6 或 7 个就位,可能产生一个几乎什么都做不了的系统。
Of course, the reality is not as strict as the simplified example in the above paragraph suggests. No AGI theorist has really posited a list of 10 crisply-defined subsystems and claimed them necessary and sufficient for AGI. We suspect there are many different routes to AGI, involving integration of different sorts of subsystems. However, if the cognitive synergy hypothesis is correct, then human-level AGI behaves roughly like the simplistic example in the prior paragraph suggests. Perhaps instead of using the 10 components, you could achieve human-level AGI with 7 components, but having only 5 of these 7 would yield drastically impaired functionality – etc. To mathematically formalize the cognitive synergy hypothesis becomes complex, but here we're only aiming for a qualitative argument. So for illustrative purposes, we'll stick with the "10 components" example, just for communicative simplicity.
当然,现实并不像上文简化例子那么严格。没有 AGI 理论家真正提出一份 10 个清晰定义子系统的清单并宣称它们对 AGI 必要且充分。我们怀疑通往 AGI 有很多不同路线,涉及不同类型子系统的整合。然而,如果认知协同假说正确,那么人类级 AGI 的行为就大致如上段简化例子所示。也许不是 10 个组件,你可以用 7 个组件达成人类级 AGI,但这 7 个中只有 5 个会产出严重受损的功能——等等。把认知协同假说数学形式化会变得复杂,但这里我们只求一个定性论证。因此为说明起见,我们坚持用"10 个组件"的例子,仅为交流简单。
Next, let's additionally suppose that for any given task, there are ways to achieve this task using a system that is much simpler than any subset of size 6 drawn from the set of 10 components needed for human-level AGI, but works much better for the task than this subset of 6 components (assuming the latter are used as a set of only 6 components, without the other 4 components).
Note that this additional supposition is a good bit stronger than mere cognitive synergy. For lack of a better name, I have called this hypothesis "tricky cognitive synergy". Tricky cognitive synergy would be the case if, for example, the following possibilities were true:
• creating components to serve as parts of a synergetic AGI is harder than creating components intended to serve as parts of simpler AI systems without synergetic dynamics
• components capable of serving as parts of a synergetic AGI are necessarily more complicated than components intended to serve as parts of simpler AI systems
接下来,让我们再假设:对于任何给定任务,存在用某个系统完成该任务的方法——该系统比从"人类级 AGI 所需的 10 个组件"中任意取出的 6 元子集简单得多,但对该任务的表现却比这 6 个组件(假设它们只作为 6 个组件、没有其余 4 个)好得多。
注意,这一额外假设比单纯认知协同强得多。由于缺乏更好的名字,我把这一假说称为"棘手的认知协同"(tricky cognitive synergy)。如果例如以下可能性成立,就属于棘手的认知协同:
• 创造作为协同 AGI 一部分的组件,比创造用于无协同动力学、更简单 AI 系统一部分的组件更难;
• 能够作为协同 AGI 一部分的组件,必然比用于更简单 AI 系统一部分的组件更复杂。
These certainly seem reasonable possibilities, since to serve as a component of a synergetic AGI system, a component must have the internal flexibility to usefully handle interactions with a lot of other components as well as to solve the problems that come its way.
If tricky cognitive synergy holds up as a property of human-level general intelligence, the difficulty of formulating tests for intermediate progress toward human-level AGI follows as a consequence. Because, according to the tricky cognitive synergy hypothesis, any test is going to be more easily solved by some simpler narrow AI process than by a partially complete human-level AGI system.
At the current stage in the development of AGI, we don't really know how big a role "tricky cognitive synergy" plays in the general intelligence. Quite possibly, 5 or 10 years from now someone will have developed wonderfully precise and practical metrics for the evaluation of incremental progress toward human-level AGI. However, it's worth carefully considering the possibility that fundamental obstacles, tied to the nature of general intelligence, stand in the way of this possibility.
这些当然是合理的可能性——因为要作为协同 AGI 系统的一个组件,组件必须具有内部灵活性:既要有效处理与大量其他组件的交互,也要解决迎面而来的问题。
如果"棘手的认知协同"作为人类级通用智能的一个属性成立,那么"为朝人类级 AGI 的中期进展制定测试"的困难就顺理成章了。因为根据这一假说,任何测试都更容易被某个更简单的窄 AI 过程解决,而不是被一个部分完成的人类级 AGI 系统解决。
在 AGI 发展的当前阶段,我们并不真正知道"棘手的认知协同"在通用智能中扮演多大角色。很可能 5 或 10 年后,有人会开发出奇妙精确且实用的度量,用于评估朝人类级 AGI 的增量进展。然而,值得仔细考虑这种可能性:与通用智能本质相连的根本性障碍,可能挡在这种可能性的路上。