← 中英对照目录 · ← 书架
第一章 · Section 1
Introduction · 引言
Artificial General Intelligence (AGI) may become the most significant technological development in human history, yet the term itself remains frustratingly nebulous, acting as a constantly moving goalpost. As specialized AI systems master tasks once thought to require human intellect—from mathematics to art—the criteria for "AGI" continually shift. This ambiguity fuels unproductive debates, hinders discussions about how far AGI is, and ultimately obscures the gap between today's AI and AGI.
通用人工智能(AGI)可能成为人类历史上最重大的技术发展,但"AGI"这个术语本身却仍然令人沮丧地模糊,像一根不断移动的标杆。随着专用 AI 系统攻克曾经被认为需要人类智能的任务——从数学到艺术——"AGI"的标准不断上移。这种模糊性助长了无益的争论,阻碍了关于"AGI 离我们多远"的讨论,最终模糊了今天的 AI 与 AGI 之间的差距。
This paper provides a comprehensive, quantifiable framework to cut through the ambiguity. Our framework aims to concretely specify the informal definition:
本文提供了一个全面、可量化的框架来穿透这种模糊性。我们的框架旨在把下面这个非正式定义具体化:
AGI is an AI that can match or exceed the cognitive versatility and proficiency of a well-educated adult.
AGI 是一种能匹配或超越"受过良好教育的成年人"的认知通用性与熟练度的人工智能。
This definition emphasizes that general intelligence requires not just specialized performance in narrow domains, but the breadth (versatility) and depth (proficiency) of skills that characterize human cognition.
这一定义强调:通用智能需要的不仅是在狭窄领域的专项表现,更是构成人类认知特征的技能的广度(通用性)与深度(熟练度)。
To operationalize this definition, we must look to the only existing example of general intelligence: humans. Human cognition is not a monolithic capability; it is a complex architecture composed of many distinct abilities honed by evolution. These abilities enable our remarkable adaptability and understanding of the world.
要把这一定义落地为可操作的标准,我们必须看向通用智能唯一现存的实例:人类。人类认知不是一种单一的能力;它是由进化打磨出的许多不同能力组成的一台复杂架构。正是这些能力赋予了我们非凡的适应力与对世界的理解。
To systematically investigate whether AI systems possess this spectrum of abilities, we ground our approach in the Cattell-Horn-Carroll (CHC) theory of cognitive abilities (Carroll, 1993; McGrew, 2009; Schneider and McGrew, 2018; McGrew, 2023; McGrew et al., 2023), the most empirically validated model of human intelligence. CHC theory is primarily derived from the synthesis of over a century of iterative factor analysis of diverse collections of cognitive ability tests. In the late 1990's to 2000's almost all major clinical, individually administered tests of human intelligence have iterated towards test revisions that were either explicitly or implicitly based on CHC model test design blueprints (Keith and Reynolds, 2010; Schneider and McGrew, 2018). CHC theory provides a hierarchical taxonomic map of human cognition. It breaks down general intelligence into distinct broad abilities and numerous narrow abilities (such as induction, associative memory, or spatial scanning). Readers interested in the strengths and limitations of the CHC framework are directed to further scholarly discussions (Wasserman, 2019; Canivez and Youngstrom, 2019).
为了系统性地考察 AI 系统是否具备这整套能力,我们以卡特-霍恩-卡罗尔(CHC)认知能力理论为方法根基——这是实证验证最充分的人类智能模型。CHC 理论主要源自对一百多年间、多批认知能力测验反复做因子分析的合成结果。1990 年代末到 2000 年代,几乎所有主要临床、个体施测的人类智力测验,都迭代向"明确或隐式基于 CHC 模型测验设计蓝图"的修订版靠拢。CHC 理论提供了一张人类认知的层级分类图:它把通用智能拆解为不同的宽泛能力(broad abilities)与大量狭窄能力(narrow abilities,如归纳、联想记忆、空间扫描)。对 CHC 框架优势与局限感兴趣的读者,可参阅相关学术讨论。
把人的测验用于 AI:从"模糊概念"到"AGI 分数"
From Humans to AI: Turning Vague Notions into an AGI Score
Decades of psychometric research have yielded a vast battery of tests specifically designed to isolate and measure these distinct cognitive components in individuals. Our framework adapts this methodology for AI evaluation. Instead of relying solely on generalized tasks that might be solved through compensatory strategies, we systematically investigate whether AI systems possess the underlying CHC narrow abilities that humans have. To determine whether an AI has the cognitive versatility and proficiency of a well-educated adult, we test the AI system with the gauntlet of cognitive batteries used to test people. This approach replaces nebulous concepts of intelligence with concrete measurements, resulting in a standardized "AGI Score" (0% to 100%), in which 100% signifies AGI.
数十年的心理测量学研究产出了海量的测验套件,专门用于在个体身上分离并测量这些不同的认知成分。我们的框架把这一方法论适配到 AI 评估上。我们不是只依赖可能被"补偿性策略"钻空子的泛化任务,而是系统地考察 AI 系统是否具备人类所拥有的底层 CHC 狭窄能力。为了判断一个 AI 是否具备"受过良好教育的成年人"的认知通用性与熟练度,我们用那套用于测试人的认知测验的"全阵列"去测试 AI 系统。这一方法用具体测量取代了模糊的智能概念,得出一个标准化的"AGI 分数"(0% 到 100%),其中 100% 表示达成 AGI。
The application of this framework is revealing. By testing the fundamental abilities that underpin human cognition—many of which appear simple for humans—we find that contemporary AI systems can solve roughly half of these often-simple assessments. This indicates that despite impressive performance on complex benchmarks, current AI lacks many of the core cognitive capabilities essential for human-like general intelligence. Current AIs are narrower than well-educated humans overall but far smarter on some specific tasks.
应用这一框架的结果很有启发性。通过测试支撑人类认知的基础能力——其中许多在人类看来很简单——我们发现,当代 AI 系统只能解决这些通常很简单的测验中的大约一半。这表明,尽管当代 AI 在复杂基准上表现惊艳,它仍然缺乏类人通用智能所需的许多核心认知能力。当前 AI 整体上比受过良好教育的成年人更"窄",但在某些特定任务上又远比人聪明。
十认知域:框架的十个核心成分
The Ten Core Cognitive Components
The framework comprises ten core cognitive components, derived from CHC broad abilities and weighted equally (10%) to prioritize breadth and cover major areas of cognition:
框架由十个核心认知成分组成,它们源自 CHC 宽泛能力,并被等权加权(各占 10%),以优先保证广度、覆盖认知的主要领域:
• General Knowledge (K): The breadth of factual understanding of the world, encompassing commonsense, culture, science, social science, and history.
• 常识(K):对世界的事实性理解的广度,涵盖常识、文化、科学、社会科学与历史。
• Reading and Writing Ability (RW): Proficiency in consuming and producing written language, from basic decoding to complex comprehension, composition, and usage.
• 读写能力(RW):消费与生产书面语言的熟练度,从基础解码到复杂的理解、写作与运用。
• Mathematical Ability (M): The depth of mathematical knowledge and skills across arithmetic, algebra, geometry, probability, and calculus.
• 数学能力(M):跨算术、代数、几何、概率与微积分的数学知识与技能的深度。
• On-the-Spot Reasoning (R): The flexible control of attention to solve novel problems without relying exclusively on previously learned schemas, tested via deduction and induction.
• 即时推理(R):灵活控制注意力去解决"现场"新问题,而非只依赖先前习得的模式;通过演绎与归纳来测试。
• Working Memory (WM): The ability to maintain and manipulate information in active attention across textual, auditory, and visual modalities.
• 工作记忆(WM):在文本、听觉与视觉等模态下,于活跃注意力中保持并操作信息的能力。
• Long-Term Memory Storage (MS): The capability to continually learn new information (associative, meaningful, and verbatim).
• 长时记忆存储(MS):持续学习新信息的能力(联想式、意义式与逐字式)。
• Long-Term Memory Retrieval (MR): The fluency and precision of accessing stored knowledge, including the critical ability to avoid confabulation (hallucinations).
• 长时记忆检索(MR):存取已存储知识的流畅度与精确度,包括关键的"避免虚构(幻觉)"能力。
• Visual Processing (V): The ability to perceive, analyze, reason about, generate, and scan visual information.
• 视觉处理(V):感知、分析、推理、生成与扫描视觉信息的能力。
• Auditory Processing (A): The capacity to discriminate, recognize, and work creatively with auditory stimuli, including speech, rhythm, and music.
• 听觉处理(A):辨别、识别并创造性地处理听觉刺激的能力,包括言语、节奏与音乐。
• Speed (S): The ability to perform simple cognitive tasks quickly, encompassing perceptual speed, reaction times, and processing fluency.
• 速度(S):快速完成简单认知任务的能力,涵盖知觉速度、反应时间与加工流畅度。
This operationalization provides a holistic and multimodal (text, visual, auditory) assessment, serving as a rigorous diagnostic tool to pinpoint the strengths and profound weaknesses of current AI systems.
这一操作化方案提供了一种整体性、多模态(文本、视觉、听觉)的评估,充当一件严格的诊断工具,用来精准指出当前 AI 系统的优势与深层弱点。
首批打分结果
Initial Scores
Table 1: AGI Score Summary for GPT-4 (2023) and GPT-5 (2025). Model K RW M R WM MS MR V A S Total. GPT-4 8% 6% 4% 0% 2% 0% 4% 0% 0% 3% 27%. GPT-5 9% 10% 10% 7% 4% 0% 4% 4% 6% 3% 57%.
表 1:GPT-4(2023)与 GPT-5(2025)的 AGI 分数汇总。十域:K 常识 / RW 读写 / M 数学 / R 即时推理 / WM 工作记忆 / MS 长时记忆存储 / MR 长时记忆检索 / V 视觉 / A 听觉 / S 速度。GPT-4:8% / 6% / 4% / 0% / 2% / 0% / 4% / 0% / 0% / 3% = 27%。GPT-5:9% / 10% / 10% / 7% / 4% / 0% / 4% / 4% / 6% / 3% = 57%。
Scope:测什么、不测什么
Scope: What We Measure and What We Do Not
Scope. Our definition is not an automatic evaluation nor a dataset. Rather, it specifies a large collection of well-scoped tasks that test specific cognitive abilities. Whether AIs can solve these tasks can be manually assessed by anyone, and people could supplement their testing using the best evaluations available at the time. This makes our definition more broad and more robust than fixed automatic AI capabilities datasets. Secondly, our definition focuses on capabilities frequently possessed by well-educated individuals, not a superhuman aggregate of all well-educated individuals' combined knowledge and skills. Therefore, our AGI definition is about human-level AI, not economy-level AI; we measure cognitive abilities rather than specialized economically valuable know-how, nor is our measurement a direct predictor of automation or economic diffusion. We leave economic measurements of advanced AI to other work. Last, we deliberately focus on core cognitive capabilities rather than physical abilities such as motor skills or tactile sensing, as we seek to measure the capabilities of the mind rather than the quality of its actuators or sensors. We discuss more limitations in the Discussion.
范围。我们的定义既不是一项自动评测,也不是一个数据集。相反,它规定了一大批边界清晰的、测试特定认知能力的任务。AI 能否解决这些任务,任何人都可以人工评估;人们也可以用当时最好的评测工具来补充测试。这让我们的定义比固定的自动 AI 能力数据集更宽泛、更强健。其次,我们的定义聚焦于"受过良好教育的个体"普遍具备的能力,而不是把所有受过良好教育的个体知识技能叠加起来的"超人集合"。因此,我们的 AGI 定义是关于"人类级 AI",而不是"经济级 AI";我们测量的是认知能力,而非专门的有经济价值的技能诀窍;我们的测量也不是自动化或经济扩散的直接预测指标。我们把高级 AI 的经济测量留给其他研究。最后,我们有意聚焦于核心认知能力,而非运动技能、触觉感知等身体能力——我们要测量的是"心智"的能力,而不是其执行器或传感器的质量。更多局限我们在"讨论"一节展开。
← 主页