← 中英对照目录 · ← 书架
第十章 · Section 10
Visual Processing (V) · 视觉处理
Visual Processing (V): The ability to analyze and generate natural and unnatural images and videos.
视觉处理(V):分析并生成自然与非自然图像和视频的能力。
感知
Perception
The ability to process and interpret visual inputs from images and videos. • Image Recognition — "Identify the picture." / "What does this image depict?" • Image Captioning — "Create descriptive caption for this image." • Image Anomaly Detection — "Which is the odd one out?" • Clip Captioning — "What happens in this video?" • Video Anomaly Detection — "Is this physically plausible?" Simple Natural Images · Complicated Images · Simple Natural Videos
处理并解读来自图像与视频的视觉输入的能力。 • 图像识别——"识别这张图片。" / "这幅图像描绘了什么?" • 图像描述——"为这幅图像生成描述性文字说明。" • 图像异常检测——"哪一个是异类?" • 片段描述——"这个视频里发生了什么?" • 视频异常检测——"这在物理上合理吗?" 简单自然图像 · 复杂图像 · 简单自然视频
视觉生成
Visual Generation
The ability to synthesize images and short videos. Sample: "Generate an image of a golden retriever playing in a park." / "Generate a diagram showing the process of photosynthesis." / "Generate a short video of somebody typing on a keyboard."
合成图像与短视频的能力。 样例:"生成一张金毛在公园玩耍的图像。" / "生成一张展示光合作用过程的示意图。" / "生成一段某人在键盘上打字的短视频。"
视觉推理
Visual Reasoning
The ability to understand and make inferences about the images. • Gestalt — "Which shape on the right is the same as the shape on the left?" • Mental Rotation / Mental Folding — "Which net, when folded, cannot form the cube?" • Embodied Reasoning — "Which trajectories should the zipper follow to zip the suitcase?" • Chart and Figure Reasoning — "What is the lowest labeled tick on the y-axis?" Also: "Count the people in the picture." / "Find the path to the center of this maze."
理解图像并对其做出推断的能力。 • 完形(Gestalt)——"右边哪个形状与左边相同?" • 心理旋转 / 心理折叠——"哪个展开图折叠后不能形成立方体?" • 具身推理——"拉链应沿哪条轨迹才能拉上这个行李箱?" • 图表推理——"y 轴上标注的最低刻度是多少?" 另有:"数一数图中的人数。" / "找到通向这个迷宫中心的路径。"
空间扫描
Spatial Scanning
The speed and accuracy of visually exploring a complex field.
视觉探索复杂视场(field)的速度与准确性。
评估细节与 AI 表现
Assessment Details & AI Performance
Assessment Details. See Appendix H for further details on how to assess visual processing capabilities concretely. AI System Performance. The table summarizes current AI system performance on Visual Processing (V) tasks. GPT-4 had no ability to perceive or generate images, while GPT-5 has appreciable but highly incomplete visual processing capabilities. Model | Perception (4%) | Generation (3%) | Reasoning (2%) | Spatial Scanning (1%) | Total GPT-4 | 0% | 0% | 0% | 0% | 0% GPT-5 | 2% | 2% | 0% | 0% | 4%
评估细节:如何在具体层面评估视觉处理,参见附录 H。 AI 系统表现:下表汇总当前 AI 系统在视觉处理(V)任务上的表现。GPT-4 完全没有感知或生成图像的能力;GPT-5 有可观但高度不完整的视觉处理能力。 模型 | 感知(4%) | 生成(3%) | 推理(2%) | 空间扫描(1%) | 合计 GPT-4 | 0% | 0% | 0% | 0% | 0% GPT-5 | 2% | 2% | 0% | 0% | 4%
← 主页