第十章 · Section 10
Visual Processing (V) · 视觉处理
Visual Processing (V): The ability to analyze and generate natural and unnatural images and videos.
视觉处理(V):分析并生成自然与非自然图像和视频的能力。
The ability to process and interpret visual inputs from images and videos.
• Image Recognition — "Identify the picture." / "What does this image depict?"
• Image Captioning — "Create descriptive caption for this image."
• Image Anomaly Detection — "Which is the odd one out?"
• Clip Captioning — "What happens in this video?"
• Video Anomaly Detection — "Is this physically plausible?"
Simple Natural Images · Complicated Images · Simple Natural Videos
处理并解读来自图像与视频的视觉输入的能力。
• 图像识别——"识别这张图片。" / "这幅图像描绘了什么?"
• 图像描述——"为这幅图像生成描述性文字说明。"
• 图像异常检测——"哪一个是异类?"
• 片段描述——"这个视频里发生了什么?"
• 视频异常检测——"这在物理上合理吗?"
简单自然图像 · 复杂图像 · 简单自然视频
The ability to synthesize images and short videos.
Sample: "Generate an image of a golden retriever playing in a park." / "Generate a diagram showing the process of photosynthesis." / "Generate a short video of somebody typing on a keyboard."
合成图像与短视频的能力。
样例:"生成一张金毛在公园玩耍的图像。" / "生成一张展示光合作用过程的示意图。" / "生成一段某人在键盘上打字的短视频。"
The ability to understand and make inferences about the images.
• Gestalt — "Which shape on the right is the same as the shape on the left?"
• Mental Rotation / Mental Folding — "Which net, when folded, cannot form the cube?"
• Embodied Reasoning — "Which trajectories should the zipper follow to zip the suitcase?"
• Chart and Figure Reasoning — "What is the lowest labeled tick on the y-axis?"
Also: "Count the people in the picture." / "Find the path to the center of this maze."
理解图像并对其做出推断的能力。
• 完形(Gestalt)——"右边哪个形状与左边相同?"
• 心理旋转 / 心理折叠——"哪个展开图折叠后不能形成立方体?"
• 具身推理——"拉链应沿哪条轨迹才能拉上这个行李箱?"
• 图表推理——"y 轴上标注的最低刻度是多少?"
另有:"数一数图中的人数。" / "找到通向这个迷宫中心的路径。"
The speed and accuracy of visually exploring a complex field.
视觉探索复杂视场(field)的速度与准确性。
评估细节与 AI 表现
Assessment Details & AI Performance
Assessment Details. See Appendix H for further details on how to assess visual processing capabilities concretely.
AI System Performance. The table summarizes current AI system performance on Visual Processing (V) tasks. GPT-4 had no ability to perceive or generate images, while GPT-5 has appreciable but highly incomplete visual processing capabilities.
Model | Perception (4%) | Generation (3%) | Reasoning (2%) | Spatial Scanning (1%) | Total
GPT-4 | 0% | 0% | 0% | 0% | 0%
GPT-5 | 2% | 2% | 0% | 0% | 4%
评估细节:如何在具体层面评估视觉处理,参见附录 H。
AI 系统表现:下表汇总当前 AI 系统在视觉处理(V)任务上的表现。GPT-4 完全没有感知或生成图像的能力;GPT-5 有可观但高度不完整的视觉处理能力。
模型 | 感知(4%) | 生成(3%) | 推理(2%) | 空间扫描(1%) | 合计
GPT-4 | 0% | 0% | 0% | 0% | 0%
GPT-5 | 2% | 2% | 0% | 0% | 4%