3.Understanding is the Mainline of AGI

Published:

It is the Understanding instead of Generation really makes sense in Computer Vision.


在2023 Future Science Prize Laureate Lecture上,何恺明提及要做一个“大大大”即所谓Scalable的Computer Vision Model。而在其中往往被忽视的一点是,我们并没有一个对这个Scalable CV模型的well-defined的定义——我们所需要scalable的能力究竟是什么?分割,分类,检测还是生成?

我们先将分类和分割定义为“理解”类任务的子任务。从 Diception 到 VisionBanana,这些工作都试图通过生成的表征去辅助理解的表征。UAE(Generation is Understanding)同样试图去证明好的理解和生成是统一的。在JEPA的中,pixel本身也仍然需要进入一个encoder之后再做latent space上的对齐。我们可以大胆地做出此设想:JEPA对齐的是理解的表征而非像素级别的生成。

从一个更感性的认知来说,绝大多数人具备认出熟人面庞的能力,却鲜有重建对方脸部的能力。对于人脑这样的一个World Model来说,我们需要的是理解,而非重建与生成。

但在CV领域有一个共识,生成往往需要高分辨率的对像素的认识,而此却与理解对图片内在结构而非像素细节的需要发生了矛盾。所以更重要的一点可能是,CV的表征应该是Hierarchical的,是层次化的,生成的表征可能分布在相对浅层的layer而理解却稍晚。

与此同时,我们也应该认识到一点:纯的CV task是没有意义的,而CV只在Agentic或Robotics的语境下才真正make sense。即,CV只在Understanding的语境下才有意义。就像梁文峰的认识:生成不是AGI的主线。所以与其说我们要做的是统一理解与生成,还不如去纯粹地去理解“理解”的表征。

理解才是AGI的主线。


Generation Is Not Understanding

As someone with a JEPA bias, I find the recent “generation as understanding” story compelling — but also too easy to overstate.

One question I took away from Kaiming He’s 2023 Future Science Prize lecture is this: if computer vision is going to scale, what exactly are we trying to scale?

Segmentation? Classification? Generation?

For a long time, classification and segmentation were the operational proxies for visual understanding. Recently, the field has been moving in another direction. DICEPTION, VisionBanana, and UAE all point to a similar possibility: many perception and multimodal tasks can be expressed through generation.

This is an important shift. DICEPTION and VisionBanana show that if we represent task outputs as images — masks, depth, normals, or other visual structures — a generator can become a generalist vision learner. UAE goes further and tries to connect understanding and generation under an auto-encoding view: understanding as encoding, generation as decoding.

I think this direction is valuable. But it also makes the distinction more important, not less.

Generation is an interface. Understanding is the structure that makes the interface useful.

A model may learn strong visual representations because generation forces it to model objects, layout, geometry, and relations. But the fact that a task can be cast as generation does not mean generation itself is understanding. It only shows that generation can expose, reuse, and sometimes organize perceptual structure.

That is the nuance I think we should keep.

Computer vision may no longer have a single “pure” task. It will increasingly be coupled with language, robotics, and world models. But the center should still be the representation that supports recognition, grounding, planning, and action.

Generation can be a powerful way to learn that representation.

But generation is not understanding.