At the International Conference on Machine Learning (ICML 2026), Dezhi Luo, Yijiang Li, Maijunxian Wang, Tianwei Zhao, Bingyang Wang, Siheng Wang, Pinyuan Feng, Pooyan Rahmanzadehgervi, Ziqiao Ma, and Hokin Deng presented their work Vision Language Models Cannot Reason About Physical Transformation. In this interview, William Yijiang Li tells us more about the research, the method the team used, and the controlled experiments they carried out.
Our research asks a fairly fundamental question: do vision-language models actually understand how the physical world changes over time?
Modern vision-language models are increasingly being used in areas such as robotics, embodied agents, and video understanding. They can recognize objects, describe scenes, and answer surprisingly sophisticated questions. However, being able to describe a scene is different from maintaining an understanding of an object or physical quantity as the scene changes.
We study this through the concept of conservation. For example, if water is poured from a short, wide glass into a tall, narrow one, its appearance changes dramatically, but the amount of water remains the same. Similarly, spreading a row of coins farther apart does not change the number of coins.
These tasks are simple for humans, but they require an important form of reasoning: you need to track what happens during a transformation and distinguish changes in appearance from changes in the underlying physical state.
This is particularly interesting for embodied AI. If an AI system is expected to interact with the real world, it cannot rely only on static visual recognition. It needs to maintain and update representations of objects and their properties as events unfold.
I would summarize the contributions in three parts.
First, and most importantly, our study shows that current vision-language models still struggle to maintain and update a stable representation of the physical world as it changes over time. Although these models can often answer basic questions about physical concepts such as conservation correctly, our controlled experiments show that this does not necessarily mean they understand the underlying transformation. Models that perform well when a quantity is conserved often fail on nearly identical cases where that quantity actually changes. This leads to our central conclusion: knowing a physical rule is not the same as being able to ground and apply that rule to a dynamic visual process. The key bottleneck appears to be continuously tracking object states and incorporating new visual evidence over time.
Second, we introduce ConservationBench, a controlled benchmark designed to study this capability systematically. It covers four fundamental quantitative properties — number, length, volume, and size — and, crucially, includes a matched non-conserving control for every conserving scenario. For example, in one video a row of coins is simply spread apart, while in its matched control a coin is actually added during the transformation. This paired design allows us to distinguish genuine transformation reasoning from a simple prior that quantities usually remain unchanged.
Third, we conduct a broad set of controlled experiments to understand why models fail. We vary temporal resolution, frame selection, prompting strategies, and the availability of visual evidence, and we also examine model confidence and attention patterns. Together, these experiments show that the failures cannot be explained simply by insufficient frames, poor prompting, or model scale. Instead, the evidence consistently suggests that many models rely on strong linguistic or perceptual heuristics rather than building a persistent representation of the physical state and updating it as the event unfolds.
Overall, the contribution is therefore not just a new benchmark or the observation that models make mistakes. It is evidence for a more fundamental limitation in current multimodal models: they can possess knowledge about the physical world without reliably representing how that world evolves over time.
We build a controlled diagnostic framework for probing whether vision-language models can maintain and update representations of physical quantities through dynamic transformations. Drawing on conservation as a canonical test of state persistence, we construct paired visual scenarios that isolate genuine transformation reasoning from static perception, linguistic priors, and shortcut-based prediction.
Each example contains an initial state, a physical transformation, and a final state. For number, for example, two rows initially contain the same number of coins. One row is then spread out. In the conservation condition, no coins are added or removed. In the matched non-conserving condition, the visual setup is almost identical, but a coin is actually added during the transformation.
We follow the same principle for four domains:
For every domain, we systematically vary irrelevant visual factors such as colors, object numbers, layouts, container shapes, directions, and transformations. This helps prevent models from solving the benchmark through memorized visual templates.
We then provide models with multiple frames sampled from each video and ask them to judge the final quantities. Importantly, some tasks can partly be solved by looking at the beginning and end, while others—particularly volume and size—really require understanding what happened during the transformation.
This lets us test not simply whether a model can perceive the final image, but whether it can integrate visual information over time and update an internal representation of the physical state.
The controls are probably the most important part of the study because a high accuracy number by itself can be misleading. A central aspect of our study is that we do not treat benchmark accuracy as sufficient evidence of physical reasoning. Instead, we design a series of controlled experiments to distinguish genuine understanding of dynamic transformations from simpler strategies such as response priors, static visual cues, or insufficient temporal context.
Our primary control is the matched non-conserving condition. For every conserving example, we construct a closely matched counterpart in which the relevant quantity actually changes while the overall visual structure remains similar. A model that genuinely tracks the transformation should therefore perform well on both conditions. However, we observe the opposite pattern in many models: strong performance on conserving cases is often accompanied by poor performance on their non-conserving counterparts. This indicates that high conservation accuracy alone can substantially overestimate a model’s underlying reasoning ability.
We then systematically test several alternative explanations for this failure.
First, we vary the amount of temporal evidence available to the model by providing 3, 5, 7, 9, or 16 frames. If the primary limitation were incomplete observation of the transformation, performance should improve consistently as more frames are provided. In practice, increasing temporal resolution produces little systematic improvement.
Second, we examine whether the issue arises from suboptimal frame selection. We compare uniform sampling with frames selected by human annotators and by a learned video-localization model. Even when the model is provided with frames judged to be maximally informative, the underlying failure pattern largely persists.
Third, we evaluate whether more explicit reasoning instructions can compensate for the limitation. We test direct prompting, sequential frame-by-frame reasoning, chain-of-thought prompting, and prompts that explicitly emphasize temporal continuity across frames. While continuity-aware prompting provides modest gains in some settings, no prompting strategy consistently resolves the problem, and chain-of-thought can in some cases reduce performance.
We also conduct a particularly diagnostic set of no-visual-evidence controls, in which the original visual input is replaced by blank images or removed entirely. Under these conditions, models still exhibit a strong tendency to predict that quantities are conserved. This demonstrates that part of their apparent success in the standard setting is driven by a pre-existing response prior rather than by visual evidence.
More strikingly, on some conservation tasks, performance can actually improve when the informative visual content is removed. This suggests that the models are not simply failing to extract enough information from the video; rather, the visual evidence can conflict with a strong prior that the model is unable to appropriately revise.
Finally, we evaluate a range of additional controls, including explicit descriptions of conservation principles, textual captions of the visual sequence, an “I don’t know” response option, model scaling, and models specifically post-trained for physical reasoning. None of these interventions eliminates the core failure pattern.
Taken together, these experiments allow us to rule out several straightforward explanations—insufficient temporal coverage, poor frame selection, prompting, lack of declarative physical knowledge, or model scale—and instead point to a more fundamental limitation in how current vision-language models represent and update physical state over time.
The clearest finding is that current vision-language models do not reliably reason about physical transformations, even when the underlying task is extremely simple for humans.
Across 112 models, most systems were only modestly above the 33.3% chance level when performance was balanced across conserving and non-conserving scenarios. Human participants, by comparison, achieved approximately 98% accuracy.
A particularly important result is that conservation accuracy alone gives a misleading picture. There was actually a negative correlation between performance on conserving and non-conserving tasks: models that appeared very good at recognizing conservation were often particularly bad at noticing when the quantity really changed.
When we required a model to answer both members of a matched conserving/non-conserving pair correctly, 82 of the 112 models achieved below 10% accuracy. Only a handful of the strongest models performed above chance under this stricter criterion.
We also found very little evidence that simply scaling the models solves the problem. Performance on conservation tasks was almost unrelated to parameter count. Adding more frames, using more sophisticated prompts, or selecting more informative frames also failed to produce consistent improvements.
Our mechanistic analysis provides a possible explanation. In one model we studied in detail, incorrect non-conservation judgments were not hesitant guesses—they were often made with higher confidence than correct judgments. Attention analysis also showed that, during these failures, the model disproportionately focused on the initial frame and failed to sufficiently incorporate information appearing later in the transformation.
So our interpretation is that the fundamental problem is not simply that these models lack a fact such as “volume is conserved.” They often possess that linguistic knowledge. The harder problem is that they do not reliably maintain and update an object-state representation over time.
That distinction is important. A model can know the rule in language without being able to ground and apply that rule to a dynamic visual event.
There are several directions that we think are particularly important.
One is to move beyond these deliberately controlled experiments toward more realistic physical environments. Our current benchmark focuses on four basic quantities under relatively clean conditions. Real-world physical reasoning also involves occlusion, deformable objects, noisy observations, interactions between multiple objects, and transformations whose consequences may be uncertain.
A second direction is to understand the failure mechanistically. Our initial attention and confidence analyses suggest that models often fail to update their representation after observing a transformation, but this needs to be tested more systematically across different model architectures and with causal interventions.We are also interested in studying whether these seemingly simple conservation failures propagate into more consequential embodied tasks—for example, planning, manipulation, tool use, or predicting the consequences of actions. A model might perform well on a high-level robotics benchmark while still relying on shortcuts that fail under small changes in the environment.
More broadly, I think this work raises a question about the representations that future multimodal models should learn. Most current systems are extremely strong at extracting static semantic features from images. But physical reasoning may require something different: persistent, predictive representations of objects and their states that can be updated as the world changes. Developing models with that kind of temporally grounded internal representation is one of the directions we are particularly interested in pursuing.
|
William Yijiang Li is a 2nd year PhD student and a Jacobs Fellow at UC San Diego. His research focuses on multi-modal LLM, Agents, World models and Embodiment, specifically the learning aspects of AI – to enable efficient (e.g. label efficiency, sample effiency, and synthetic data) and robust learningin multi-modal, interactive and 3D embodied environments. He has published over 10 top conference papers on these topics including ICML, Neurips, ICLR, NACCL, ACL, ICCV, ICLR, CVPR, and EMNLP, organized several workshops in leading conferences such as ES-Reasoning @ ICLR 2026, the 2nd Efficient-Reasoning @ COLM 2026, and the 2nd MMRAgI @ CVPR 2026, and co-founded the GrowAI, a research organization and community aiming to promote the research of AI that can learn and grow like humans. |