ΑΙhub.org
 

Interview with William Yijiang Li: vision language models and the physical world


by
05 October 2026



share this:

At the International Conference on Machine Learning (ICML 2026), Dezhi Luo, Yijiang Li, Maijunxian Wang, Tianwei Zhao, Bingyang Wang, Siheng Wang, Pinyuan Feng, Pooyan Rahmanzadehgervi, Ziqiao Ma, and Hokin Deng presented their work Vision Language Models Cannot Reason About Physical Transformation. In this interview, William Yijiang Li tells us more about the research, the method the team used, and the controlled experiments they carried out.

What is the topic of the research in your paper and why is it an interesting area for study?

Our research asks a fairly fundamental question: do vision-language models actually understand how the physical world changes over time?

Modern vision-language models are increasingly being used in areas such as robotics, embodied agents, and video understanding. They can recognize objects, describe scenes, and answer surprisingly sophisticated questions. However, being able to describe a scene is different from maintaining an understanding of an object or physical quantity as the scene changes.

We study this through the concept of conservation. For example, if water is poured from a short, wide glass into a tall, narrow one, its appearance changes dramatically, but the amount of water remains the same. Similarly, spreading a row of coins farther apart does not change the number of coins.

These tasks are simple for humans, but they require an important form of reasoning: you need to track what happens during a transformation and distinguish changes in appearance from changes in the underlying physical state.

This is particularly interesting for embodied AI. If an AI system is expected to interact with the real world, it cannot rely only on static visual recognition. It needs to maintain and update representations of objects and their properties as events unfold.

What are the main contributions of your research?

I would summarize the contributions in three parts.

First, and most importantly, our study shows that current vision-language models still struggle to maintain and update a stable representation of the physical world as it changes over time. Although these models can often answer basic questions about physical concepts such as conservation correctly, our controlled experiments show that this does not necessarily mean they understand the underlying transformation. Models that perform well when a quantity is conserved often fail on nearly identical cases where that quantity actually changes. This leads to our central conclusion: knowing a physical rule is not the same as being able to ground and apply that rule to a dynamic visual process. The key bottleneck appears to be continuously tracking object states and incorporating new visual evidence over time.

Second, we introduce ConservationBench, a controlled benchmark designed to study this capability systematically. It covers four fundamental quantitative properties — number, length, volume, and size — and, crucially, includes a matched non-conserving control for every conserving scenario. For example, in one video a row of coins is simply spread apart, while in its matched control a coin is actually added during the transformation. This paired design allows us to distinguish genuine transformation reasoning from a simple prior that quantities usually remain unchanged.

Third, we conduct a broad set of controlled experiments to understand why models fail. We vary temporal resolution, frame selection, prompting strategies, and the availability of visual evidence, and we also examine model confidence and attention patterns. Together, these experiments show that the failures cannot be explained simply by insufficient frames, poor prompting, or model scale. Instead, the evidence consistently suggests that many models rely on strong linguistic or perceptual heuristics rather than building a persistent representation of the physical state and updating it as the event unfolds.

Overall, the contribution is therefore not just a new benchmark or the observation that models make mistakes. It is evidence for a more fundamental limitation in current multimodal models: they can possess knowledge about the physical world without reliably representing how that world evolves over time.

Could you explain your methodology?

We build a controlled diagnostic framework for probing whether vision-language models can maintain and update representations of physical quantities through dynamic transformations. Drawing on conservation as a canonical test of state persistence, we construct paired visual scenarios that isolate genuine transformation reasoning from static perception, linguistic priors, and shortcut-based prediction.

Each example contains an initial state, a physical transformation, and a final state. For number, for example, two rows initially contain the same number of coins. One row is then spread out. In the conservation condition, no coins are added or removed. In the matched non-conserving condition, the visual setup is almost identical, but a coin is actually added during the transformation.

We follow the same principle for four domains:

  • Number: objects are rearranged, with or without objects being added or removed.
  • Length: an object is repositioned, with or without its physical length being changed.
  • Volume: liquid is poured between differently shaped containers, either completely or with some liquid deliberately left behind.
  • Size: playdough is reshaped, either preserving all the material or leaving some material out.

For every domain, we systematically vary irrelevant visual factors such as colors, object numbers, layouts, container shapes, directions, and transformations. This helps prevent models from solving the benchmark through memorized visual templates.

We then provide models with multiple frames sampled from each video and ask them to judge the final quantities. Importantly, some tasks can partly be solved by looking at the beginning and end, while others—particularly volume and size—really require understanding what happened during the transformation.

This lets us test not simply whether a model can perceive the final image, but whether it can integrate visual information over time and update an internal representation of the physical state.

Can you talk about the controlled experiments you carried out?

The controls are probably the most important part of the study because a high accuracy number by itself can be misleading. A central aspect of our study is that we do not treat benchmark accuracy as sufficient evidence of physical reasoning. Instead, we design a series of controlled experiments to distinguish genuine understanding of dynamic transformations from simpler strategies such as response priors, static visual cues, or insufficient temporal context.

Our primary control is the matched non-conserving condition. For every conserving example, we construct a closely matched counterpart in which the relevant quantity actually changes while the overall visual structure remains similar. A model that genuinely tracks the transformation should therefore perform well on both conditions. However, we observe the opposite pattern in many models: strong performance on conserving cases is often accompanied by poor performance on their non-conserving counterparts. This indicates that high conservation accuracy alone can substantially overestimate a model’s underlying reasoning ability.

We then systematically test several alternative explanations for this failure.

First, we vary the amount of temporal evidence available to the model by providing 3, 5, 7, 9, or 16 frames. If the primary limitation were incomplete observation of the transformation, performance should improve consistently as more frames are provided. In practice, increasing temporal resolution produces little systematic improvement.

Second, we examine whether the issue arises from suboptimal frame selection. We compare uniform sampling with frames selected by human annotators and by a learned video-localization model. Even when the model is provided with frames judged to be maximally informative, the underlying failure pattern largely persists.

Third, we evaluate whether more explicit reasoning instructions can compensate for the limitation. We test direct prompting, sequential frame-by-frame reasoning, chain-of-thought prompting, and prompts that explicitly emphasize temporal continuity across frames. While continuity-aware prompting provides modest gains in some settings, no prompting strategy consistently resolves the problem, and chain-of-thought can in some cases reduce performance.

We also conduct a particularly diagnostic set of no-visual-evidence controls, in which the original visual input is replaced by blank images or removed entirely. Under these conditions, models still exhibit a strong tendency to predict that quantities are conserved. This demonstrates that part of their apparent success in the standard setting is driven by a pre-existing response prior rather than by visual evidence.

More strikingly, on some conservation tasks, performance can actually improve when the informative visual content is removed. This suggests that the models are not simply failing to extract enough information from the video; rather, the visual evidence can conflict with a strong prior that the model is unable to appropriately revise.

Finally, we evaluate a range of additional controls, including explicit descriptions of conservation principles, textual captions of the visual sequence, an “I don’t know” response option, model scaling, and models specifically post-trained for physical reasoning. None of these interventions eliminates the core failure pattern.

Taken together, these experiments allow us to rule out several straightforward explanations—insufficient temporal coverage, poor frame selection, prompting, lack of declarative physical knowledge, or model scale—and instead point to a more fundamental limitation in how current vision-language models represent and update physical state over time.

What were your main findings?

The clearest finding is that current vision-language models do not reliably reason about physical transformations, even when the underlying task is extremely simple for humans.

Across 112 models, most systems were only modestly above the 33.3% chance level when performance was balanced across conserving and non-conserving scenarios. Human participants, by comparison, achieved approximately 98% accuracy.

A particularly important result is that conservation accuracy alone gives a misleading picture. There was actually a negative correlation between performance on conserving and non-conserving tasks: models that appeared very good at recognizing conservation were often particularly bad at noticing when the quantity really changed.

When we required a model to answer both members of a matched conserving/non-conserving pair correctly, 82 of the 112 models achieved below 10% accuracy. Only a handful of the strongest models performed above chance under this stricter criterion.

We also found very little evidence that simply scaling the models solves the problem. Performance on conservation tasks was almost unrelated to parameter count. Adding more frames, using more sophisticated prompts, or selecting more informative frames also failed to produce consistent improvements.
Our mechanistic analysis provides a possible explanation. In one model we studied in detail, incorrect non-conservation judgments were not hesitant guesses—they were often made with higher confidence than correct judgments. Attention analysis also showed that, during these failures, the model disproportionately focused on the initial frame and failed to sufficiently incorporate information appearing later in the transformation.

So our interpretation is that the fundamental problem is not simply that these models lack a fact such as “volume is conserved.” They often possess that linguistic knowledge. The harder problem is that they do not reliably maintain and update an object-state representation over time.

That distinction is important. A model can know the rule in language without being able to ground and apply that rule to a dynamic visual event.

What further work are you planning in this area?

There are several directions that we think are particularly important.

One is to move beyond these deliberately controlled experiments toward more realistic physical environments. Our current benchmark focuses on four basic quantities under relatively clean conditions. Real-world physical reasoning also involves occlusion, deformable objects, noisy observations, interactions between multiple objects, and transformations whose consequences may be uncertain.

A second direction is to understand the failure mechanistically. Our initial attention and confidence analyses suggest that models often fail to update their representation after observing a transformation, but this needs to be tested more systematically across different model architectures and with causal interventions.We are also interested in studying whether these seemingly simple conservation failures propagate into more consequential embodied tasks—for example, planning, manipulation, tool use, or predicting the consequences of actions. A model might perform well on a high-level robotics benchmark while still relying on shortcuts that fail under small changes in the environment.

More broadly, I think this work raises a question about the representations that future multimodal models should learn. Most current systems are extremely strong at extracting static semantic features from images. But physical reasoning may require something different: persistent, predictive representations of objects and their states that can be updated as the world changes. Developing models with that kind of temporally grounded internal representation is one of the directions we are particularly interested in pursuing.

About Yijiang Li

William Yijiang Li is a 2nd year PhD student and a Jacobs Fellow at UC San Diego. His research focuses on multi-modal LLM, Agents, World models and Embodiment, specifically the learning aspects of AI – to enable efficient (e.g. label efficiency, sample effiency, and synthetic data) and robust learningin multi-modal, interactive and 3D embodied environments. He has published over 10 top conference papers on these topics including ICML, Neurips, ICLR, NACCL, ACL, ICCV, ICLR, CVPR, and EMNLP, organized several workshops in leading conferences such as ES-Reasoning @ ICLR 2026, the 2nd Efficient-Reasoning @ COLM 2026, and the 2nd MMRAgI @ CVPR 2026, and co-founded the GrowAI, a research organization and community aiming to promote the research of AI that can learn and grow like humans.



tags: ,


Lucy Smith is Senior Managing Editor for AIhub.
Lucy Smith is Senior Managing Editor for AIhub.

            AUAI is supported by:



Subscribe to AIhub newsletter on substack



Related posts :

Forthcoming machine learning and AI seminars: October 2026 edition

  02 Oct 2026
A list of free-to-attend AI-related seminars that are scheduled to take place in the next couple of months.

Rebuilding the brain with neuromorphic computing: an interview with Oliver Rhodes

  01 Oct 2026
Neuromorphic computing takes inspiration from biology to build faster, more energy efficient systems.
monthly digest

AIhub monthly digest: September 2026 – tracking animal populations, recommender systems, and an interview with Ken Goldberg

  29 Sep 2026
Welcome to our monthly digest, where you can catch up with AI research, events and news from the month past.

AI-powered platforms uncover proteins that organise cellular compartments

Two platforms could enable researchers to more accurately predict proteins that undergo phase separation.

When compression techniques don’t just add up: Interaction effects in hybrid LLM compression

  25 Sep 2026
New research finds that combining common LLM compression techniques doesn't just add up — sometimes it backfires, sometimes it surprises.

What’s coming up at #IROS2026?

  24 Sep 2026
Find out what the International Conference on Intelligent Robots and Systems has in store.

IJCAI-ECAI 2026 tutorial / workshop round-up part 1

  23 Sep 2026
Find out more about two sessions on data-centric AI and a trustworthy agentic AI roadmap.


↑


AUAI is supported by:







Subscribe to AIhub newsletter on substack




 















©2026.05 - Association for the Understanding of Artificial Intelligence