Imagine you are driving toward an intersection. A truck parked on the corner blocks your view of a car speeding toward the intersection from the right. You can’t see it, but the car next to you can. What if that neighboring vehicle could simply tell your car,
“Vehicle merging from the right, about 15 meters away, moving quickly.”
Your car now knows something its own sensors could never observe.
This simple idea of cars helping each other by describing what they see in natural language motivates UNCAP, our new framework for cooperative autonomous driving. But there’s an important twist: not every neighboring vehicle is worth listening to, and not every message deserves the same level of trust.
UNCAP teaches autonomous vehicles both who to talk to as well as how much to trust what they hear.
Researchers have since long envisioned fleets of autonomous vehicles cooperating to make driving safer. Blind corners, occlusions, and hidden pedestrians will become much less dangerous if nearby cars can share what they observe.
The traditional approach is straightforward: cars can exchange raw sensor data such as camera images, LiDAR point clouds, or radar measurements. However, this quickly becomes impractical to be implemented at scale for a variety of reasons. Raw sensor streams consume enormous bandwidth, require significant computation to process stream data, and assume that every vehicle has comparable sensing hardware. In reality, with the rise of autonomous driving, the roads are instead observing a mixture of manufacturers, sensor configurations, and heterogenous software stacks. Natural language offers an intriguing alternative to sensor data sharing. Instead of transmitting megabytes of images, a vehicle might simply say, “pedestrian crossing behind bus” or “motorcycle approaching from the left, fast-moving”. These messages are tiny, hardware-agnostic, and immediately understandable by modern vision-language models across the spectrum.
But language introduces a new problem. At a busy intersection, a vehicle might have ten or twenty neighbors. Should it listen to all of them? And if two vehicles disagree, whose description should it believe? UNCAP answers all these questions!
An appeal of UNCAP is that it works entirely in a zero-shot manner. Instead of training a new model (which is resource intensive), UNCAP has a “Bring-your-own-model” strategy. It builds on existing vision-language models and adds a communication strategy that allows vehicles to cooperate intelligently. The UNCAP framework includes four stages.
UNCAP helps each vehicle decide which neighbors matter. Those messages describe nearby vehicles, motion, and confidence, which UNCAP combines before making a driving decision.
Each vehicle periodically broadcasts a tiny “heartbeat” of information containing only its position and direction of travel. This heartbeat consists of only a few bytes of data. The aim is to provide just enough context for nearby vehicles to understand who is around them.
On the road, most nearby vehicles are irrelevant. A car traveling away from the intersection provides little useful information, whereas a vehicle approaching the same intersection may offer exactly the observations we need. UNCAP reasons about which neighbors are both close enough and likely to affect the vehicle’s future path. This is crucial as it dramatically reduces the bandwidth of unnecessary communication as well as the computation required by downstream stages. It enables UNCAP to scale to handle large teams.
Relevant vehicles then describe what they see using natural language, accompanied by calibrated confidence estimates. Calibration is desirable as it produces reliable uncertainty estimates. UNCAP uses conformal prediction, which is a lightweight calibration method that ensures confidence values reflect the actual reliability of the information being received. Every vehicle then reasons about whether an incoming message actually reduces its uncertainty about the scenario. To answer this, UNCAP measures the mutual information contributed by each incoming observation. Mutual information ensures that the vehicle can give greater weightage to messages that make the scene clearer, while uninformative or potentially misleading observations are discounted. Rather than blindly treating every neighbor homogenously, each vehicle builds an uncertainty-aware understanding of the environment.
Finally, the scene description generated by reasoning about all incoming messages is passed to an off-the-shelf vision-language model, which produces a driving decision such as merge, wait, or yield, along with a corresponding confidence score.
We evaluated UNCAP in the popular CARLA driving simulator using the OPV2V benchmark, which is a widely used benchmark by autonomous driving researchers. OPV2V includes highway merging, urban intersections, and other scenarios where vehicles frequently experience occlusions.
These bird’s-eye-view (BEV) examples compare driving with and without UNCAP. With selected helper messages, vehicles behave more proactively in merging and intersection scenes.
Compared with non-communicating vehicles, and compared with prior language-based approaches that blindly broadcast every observation to every other car, UNCAP consistently improved both efficiency and safety. Over various scenarios, our experiments yielded key qualitative gains of UNCAP over traditional cooperation schemes include:
UNCAP avoids broadcasting every observation to every nearby vehicle. It sharply reduces bandwidth while preserving the information needed for a driving decision.
Perhaps the most surprising result was that sharing raw images actually performed worse than sharing language.
Modern vision-language models appear to struggle when jointly reasoning over multiple camera feeds captured from different viewpoints. Translating those observations into concise natural-language descriptions first seems to act as a form of structured reasoning, distilling the essential information before it reaches the decision-making model.
Equally encouraging, the improvements were consistent across several underlying vision-language models, including GPT-4o, GPT-4.1, GPT-4o-mini, and GPT-5. This suggests that the benefits stem from the communication strategy itself rather than from any particular model.
UNCAP is an early step toward autonomous vehicles that cooperate more like human drivers.
Today’s implementation has important limitations. Our experiments are limited to simulation, so real-world constraints such as noisy sensing, unreliable wireless communication, network delays and packet losses remain to be explored. Finally, because querying today’s vision-language models takes roughly a second, decisions are made at key moments rather than continuously.
That being said, the broader message extends well beyond autonomous driving.
As AI systems increasingly operate in teams, whether fleets of robots, drones, or software agents, the challenge is no longer simply making each agent smarter. The critical need of the hour is to ensure that each agent in the team learns how to communicate efficiently, decide which information to trust, and build a shared understanding of the world.
In a nutshell, UNCAP suggests that natural language, paired with principled reasoning about uncertainty, can provide a surprisingly effective foundation for that future.
UNCAP: Uncertainty-Guided Neurosymbolic Planning Using Natural Language Communication for Cooperative Autonomous Vehicles, Neel P. Bhatt, Po-han Li, Kushagra Gupta, Rohan Siva, Daniel Milan, Alexander T. Hogue, Sandeep P. Chinchali, David Fridovich-Keil, Zhangyang Wang, Ufuk Topcu.
This work was nominated for the AAMAS 2026 best paper award.