2nd Edition — IEEE Intelligent Vehicles Symposium 2026
Vision, Language, and Multimodal Human Instructions
for Interactive Intelligent Vehicles
13:30–17:30 • LaSalle B • IEEE IV 2026 • VL-IIV 2026
About the Workshop
The Vision, Language, and Multimodal Human Instructions for Interactive Intelligent Vehicles (VL-IIV 2026) workshop explores the intersection of computer vision, language understanding, and multimodal reasoning for human-in-the-loop autonomous driving. The workshop focuses on systems and datasets that allow vehicles to perceive, interpret, and respond to visual and linguistic instructions.
Interactive autonomous systems capable of interpreting multimodal human instructions are critical to the next generation of safe and trustworthy transportation. This workshop promotes human-centered autonomy, reducing risks from fully unsupervised systems while enhancing transparency and user control.
This workshop was formerly organized under the name VALOR (Vision and Language Oriented Representations for Intelligent Vehicles / ITS), with a similar emphasis and scope on vision-language understanding in the context of human interactive driving. As the vision-language field has grown rapidly and many workshops have emerged addressing general vision-language and foundation model topics, we updated the title to more clearly emphasize what has always been the distinctive focus of this workshop: the interactions of humans with and within these AI-driven systems.
Topics
We welcome contributions with a strong focus on — but not limited to — the following topics:
Speakers
doScenes Instructed Driving Challenge
VL-IIV 2026 hosts the doScenes Instructed Driving Challenge
The challenge evaluates how well vision-language models predict trajectories conditioned on human driving instructions. The dataset contains scene-level captions, driver intent labels, and natural-language instructions for upcoming maneuvers — all human-generated and labeled by multiple annotators, creating a diverse set of descriptors mapping to the same maneuver.
Participants predict the vehicle's future trajectory conditioned on any combination of (1) visual scene input (multi-camera), (2) language instruction, and (3) scene context (history + map), evaluated using displacement error, visualization, and explainability.
View Challenge Details →Schedule
Half-day workshop, 13:30–17:30 — LaSalle B.
| 13:30 | Welcome |
Opening Remarks
Ross Greer & Mohan Trivedi
|
10 min |
| 13:40 | Invited Talk |
Building Trustworthy Physical AI Datasets with Agentic Evidence Gathering
Varun Krishnan — NomadicML
Abstract & BioPhysical AI companies collect enormous volumes of data, but identifying the moments that actually improve models remains a major bottleneck. We present Nomadic, an agentic data intelligence platform that combines VLMs with targeted evidence gathering to discover, validate, and explain complex events in robotics and autonomous driving datasets. By grounding model predictions in visual and motion evidence, Nomadic improves event accuracy, reduces manual review, and enables scalable curation of high-value training data for physical AI. Bio: Varun Krishnan is Cofounder / CTO at Nomadic AI. He was previously a Research Scientist on Lyft's Autonomous Navigation Team. |
40 min |
| 14:20 | Invited Talk |
Beyond Visual Question Answering: Context-Grounded LVLMs for Safer Transportation Perception
Abhijit Sarkar, PhD & Heesang Han — Virginia Tech Transportation Institute
AbstractLarge Vision-Language Models have created new opportunities for transportation scene understanding by allowing researchers and practitioners to query complex traffic scenes using natural language. However, transportation perception is not simply an image-understanding problem. Safety-relevant reasoning often depends on contextual information not fully captured by a single RGB frame, including temporal motion, 3D spatial structure, road-user interactions, environmental conditions, and visibility constraints. This talk presents recent work on context-grounded LVLMs for transportation through two complementary directions. The first is 3D spatial grounding, where 2D visual inputs are augmented with LiDAR-, stereo-, or tracking-derived information so that LVLMs can reason about depth, object position, motion, and vehicle-to-vehicle interaction — moving beyond visual description toward spatially informed traffic-scene reasoning critical for safety questions such as lane-change feasibility, vulnerable-road-user presence, and interaction risk. The second is concept grounding, where LVLMs are encouraged to reason through human-interpretable evidence rather than directly predicting a final label, making outputs more interpretable, auditable, and robust for safety-critical transportation applications. |
40 min |
| 15:00 | Invited Talk |
Systematizing the Unusual: A Taxonomy-Driven Dataset for Vision–Language Model Reasoning About Edge Cases in Traffic
Krzysztof Czarnecki — University of Waterloo
AbstractOne of the central challenges in developing robust vision–language models (VLMs) for real-world autonomy is their ability to recognize, interpret, and reason about rare and hazardous situations—so-called edge cases. Unlike routine traffic patterns, which are well-represented in large-scale datasets, these scenarios occur infrequently, are highly diverse, and often involve subtle contextual cues that challenge both object detection and semantic understanding. EdgeScenes is a new dataset under development at the WISE Lab that directly targets this limitation. It systematically captures and organizes rare road situations using a structured, ontology-driven taxonomy of hazardous conditions spanning infrastructure anomalies, abnormal road-user behavior, foreign objects, environmental extremes, and complex interactions. The dataset is constructed from crowdsourced video footage and annotated with a rich multimodal schema covering over 300 fine-grained hazardous conditions, along with temporal extents and contextual tags. As such, it establishes a testbed specifically designed for evaluating VLMs on out-of-distribution, safety-critical driving scenarios. This talk will present the motivation, design principles, and annotation methodology behind EdgeScenes, with particular emphasis on how taxonomy-based labeling enables systematic identification of edge cases and gaps. I will discuss key insights gathered during dataset construction, including patterns in real-world hazard emergence and challenges in visual–semantic grounding. Finally, I will report early experimental results using frontier VLMs, showing that while current models can detect a broad range of hazards beyond closed-vocabulary vision systems, they still exhibit severe limitations, showing both high false positive and false negative rates in recognizing road hazards. These findings point to an urgent need for progress in improving the reliability of VLMs to support trustworthy multimodal perception in autonomous systems. Speaker BioKrzysztof Czarnecki is a Professor of Electrical and Computer Engineering and a University Research Chair at the University of Waterloo, where he leads the Waterloo Intelligent Systems Engineering (WISE) Laboratory. He currently also serves as an Associate Director of the Waterloo Centre for Automotive Research (WatCAR). His research focuses on assuring the safety of AI systems and driving behavior. In 2018, he co-led the development of the first autonomous vehicle tested on public roads in Canada. He has made significant contributions to automotive AI and software safety standards, including SAE J3164 and ISO 8800. Before joining the University of Waterloo, he worked at DaimlerChrysler Research in Germany (1995–2002), where he advanced software development practices and technologies for enterprise, automotive, and aerospace sectors. His work has been recognized with numerous awards, including the Premier's Research Excellence Award (2004) and the British Computing Society’s Upper Canada Award for Outstanding Contributions to the IT Industry (2008). He has also received twelve Best Paper Awards, three ACM Distinguished Paper Awards, and five Most Influential Paper Awards. |
40 min |
| 15:40 | Break |
Coffee Break
|
15 min |
| 15:55 | Invited Talk |
Making of Trustworthy Autonomous Vehicles: a brief overview of an unfinished journey
Prof. Mohan Trivedi — University of California, San Diego
|
25 min |
| 16:20 | Challenge |
doScenes Instructed Driving Challenge — Overview
Angel Martinez, Parthib Roy
|
5 min |
| 16:25 | Challenge |
Language + History Track — 1st Place Presentation
Md Thamed Bin Zaman Chowdhury — UCF UrbanITY Lab
|
10 min |
| 16:35 | Oral Papers |
INSIGHT: Enhancing Autonomous Driving Safety through Vision-Language Models on Context-Aware Hazard Detection and Reasoning
Dianwei Chen, Zifan Zhang, Lei Cheng, Yuchen Liu, Xianfeng Yang
|
25 min |
| 17:00 | Closing |
Closing Remarks
Organizers
|
10 min |