Overview
VLALight represents a novel lightweight, end-to-end vision-language-action framework specifically designed for traffic signal control (TSC). This system aims to address challenges in urban congestion by enabling visual-context-aware TSC. It directly translates intersection observations and signal-phase information into discrete signal actions, moving away from intermediate image-to-text descriptions or handcrafted traffic-state representations.
The framework integrates multiple directional camera views into a unified visual input, utilizing textual instructions to establish correspondence between these views, traffic movements, and signal phases. This design facilitates direct action prediction with a compact 0.5 billion-parameter model.
Research Context
Traffic signal control is identified as essential for mitigating urban congestion. Recent advancements in vision-language models (VLMs) have enabled a richer interpretation of intersection scenes, creating new opportunities for visual-context-aware TSC. However, existing approaches often face limitations such as loose coupling between modules, repeated information conversion, and potential loss of fine-grained visual details. Sequential inference in these systems can also introduce substantial latency.
Approach
The VLALight framework was developed to overcome limitations associated with loose coupling, information conversion, and sequential inference in previous VLM-based TSC systems. Its design is characterized by an end-to-end architecture, directly mapping inputs to outputs.
- Input Integration: VLALight combines multiple directional camera views into a singular, unified visual input.
- Correspondence Establishment: Textual instructions are employed to establish the correspondence between these combined camera views, traffic movements, and signal phases.
- Direct Action Prediction: The framework directly predicts discrete signal actions based on the processed observations and signal-phase information.
- Model Size: The architecture incorporates a compact 0.5 billion-parameter model for this prediction.
- Mechanism Avoidance: It bypasses the need for intermediate image-to-text descriptions or the use of handcrafted traffic-state representations.
Findings
Experiments conducted with VLALight demonstrated specific performance advantages in emergency-aware traffic signal control:
- Emergency Vehicle Service: VLALight delivered the best emergency-vehicle service among all compared methods.
- Waiting Time Reduction: It reduced pooled emergency waiting time by 21.1% when compared to the cascaded VLMLight system.
- Real-time Operation: The system was capable of running in real time on local hardware.
- Generalization: VLALight demonstrated an ability to generalize to previously unseen intersection topologies and traffic-flow patterns.
Why This Matters
The development of VLALight directly addresses the critical need for more efficient urban traffic management, particularly concerning emergency response. By significantly reducing emergency vehicle waiting times and operating in real time, the system offers a pathway to improve urban safety and traffic flow dynamics without requiring extensive computational resources or pre-defined traffic states.