Overview
The World Embedding Benchmark (WEB) has been introduced to assess the encoding of physical information within video representations. The benchmark addresses the increasing focus on physical fidelity in world models and video generation by systematically investigating how video representations capture and process physical data.
Research Context
Recent advancements in world models and video generation have highlighted the importance of physical fidelity. However, the mechanisms by which video representations encode physical information have remained less understood. This gap necessitates tools and methodologies to quantitatively evaluate the physical understanding embedded within video models.
Approach
The World Embedding Benchmark comprises 8,000 controlled simulation cases, drawn from 80 distinct families. These cases span four primary domains of physics: fluid mechanics, solid mechanics, dynamics, and optics-electromagnetism. Each simulation case is structured to pair a rendered video with corresponding simulation-derived physical annotations. The benchmark facilitates three distinct tasks designed to evaluate different aspects of physical information encoding:
- Text-video retrieval: This task assesses the model's ability to retrieve relevant videos based on textual queries, indicating cross-modal physical alignment.
- Physical-property regression: This task measures the recoverability of quantitative physical information directly from video embeddings.
- Multiple-choice video-description pair classification: This task evaluates the model's capacity to distinguish correct video-description pairs, further exploring cross-modal physical alignment.
These tasks are designed to differentiate between cross-modal physical alignment and the recoverability of quantitative physical information. The study further investigated the utility of these embeddings in retrieval-augmented generation.
Findings
Evaluation of pre-trained omnimodal embedding models using the World Embedding Benchmark yielded several key observations:
- These models exhibited weak performance in text-video retrieval tasks.
- Their performance in within-family multiple-choice video-description pair classification was near chance levels.
- Despite these limitations, lightweight probes were capable of recovering useful physical information from frozen video embeddings, suggesting that some physical data is implicitly present.
- Continual contrastive training, specifically using physics-specific video-text pairs, resulted in improvements in both retrieval and pair classification tasks.
- Conversely, this same continual contrastive training led to a degradation in performance on the physical-property regression task. This finding revealed a trade-off between enhancing cross-modal physical alignment and maintaining the recoverability of quantitative physical information.
- The study also explored the application of these embeddings in retrieval-augmented generation (RAG) with MiniMax-H3. It was observed that retrieved references improved the physical fidelity of generated videos.
- Stronger retrieval models, when used in this RAG setup, yielded larger gains in the physical fidelity of generated videos in the experiments conducted.
Why This Matters
The findings from the World Embedding Benchmark highlight a critical need to jointly evaluate both physical alignment and quantitative property recoverability in video models. The demonstrated utility of physical representations for enhancing video generation suggests that improving how models encode and utilize physical information can lead to more physically accurate generated content.
Potential Applications
The study suggests that the developed embeddings can be used to retrieve reference videos for retrieval-augmented generation. This application has been shown to improve the physical fidelity of generated videos. The specific gains observed were larger when stronger retrieval models were employed in this context.