A Very Big Video Reasoning Suite

We bet on a future that video reasoning is the next fundamental intelligence paradigm, after language reasoning, where spatiotemporal embodied world experiences could be more naturally captured.

Data Engines

View All
gravity_physics
GitHub
Knowledge out-of-domain testset
A ball at height 14.9m with initial downward velocity 4.1 m/s falls under gravity 12.8 m/s² and bounces on ground with elasticity 0.80. Show the full trajectory with velocity arrows (direction and magnitude) updating throughout until the ball stops.
First Frame
Last Frame
construction_blueprint
GitHub
Abstraction in-domain testset
In the scene, the upper structure has a missing piece outlined with a dashed line. There are 4 candidate pieces below. The video sequentially checks each candidate from left to right: highlights the current candidate being examined with a frame, previews how the piece fits in the gap, marks it with ✓ if the shape matches or ✗ if it doesn't, then moves to the next candidate. Once the matching piece is found, an animation demonstrates it moving into the gap to complete the structure.
First Frame
Last Frame
key_door_matching
GitHub
Spatiality in-domain testset
The scene shows a maze with a green circular agent, colored diamond-shaped keys, and colored hollow rectangular doors. Find the Blue key and then navigate to the matching Blue door, showing the complete movement process step by step.
First Frame
Last Frame
symbol_delete
GitHub
Transformation out-of-domain testset
Delete red ◆ at position 3. The animation shows the target symbol fading out, then the remaining symbols shifting left to close the gap.
First Frame
Last Frame
stable_sort
GitHub
Perception in-domain testset
The scene contains two types of shapes, each type has three shapes of different sizes arranged randomly. Keep all shapes unchanged in appearance (type, size, and color). Only rearrange their positions: first group the shapes by type, then within each group, sort the shapes from smallest to largest (left to right), and arrange all shapes in a single horizontal line from left to right.
First Frame
Last Frame

Inference Results

View Full Bench
Mirror Reflection - Samples
00
01
02
03
04
Task Domains 1/5
Mirror Reflection
Knowledge in-domain testset
Symmetry Completion
Abstraction out-of-domain testset
Select Leftmost Shape
Spatiality out-of-domain testset
Rotation Puzzle
Transformation in-domain testset
Majority Color
Perception in-domain testset
Prompt
Loading...
Ground Truth
First
First Frame
Final
Final Frame
Model Outputs
1/
VBVR-Wan2.2
VBVR-Wan2.2
CogVideoX 1.5
Kling 2.6
LTX-2
Runway Gen-4
Sora 2
Veo 3
Wan 2.2 I2V
Hunyuan I2V
Seedance 2.0

Leaderboard

Modality
Split
Type
Category
2026-04-28