Bind symbols to the scene
Identifies the physical quantities and experiment-relevant objects in the annotated first frame before any law is invoked.
Benchmarking Thinking with Video
Towards Law-Grounded Physical Intelligence
1S-Lab, Nanyang Technological University ยท 2The Chinese University of Hong Kong
Built on the 400-video Orchard mechanics dataset, Apple-π evaluates Perception, Formulation, and Deduction through chain-of-frames traces, combining MLLM-based judging with physics-law-grounded metrics to diagnose where physical reasoning breaks.
Apple-π asks the question existing physical-video benchmarks leave unresolved: when a model generates plausible motion, has it actually invoked the governing law, or merely learned an intuitive visual prior?
We turn Newton-style scientific reasoning into an auditable video protocol: a model must perceive physical quantities, formulate the governing law, and deduce law-consistent dynamics over time.
Apple-π pairs an infographic-style first frame with a chain-of-frames prompt, then evaluates the response through five subtracks that expose where physical reasoning breaks.
Identifies the physical quantities and experiment-relevant objects in the annotated first frame before any law is invoked.
Tests whether the model selects the symbolic equation and predicts a law-grounded target state at a specified time.
Generates the full video sequence and compares the resulting trajectory frame by frame with physics-derived ground truth.
Orchard organizes 400 videos across 10 classical mechanics tasks for clean single-law diagnosis and multi-law composition.
Scenarios are organized around explicit classical-mechanics laws, so failures map back to concrete physical principles.
Controlled cases reduce confounders and localize failures across perception, formulation, and deduction.
Composition tasks test whether models preserve physical states while switching between laws.
Apple-π reports Avg. together with track-wise, pillar-wise, and source-wise scores, so each ranking exposes not only which model wins, but where physical reasoning breaks: across perception, formulation, deduction, physical-law family, and sim-to-real transfer under controlled law-grounded evaluation.
Best video model trails the top overall model, exposing a gap between video synthesis and law-grounded reasoning.
P-T/P-G denote Perception-Text/Graphic; F-T/F-G denote Formulation-Text/Graphic; Ded. denotes Deduction.
Apple-π turns the leaderboard into diagnosis: the important signal is not only who scores higher, but which stage, law family, and visual source breaks the physical reasoning chain.
Unified models dominate Perception and Formulation, suggesting that explicit understanding and controllable generation help bind annotations, objects, equations, and target states.
The stage radar compares Video avg., Unified avg., and the strongest model from each family; even the best unified model stays near 0.40 on Deduction, revealing a temporal dynamics bottleneck.
Scores drop along the protocol: models often read local cues, partially formulate laws, and then fail when those laws must be rolled forward.
The reasoning funnel makes the decline explicit: earlier success is necessary but not sufficient; good perception does not guarantee law-consistent temporal dynamics.
Single-law pillars are easier than Multi because multi-law cases require one law's output state to become the next law's initial condition.
The pillar heatmap contrasts single-law tasks with composition: current models learn isolated motion priors, but struggle to carry position, velocity, direction, and contact across law transitions.
Both video and unified models perform worse on real-world cases than simulated cases, even though the governing laws are unchanged.
The source-transfer slope shows the Sim-to-Real drop, which mainly reflects failures in visual grounding, tracking, and robust law execution under realistic appearance variation.
Future video world models need more than plausible motion: they need explicit physical understanding, richer reasoning data, and post-training that rewards law-consistent dynamics over time across stages, laws, and sources.
Open a case to inspect the first frame, the track-specific prompts for unified and video models, the physics-derived ground truth, and model outputs across five visible reasoning tracks.
@misc{yao2026applepibenchmarkingthinkingvideo,
title={Apple-$\pi$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence},
author={Runmao Yao and Kairui Hu and Yukang Cao and Ruisi Wang and Shulin Tian and Ziang Cao and Weichen Fan and Ziqi Huang and Yuhao Dong and Hao Li and Zhaoxi Chen and Zhongang Cai and Lei Yang and Ziwei Liu},
year={2026},
eprint={2607.16401},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2607.16401},
}