Apple-π

Benchmarking Thinking with Video
Towards Law-Grounded Physical Intelligence

Runmao Yao1,*, Kairui Hu1,*, Yukang Cao1, Ruisi Wang1, Shulin Tian1,
Ziang Cao1, Weichen Fan1, Ziqi Huang1, Yuhao Dong1, Hao Li1, Zhaoxi Chen1, Zhongang Cai1, Lei Yang2, Ziwei Liu1,โœ‰

1S-Lab, Nanyang Technological University ยท 2The Chinese University of Hong Kong

*Equal contribution ยท โœ‰Corresponding author

Nanyang Technological University The Chinese University of Hong Kong
TL;DR

Apple-π is the first benchmark that asks video models to reason through explicit physical laws, turning motion generation into an auditable test of physical intelligence.

Built on the 400-video Orchard mechanics dataset, Apple-π evaluates Perception, Formulation, and Deduction through chain-of-frames traces, combining MLLM-based judging with physics-law-grounded metrics to diagnose where physical reasoning breaks.

01 The thesis

Beyond visual plausibility, toward law-grounded physical intelligence.

Apple-π asks the question existing physical-video benchmarks leave unresolved: when a model generates plausible motion, has it actually invoked the governing law, or merely learned an intuitive visual prior?

We turn Newton-style scientific reasoning into an auditable video protocol: a model must perceive physical quantities, formulate the governing law, and deduce law-consistent dynamics over time.

Apple-pi teaser contrasting plausible video with law-grounded deduction
02 Benchmark protocol

Make generated video a visible reasoning trace.

Apple-π pairs an infographic-style first frame with a chain-of-frames prompt, then evaluates the response through five subtracks that expose where physical reasoning breaks.

Apple-pi benchmark protocol with five reasoning subtracks
(A) Perception

Bind symbols to the scene

Identifies the physical quantities and experiment-relevant objects in the annotated first frame before any law is invoked.

(B) Formulation

Formulate governing laws

Tests whether the model selects the symbolic equation and predicts a law-grounded target state at a specified time.

(C) Deduction

Roll the law through time

Generates the full video sequence and compares the resulting trajectory frame by frame with physics-derived ground truth.

03 Orchard Dataset

A law-first orchard of physical scenarios.

Orchard organizes 400 videos across 10 classical mechanics tasks for clean single-law diagnosis and multi-law composition.

Apple-pi Orchard taxonomy of classical mechanics tasks
400videos
10mechanics tasks
3sources
Law-first taxonomy

Mechanics before miscellany

Scenarios are organized around explicit classical-mechanics laws, so failures map back to concrete physical principles.

Clean diagnosis

Single-law cases isolate errors

Controlled cases reduce confounders and localize failures across perception, formulation, and deduction.

Compositional transfer

Multi-law cases test chaining

Composition tasks test whether models preserve physical states while switching between laws.

Data-source composition

Diverse sources for controlled diagnosis.

Simulated Self-recorded Internet-sourced
Free Fall
2536
61
Projectile Motion
3159
90
Inclined Plane
24
24
Circular Motion
26
26
Elastic Collision
2615
41
Perfectly Inelastic Collision
2510
35
Inelastic Collision
2516
41
At Rest
2511
36
Uniform Linear
25
25
Composition
1110
21
04 Leaderboard

Leaderboard for law-grounded reasoning.

Apple-π reports Avg. together with track-wise, pillar-wise, and source-wise scores, so each ranking exposes not only which model wins, but where physical reasoning breaks: across perception, formulation, deduction, physical-law family, and sim-to-real transfer under controlled law-grounded evaluation.

Overall ranking Avg. score
Video gap 0.000

Best video model trails the top overall model, exposing a gap between video synthesis and law-grounded reasoning.

Video models Avg. score
Unified models Avg. score
Diagnostic view

P-T/P-G denote Perception-Text/Graphic; F-T/F-G denote Formulation-Text/Graphic; Ded. denotes Deduction.

05 Results

Four findings on physical reasoning failures.

Apple-π turns the leaderboard into diagnosis: the important signal is not only who scores higher, but which stage, law family, and visual source breaks the physical reasoning chain.

Finding 1

Understanding-before-generation raises the ceiling.

Unified models dominate Perception and Formulation, suggesting that explicit understanding and controllable generation help bind annotations, objects, equations, and target states.

The stage radar compares Video avg., Unified avg., and the strongest model from each family; even the best unified model stays near 0.40 on Deduction, revealing a temporal dynamics bottleneck.

Finding 2

Reasoning narrows from perception to deduction.

Scores drop along the protocol: models often read local cues, partially formulate laws, and then fail when those laws must be rolled forward.

The reasoning funnel makes the decline explicit: earlier success is necessary but not sufficient; good perception does not guarantee law-consistent temporal dynamics.

Finding 3

Multi-law composition exposes weak state transfer.

Single-law pillars are easier than Multi because multi-law cases require one law's output state to become the next law's initial condition.

The pillar heatmap contrasts single-law tasks with composition: current models learn isolated motion priors, but struggle to carry position, velocity, direction, and contact across law transitions.

Finding 4

Real-world videos reveal a persistent grounding gap.

Both video and unified models perform worse on real-world cases than simulated cases, even though the governing laws are unchanged.

The source-transfer slope shows the Sim-to-Real drop, which mainly reflects failures in visual grounding, tracking, and robust law execution under realistic appearance variation.

Key takeaway

Future video world models need more than plausible motion: they need explicit physical understanding, richer reasoning data, and post-training that rewards law-consistent dynamics over time across stages, laws, and sources.

06 Case Studies

Inspect ten law-grounded reasoning cases.

Open a case to inspect the first frame, the track-specific prompts for unified and video models, the physics-derived ground truth, and model outputs across five visible reasoning tracks.

Cite this work
@misc{yao2026applepibenchmarkingthinkingvideo,
      title={Apple-$\pi$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence},
      author={Runmao Yao and Kairui Hu and Yukang Cao and Ruisi Wang and Shulin Tian and Ziang Cao and Weichen Fan and Ziqi Huang and Yuhao Dong and Hao Li and Zhaoxi Chen and Zhongang Cai and Lei Yang and Ziwei Liu},
      year={2026},
      eprint={2607.16401},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2607.16401},
}