WorldSimProbe

Diagnosing Simulator Faithfulness in Action-Conditioned World Models for Embodied Manipulation

Peterson Co1,2,3,*,†, Sicheng Hu1,2,*, Chunxuan Jiao1,2,*, Hongyang Cheng2, Yulin Luo1, Yijie Xu2,4 Sixiang Chen1, Zhongxia Zhao2, Zihao Wang5, DaFeng Chi3, Peidong Liu3, YuTong Chen2,6, Henghua Liu2,6 Zhihao Yuan3, Huizhu Jia1, Yuzheng Zhuang3, Tianle Zhang3, Liang Lin3, Huajie Tan2,†, Shanghang Zhang1,✉

1State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University, 2EvoPhys AI, 3Joy Future Academy, JD 4The University of Sydney, 5The Hong Kong University of Science and Technology, 6Beijing Institute of Technology

* Equal contribution. Project leaders. Corresponding author: shanghang@pku.edu.cn

Evaluation gap

A rollout can look right for the wrong reason.

Most ACWM evaluations focus on visual quality, task success, or coarse responsiveness, leaving two gaps: action realization is treated as a monolithic capability rather than a multi-dimensional process spanning local, global, and source-specific control variations, while interaction response is rarely checked for grounding in realized motion, valid contact, and correct dynamics. Yet a faithful simulator must preserve the full chain from supplied action to agent motion to environment response; otherwise, a rollout can look plausible or succeed while still ignoring controls or hallucinating unsupported interactions.

PCA view comparing task-conditioned, counterfactual, and source-diverse action trajectories
There exists a distribution gap between what simulators can do and what ACWMs are tested on (task distribution).

Failure modes

Where the causal chain breaks.

Action Calibration Failure

Action calibration is the alignment between the magnitude of an action change and its physical effect. Miscalibration shifts or obscures the failure threshold—the smallest perturbation beyond which task failure or rollout outcome change persists.

One standard. One chain.

A faithful simulator must preserve both links.

WorldSimProbe defines the Observable Simulator Contract, a single standard for simulator faithfulness: supplied actions must produce corresponding agent motion, and environment responses must be grounded in that motion. It traces this contract along one chain—from action realization to interaction response—through five controlled diagnostic suites.

Realized motion agent future
contact gate

≈ ΦR(r, e, a) ≈ ΦE(e, )

Interactive probes

Test the causal chain yourself.

Offline -- ms
Simulator Waiting
Connect to begin
Arm Left Command -- Steps 0
LingBot-VA Ready
Waiting for an action The selected model will run with the same trace.
Model ready

Operational overview

Five probes reveal where the chain fails.

Five diagnostic suites

Inspect each failure in aligned rollouts.

Benchmark results

One score cannot capture every failure.

LingBot-VA · RoboTwin
Five-suite diagnostic profile Radar chart comparing the current platform's top three models with the selected model.
Action realization Interaction response
Benchmark scores by model and diagnostic suite
Model T1 T2 T3 T4 T5

We’ve open-sourced the evaluation code and sample data. Run it on your model.

Citation

BibTeX
@misc{co2026worldsimprobe,
  title        = {{WorldSimProbe}: Diagnosing Simulator Faithfulness in Action-Conditioned World Models for Embodied Manipulation},
  author       = {Co, Peterson and Hu, Sicheng and Jiao, Chunxuan and Cheng, Hongyang and Luo, Yulin and Xu, Yijie and
                  Chen, Sixiang and Zhao, Zhongxia and Wang, Zihao and Chi, DaFeng and Liu, Peidong and Chen, YuTong and Liu, Henghua and
                  Yuan, Zhihao and Jia, Huizhu and Zhuang, Yuzheng and Zhang, Tianle and Lin, Liang and Tan, Huajie and Zhang, Shanghang},
  year         = {2026},
  eprint       = {2608.09298},
  archivePrefix = {arXiv},
  primaryClass = {cs.RO},
  url          = {https://arxiv.org/abs/2608.09298}
}

Operational overview of the five WorldSimProbe suites and their evaluators.