Autoregressive video · Prospective reasoning

ProARLearning Prospective Reasoning
with Autoregressive Video Models

Linghui Shen1·Tinghui Zhu2·Sheng Zhang3·Muhao Chen2

1 The Hong Kong Polytechnic University2 University of California, Davis3 Microsoft

ProAR transforms autoregressive video generation into a visual reasoning process through prospective guidance.

ProAR promotional teaser featuring shape sorting, FlowFree, maze navigation, and robotic manipulation scenes.

01 / The idea

From Reactive to
Prospective Generation

Video-native reasoning solves tasks by directly generating a sequence of visual states, with applications in physical reasoning and embodied intelligence.

Stable SortGround truth
Group shapes by type and order each group from smallest to largest.

Two requirements for reasoning

01

Goal completion

The final state must satisfy the task objective.

The right destination
02

Progression validity

Intermediate transitions must follow task logic or physical constraints.

Valid steps along the way

Standard autoregressive models react to past history through next-chunk prediction.
Locally plausible predictions therefore do not ensure goal completion or valid progression.

Standard AR predicts the next chunk reactively; ProAR anticipates future transitions and outcomes to guide generation toward task completion.
From reacting to the past to anticipating the future.

02 / The method

Learning Prospective Reasoning

Prospective reasoning anticipates future outcomes and state transitions from observed history and uses these expectations to guide current generation.

ProAR introduces two complementary mechanisms: outcome guidance and transition guidance.

ProAR architecture: history, current and goal streams; asymmetric attention for outcome guidance; training-only future representation alignment for transition guidance.
A shared backbone. Two complementary forms of guidance.

Outcome guidance

Goal Belief for
Long-Horizon Planning

“What should the generation achieve?”

ProAR jointly predicts the current chunk and an evolving goal frame. Asymmetric attention lets the goal belief guide current generation while shielding it from the noisy current chunk.

Explicit, sparse, long-range → Goal completion

Transition guidance

Representation Self-Alignment for
Progressive Reasoning

“What should happen next?”

Leveraging teacher forcing in AR training, ProAR extracts clean next-chunk representations in a single forward pass—without an external representation encoder or an additional backbone pass.
A lightweight predictor aligns current hidden states with these
future targets, encouraging them to anticipate upcoming temporal dynamics.

Late-start alignment. Transition guidance is activated after an initial training warm-up, once the backbone has learned task-relevant representations.

Implicit, dense, short-range → Progression validity

At inference ProAR predicts its own goal belief without ground-truth future frames; the alignment predictor is removed.
Overall inference time increases by only about 10% over standard AR.

03 / Results in action

Reasoning through video.

From perceptual, spatial, and abstract reasoning tasks to embodied simulation.

04 / Training efficiency

Better reasoning. Fewer training steps.

VBVR 10-task mean score versus training steps: ProAR exceeds the fully trained AR baseline at 2,500 steps.
Mean score on the VBVR 10-task subset.
25%

of the training steps

ProAR surpasses the fully trained AR baseline after only 2,500 training steps on the VBVR 10-task subset.

ProAR at 2.5k steps0.7079
AR at 10k steps0.6635

A further gain from transition guidance

Late start. Continued improvement.

After Transition guidance is activated at 7,500 steps, the mean score rises from 0.7666 to 0.8010 by step 10,000.

0.76667.5k · Transition starts
0.801010k · Final score

05 / Citation

BibTeX

If you find this work useful, please consider citing it.

@misc{shen_proar,
  title  = {ProAR: Learning Prospective Reasoning with
            Autoregressive Video Models},
  author = {Shen, Linghui and Zhu, Tinghui and
            Zhang, Sheng and Chen, Muhao}
}

Publication details will be updated with the public release.