Phi-Zero is a video world model built on what the authors call physical language. The idea is to reverse the order of two steps that usually happen together.
Instead of generating the next frames right away, the model first reasons about how everything should move and change physically. Only then does that reasoning go through the video generator, which renders the frames.
Why this matters
- The separation greatly improves prediction of what happens next.
- In physical coherence and understanding, it outperforms comparable world models on average.
- It serves three distinct fronts: interactive worlds, autonomous driving, and robot training.
1.Why separate reasoning from rendering
2 minA typical video generator predicts pixels from pixels. It learns that certain images tend to follow others, without ever representing why — and that's why it slips up on physics: objects going through walls, liquid flowing upward, shadows that don't track.
Phi-Zero inserts a step before: describing what should physically happen. Rendering becomes a translation of a decision already made, not a guess. It's the same principle that makes a reasoning model get more math answers right than one that responds directly.
Spec sheet
- Status
- Code "coming soon"
- Approach
- Physical language before rendering
- Applications
- Interactive worlds, autonomous driving, robots

Continue in the full microcourse
You've read the opening of 3 classes
The microcourse covers the complete step-by-step, the selection criteria, where the tool fails, who it's really for — and, in the Expert version, the official address to start today.
- The three applications2 min
- Current state1 min
How we verified
We track releases straight from primary sources, transcribe what's demonstrated, check every name and number against the manufacturer's official documentation, and rewrite it in Portuguese — with what the tool no do it together, which is the part the ad leaves out.


