Technology

World models, not video models

Primate Intelligence builds Darwin, a JEPA-based predictive world model for video.

Generative video models learn to render pixels; world models learn what happens next and why — the actual substrate of understanding. The difference is architectural. Generative models optimize for visual plausibility; world models optimize for predictive accuracy in representation space.

Prediction in representation space — not pixel space — is what makes Darwin fast, deterministic, and deployable at the edge. There is no decoder generating hallucinated frames, no language model sampling stochastically at the perception layer. Just a model that watches video the way a systems engineer needs: deterministic, timestamped, confidence-scored, queryable in plain English.

The full 'World models, not video models' essay is coming during our August launch.

Explore the technology